7 Commits
Author SHA1 Message Date
Daniel Samson 24c49f56e1 kernel: fix pci-scan triple fault and flaky restart-drill checks
pciScanTest put three [64]DeviceDescriptor arrays on the stack; at 328
bytes each that is a 66,096-byte frame, larger than the 64 KiB bootstrap
stack the boot context runs on, so the prologue overflowed the stack and
triple-faulted on entry before the test could print anything — the CPU
reset with no exception line. Reuse one descriptor buffer (frame ~21 KiB).

With the fault gone, the restart-drill checks proved flaky. They polled
process.write_buffer (which holds only the latest write-syscall message)
for the transient restarting/re-scan lines, which the concurrent
class-driver output overwrote before the starved low-priority poller could
observe them. The kernel broker count can't witness the restart either:
register is idempotent and the table has no unregister, so the count is
invariant across kill/prune/respawn.

Follow the sibling restart drills (usbReportTest/driverRestartTest): assert
the ordered drill in the harness regex over the full serial log (a
backreference requires the respawn to re-scan the same count), and keep only
race-free broker checks in the kernel — empty before the scan, populated
once it settles, never growing past N — sampled via scheduler.sleep at a
fixed cadence instead of a busy-poll.
2026-07-14 13:48:33 +01:00
Daniel Samson 0217662808 display: resilient scanout — supervised driver, survive loss, re-attach (v2 V6)
The compositor now survives the virtio-gpu driver dying and re-attaches when device-manager
restarts it — the last piece of display v2.

Three parts:

- The driver hellos the device manager (role: bus). It never did, so the manager — which
  spawns it supervised and expects a hello — was stopping it at the 3s hello deadline every
  run (the gate markers just printed first). Now it is properly supervised: not stopped for
  silence, and restarted on death.

- A kernel IPC fix so a call to a dead service errors instead of hanging. An endpoint records
  its owner; when that task dies, its registered endpoints are marked dead (and any parked
  senders woken with -EPEER), so ipc_call returns -EPEER rather than blocking on a reply that
  will never come. Without this the compositor's first present after the driver died blocked
  forever. General robustness — any client of any service benefits.

- The compositor re-attaches. Its .scanout calls now fail cleanly (caught), freezing the last
  frame; when the restarted driver re-announces, attach_scanout detects the backend is already
  virtio and logs "scanout re-attached", mapping the fresh shared surface and re-looking-up
  .scanout. (The previous shm mapping leaks — no shm_unmap syscall yet — but its frames are the
  dead driver's, reclaimed on exit.)

- device-manager gains a "test-scanout-restart" mode (like test-usb-restart) that kills the
  virtio-gpu driver once after it hellos; the displayReattachTest kernel scenario drives it.

Gate: python3 test/qemu_test.py display-reattach — "scanout upgraded to virtio-gpu" then
"scanout re-attached", no CPU exception, passing 3/3. host tests, ipc/ipc-call/ipc-cap,
supervision, shm, display-service, display-demo, virtio-gpu, display-native, and
display-modeset all pass; default zig build clean. v2 (V1-V6) complete.
2026-07-14 13:10:21 +01:00
Daniel Samson 4231301896 display: runtime mode-setting, EDID, and fenced (vsync) present (v2 V5)
The native backend can now change resolution and presents tear-free.

Mode-setting without churn. The driver sizes its scanout resource + shared surface to the
largest mode it offers and treats a mode change as re-pointing the scanout rectangle within
that surface — so the resource, its backing, and the shared mapping never change, and the
surface's row stride (the max width) is fixed while the active width/height move. The
compositor is handed that stride in the announce and composes at it; a smaller mode just
paints the top-left rectangle. This sidesteps the surface re-share a true resolution change
would otherwise need (the service harness can't reply with a capability).

- scanout-protocol gains get_modes + set_mode; the driver offers {640x480, 800x600} and
  re-points set_scanout on set_mode.
- backend.VirtioGpu carries the surface stride, exposes modes()/setMode(), and reports
  canModeSet = hasVsync = true.
- runtime.display gains modes()/setMode() (display-protocol get_modes/set_mode, forwarded to
  the backend) — the client-facing API.
- EDID: the driver negotiates VIRTIO_GPU_F_EDID when the device offers it, reads the monitor's
  EDID, and logs its preferred mode (parsed from the first detailed timing descriptor).
- vsync: every resource_flush is fenced (VIRTIO_GPU_FLAG_FENCE); the device signals the fence
  when the frame is on screen, which the used-ring ack the synchronous present already waits
  on gates — so a completed present is a tear-free one.

After the native upgrade the compositor runs a one-shot mode-set self-check: query the modes,
switch to a different one, re-composite, and confirm the backend reports the new geometry —
the gate's markers.

Gate: python3 test/qemu_test.py display-modeset (reuses the display-native boot) — "display:
mode set to 800x600, verified" + "display: vsync present ok", passing 3/3. host tests,
display-service, display-demo, shm, virtio-gpu, and display-native still pass.
2026-07-14 12:44:21 +01:00
Daniel Samson 58927ed7e5 display: hot-attach a virtio-gpu native backend over the GOP floor (v2 V4)
The compositor now boots on the GOP framebuffer and upgrades to the virtio-gpu driver
the moment it announces itself — the pluggable-scanout payoff.

The shared surface. The scanout resource is an shm region the driver creates
(shm_physical, a new syscall, hands it the guest-physical for attach_backing) and passes
to the compositor as a capability. The compositor maps it and composites straight into
it: on x86 DMA is cache-coherent, so the cacheable shared pages the CPU paints are exactly
what the device transfers-and-flushes — no copy, no explicit flush.

The handshake. After bring-up the driver looks up .display and sends attach_scanout with
the geometry + the surface capability. The compositor maps the surface, looks up the
driver's .scanout endpoint itself (the driver registered it — no need to pass it), switches
to backend.VirtioGpu, and re-composites the current frame. present() over the native
backend is a present request on .scanout -> transfer-to-host + resource flush. The first
native present is deferred to a one-shot timer: presenting inline from the announce handler
would deadlock, since the driver is still blocked on our reply and not yet serving .scanout.
After it lands, the compositor reads a pixel back from the shared surface to confirm the
frame reached the device's backing.

- shm_physical (syscall 36) + runtime.shm.physical.
- scanout-protocol (the compositor->driver present channel), separate from the
  client-facing display protocol; the display protocol gains attach_scanout.
- backend.VirtioGpu joins backend.Gop in the tagged union; select() still boots GOP.
- the virtio-gpu driver's scanout backing is now shm (was DMA); it announces + serves
  .scanout present requests (transfer-to-host + flush of the shared surface).

Also fixes a latent framebuffer-geometry corruption the display service hit only when it
enumerated the device tree alongside a busy device-manager: Gop.init now captures the
geometry into a small value the instant device_enumerate returns (rather than re-reading
the 328-byte descriptor across the later claim/mmio_map syscalls) and retries on a zero
geometry. The underlying device-table clobber is a separate kernel bug, tracked apart.

Gate: python3 test/qemu_test.py display-native (QEMU -device virtio-gpu-pci) — "display:
scanout upgraded to virtio-gpu" + "display: native present verified" + "display-demo: ok",
passing 3/3. host tests, display-service, display-demo, shm, and virtio-gpu still pass.
2026-07-14 12:21:37 +01:00
Daniel Samson 6e0e0a62c6 virtio-gpu: a modern virtio 1.0 display driver, bring-up to a flushed frame (v2 V3)
The first native scanout backend's driver half. The device manager matches the
virtio-gpu PCI function by its Display/Other class triple and spawns the driver,
which claims the device, confirms vendor 0x1AF4/device 0x1050 from config space,
enables memory-space + bus mastering, and walks the virtio vendor capabilities to
find the common-config and notify structures in its BAR.

From there it is the standard modern-virtio bring-up: reset, negotiate VERSION_1,
stand up the control virtqueue in coherent DMA, then drive the GPU end to end —
RESOURCE_CREATE_2D, ATTACH_BACKING (a coherent DMA buffer; V4 swaps in the shm-
shared surface), SET_SCANOUT, paint a known pattern, TRANSFER_TO_HOST_2D,
RESOURCE_FLUSH, and wait for the device's used-ring ack. Reading the backing back
proves it is CPU-visible RAM; the ack proves the device consumed the frame —
together the automated stand-in for "it's on screen", no screenshot.

- system/drivers/virtio-gpu/: the driver, plus virtio-gpu-protocol.zig (control
  commands) and virtio-pci.zig (the 1.0 PCI transport + split-virtqueue), both with
  host-tested struct sizes.
- device-manager matches the display/other class triple to "virtio-gpu"; the driver
  self-confirms the vendor/device id, since the class alone cannot distinguish it.
- ServiceId.scanout (11): the driver registers it so the compositor finds it in V4.

Gate: python3 test/qemu_test.py virtio-gpu (QEMU -device virtio-gpu-pci) — the
device-manager stack discovers the function, the driver brings up a 640x480 scanout
and flushes a test pattern: "virtio-gpu: scanout 640x480 online" + "flush acked,
pixel check ok". host tests, display-service, shm, and device-list still pass.

Note: the pre-existing pci-scan case triple-faults on main (verified at 88ad432,
before this change); it is unrelated and tracked separately.
2026-07-14 11:29:53 +01:00
Daniel Samson 88ad432758 kernel: shm cross-process shared memory capability (v2 V2)
Generalize capability passing from endpoints to memory objects. The per-task
handle table now holds kind-tagged entries (scheduler.HandleObject{kind, ptr});
closeHandles and shareCapability dispatch by kind, so a shared-memory object
rides an ipc_call send_cap exactly like an endpoint and is refcount-freed only
when its last capability drops.

- shm_create(len) -> vaddr, handle: contiguous, zeroed, cacheable frames wrapped
  in a refcounted ShmObject, mapped into the caller's shm arena (PML4[230]).
- shm_map(cap) -> vaddr: map the same physical pages into a receiver that got the
  capability. mapUserSharedInto maps WB-cacheable + device_grant, so a sharer's
  teardown never frees the shared frames — the object owns them.
- runtime.shm: create(len) -> Region{ptr, handle, len}, map(handle) -> ptr.

Gate: qemu_test.py shm — shm-client creates a region, writes a pattern, passes
its capability to shm-server, which maps it and reads the same bytes back
(shm: shared 4096 bytes ok). ipc/ipc-call/ipc-cap/supervision/dma/usermem/
display-service and host tests all still pass — the handle change broke no IPC.
2026-07-14 10:57:14 +01:00
Daniel Samson 9333d0572f display: pluggable scanout backend seam (v2 V1)
Extract scanout from the compositor into backend.zig — a `Backend` tagged union
with info()/surface()/present(damage) and canModeSet/hasVsync flags. The v1 GOP
path becomes `backend.Gop` (claim the display node, WC-map the LFB, keep the
cacheable back buffer, present = the damage-rect WC copy); display.zig now
composes into backend.surface() and calls backend.present(damage), with no LFB or
framebuffer geometry left in the compositor core. The selection decision is the
pure chooseKind(native_available), split from the syscall-bound bring-up, ready
for the native-if-present branch at V4.

Pure refactor: display-service, display-demo, and zig build test all pass
unchanged.
2026-07-14 09:40:15 +01:00
25 changed files with 2307 additions and 293 deletions
+18
View File
@@ -355,6 +355,13 @@ pub fn build(b: *std.Build) void {
});
runtime_module.addImport("display-protocol", display_protocol_module);
// The scanout protocol: the compositor's outbound present channel to a native scanout
// driver (virtio-gpu), separate from the client-facing display protocol (docs/display-v2.md).
const scanout_protocol_module = b.addModule("scanout-protocol", .{
.root_source_file = b.path("system/services/display/scanout-protocol.zig"),
});
runtime_module.addImport("scanout-protocol", scanout_protocol_module);
// The power protocol: system power's domain-named surface (docs/power.md).
const power_protocol_module = b.addModule("power-protocol", .{
.root_source_file = b.path("system/services/power/protocol.zig"),
@@ -468,6 +475,9 @@ pub fn build(b: *std.Build) void {
const fat_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat", "system/services/fat/fat.zig");
const display_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display", "system/services/display/display.zig");
const display_demo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display-demo", "system/services/display-demo/display-demo.zig");
const virtio_gpu_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "virtio-gpu", "system/drivers/virtio-gpu/virtio-gpu.zig");
const shm_server_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "shm-server", "system/services/shm-server/shm-server.zig");
const shm_client_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "shm-client", "system/services/shm-client/shm-client.zig");
const fat_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat-test", "system/services/fat/fat-test.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
// The PCI bus driver decodes each function's class triple to human names in its
@@ -539,6 +549,12 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(display_exe.getEmittedBin());
mk_run.addArg("display-demo");
mk_run.addFileArg(display_demo_exe.getEmittedBin());
mk_run.addArg("virtio-gpu");
mk_run.addFileArg(virtio_gpu_exe.getEmittedBin());
mk_run.addArg("shm-server");
mk_run.addFileArg(shm_server_exe.getEmittedBin());
mk_run.addArg("shm-client");
mk_run.addFileArg(shm_client_exe.getEmittedBin());
mk_run.addArg("pci-bus");
mk_run.addFileArg(pci_bus_exe.getEmittedBin());
mk_run.addArg("crash-test");
@@ -775,6 +791,8 @@ pub fn build(b: *std.Build) void {
"system/services/fat/engine.zig", // FAT read/write over a RAM-backed image
"system/services/display/compositor.zig", // Rect math + fill/composite/blit-tile
"system/services/display/protocol.zig", // pack(): native pixel encoding per format
"system/drivers/virtio-gpu/virtio-gpu-protocol.zig", // virtio-gpu command struct sizes
"system/drivers/virtio-gpu/virtio-pci.zig", // virtio 1.0 PCI transport struct sizes
}) |root| {
const mod_tests = b.addTest(.{
.root_module = b.createModule(.{
+3 -2
View File
@@ -73,8 +73,9 @@ rather than restate it. Roughly in the order things happen at runtime:
track: a user-space compositor that owns the framebuffer, composes a layer stack into
a double buffer, and presents it. Why GOP and the PCI display device are two views of
one controller, the device-node + write-combining handoff, and what flicker-free buys
that tear-free doesn't. Plan: [display-plan.md](display-plan.md). **v2** makes scanout
a pluggable backend (GOP floor + a native virtio-gpu driver, hot-attached):
that tear-free doesn't. Plan: [display-plan.md](display-plan.md). **v2** (complete) makes
scanout a pluggable backend — GOP floor + a native virtio-gpu driver, hot-attached, with
runtime mode-set, EDID, fenced vsync presents, and restart re-attach:
[display-v2.md](display-v2.md), plan [display-v2-plan.md](display-v2-plan.md).
20. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
+100 -66
View File
@@ -40,91 +40,125 @@ used ring**. Those two together (pixel-readback + flush-ack) are the automated s
---
## V1 — The scanout backend seam (refactor, no behaviour change)
## V1 — The scanout backend seam (refactor, no behaviour change) ✅
Extract scanout from the compositor so today's path becomes one backend among future ones.
- [ ] A `Backend` interface in `system/services/display/`: `surface() -> {ptr, pitch,
format, width, height}`, `present(damage: Rect)`, and capability flags
(`canModeSet`, `hasVsync`, both false for now).
- [ ] Wrap the v1 GOP path as `GopBackend` (claim the `display` node, WC-map the LFB,
`present` = the current damage-rect WC copy). The compositor composes into
`backend.surface()` and calls `backend.present(damage)` — no direct LFB references
left in the compositor core.
- [ ] Pure backend-selection logic factored so it's host-testable.
- [x] `system/services/display/backend.zig`: a `Backend` tagged union with `info()`,
`surface()` (the cacheable compose target), `present(damage)`, and capability flags
(`canModeSet`/`hasVsync`, both false for GOP).
- [x] The v1 GOP path is now `backend.Gop` (claims the `display` node, WC-maps the LFB,
keeps the cacheable back buffer, `present` = the damage-rect WC copy). display.zig
composes into `backend.surface()` and calls `backend.present(damage)` — no LFB or
framebuffer geometry left in the compositor core.
- [x] The selection decision is the pure `chooseKind(native_available)` (gop unless a
native driver announced), split from the syscall-bound `select()`/`Gop.init()`.
**Gate:** `display-service` + `display-demo` still pass unchanged (pure refactor; GOP is
the only backend), and `zig build test` stays green.
**Gate (met):** `display-service` + `display-demo` pass **unchanged** (pure refactor; GOP
is the only backend), and `zig build test` stays green.
## V2 — The `shm` cross-process memory capability (kernel)
## V2 — The `shm` cross-process memory capability (kernel) ✅
- [ ] [abi.zig](../system/abi.zig): `shm_create`, `shm_map` syscalls (+ a `ServiceId`/cap
convention if needed). Kernel handlers: `shm_create(len)` allocates page-aligned RAM,
returns a handle + maps it; passing the handle as an `ipc_call` `send_cap` shares it;
`shm_map(cap)` maps the same physical pages into the receiver. Reclaimed on death.
- [ ] `library/runtime/shm.zig` (+ barrel export): `create(len) -> Region{handle, ptr}`,
`map(cap) -> ptr`.
- [ ] Reuse the M13 capability-passing machinery (endpoints → memory objects).
- [x] [abi.zig](../system/abi.zig): `shm_create` (34) / `shm_map` (35) syscalls + a
`shm_test` service id. Handlers in process.zig: `shm_create(len)` allocates contiguous,
zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability
handle, maps them into the caller's shm arena → returns vaddr + handle; `shm_map(cap)`
maps the same physical pages into the receiver. Reclaimed on death (see below).
- [x] The capability core (ipc-synchronous.zig) is now **kind-tagged**: `scheduler.Task`'s
handle table holds `HandleObject{kind, ptr}`; `closeHandles` and `shareCapability`
dispatch by kind, so an `ShmObject` rides an `ipc_call` `send_cap` exactly like an
endpoint and frees only when its last capability drops. `mapUserSharedInto` (paging)
maps WB-cacheable + `device_grant`, so a sharer's teardown never frees the shared
frames — the object owns them.
- [x] `library/runtime/shm.zig` (+ barrel export): `create(len) -> Region{ptr, handle, len}`,
`map(handle) -> ptr`.
**Gate:** a kernel/qemu `shm` test — process A `shm_create`s a region, writes a pattern,
passes the cap to process B, which `shm_map`s it and reads the same bytes back (proving
shared physical pages, not a copy). Heartbeat `shm: shared N bytes ok`.
**Gate (met):** `python3 test/qemu_test.py shm` — `shm-client` creates a region, writes a
pattern, and passes its capability to `shm-server` as an `ipc_call` send_cap; the server
`shm_map`s it and reads the **same bytes** back → `shm: shared 4096 bytes ok`. Guardrail:
`ipc`/`ipc-call`/`ipc-cap`, `supervision`, `dma`, `usermem`, `display-service`, and host
tests all still pass — the handle-table change broke no existing IPC.
## V3 — The virtio-gpu driver: bring-up + a frame on screen
## V3 — The virtio-gpu driver: bring-up + a frame on screen ✅
- [ ] `system/drivers/virtio-gpu/`: claim the virtio-gpu PCI function (device-manager
match on its PCI/virtio id), map BARs, negotiate features, set up the control
virtqueue. `virtio-gpu-protocol.zig` for the control structs (host-tested sizes).
- [ ] Create a 2D scanout resource backed by an `shm` region, `attach_backing`,
`set_scanout` to CRTC 0, and `resource_flush` a test pattern.
- [ ] Register a `scanout` service (new `ServiceId`).
- [x] `system/drivers/virtio-gpu/`: claim the virtio-gpu PCI function (device-manager
match on the display/other class triple, driver self-confirms vendor 0x1AF4/device
0x1050 from config space), enable memory-space + bus-master, walk the vendor
capabilities in config space to find common-config + notify, map the BAR, negotiate
VERSION_1, and stand up the control virtqueue in coherent DMA. `virtio-gpu-protocol.zig`
+ `virtio-pci.zig` for the control/transport structs (host-tested sizes).
- [x] Create a 2D scanout resource backed by a coherent DMA region (V4 swaps this for the
shm-shared surface), `attach_backing`, `set_scanout` to scanout 0, `transfer_to_host_2d`
+ `resource_flush` of a test pattern, and wait on the used ring.
- [x] Register a `scanout` service (`ServiceId.scanout` = 11).
**Gate (automated):** a `virtio-gpu` case (QEMU `-device virtio-gpu`) where the driver
writes a known test pattern into the shm scanout resource, `resource_flush`es it, and
**waits for the device's used-ring ack**, then reads the resource back and checks the
pattern — logging `virtio-gpu: scanout {w}x{h} online` and `virtio-gpu: flush acked, pixel
check ok`. That proves virtqueue + resource + attach + set_scanout + flush end to end
without a screenshot (the used-ring ack is the device confirming it consumed the frame).
**Gate (met):** the `virtio-gpu` case (QEMU `-device virtio-gpu-pci`) boots the
device-manager stack, which discovers the function and spawns the driver; the driver writes
a known test pattern into the scanout backing, `transfer_to_host_2d` + `resource_flush`es
it, and **waits for the device's used-ring ack**, then reads the backing back and checks the
pattern — logging `virtio-gpu: scanout 640x480 online` and `virtio-gpu: flush acked, pixel
check ok`. That proves virtqueue + resource + attach + set_scanout + transfer + flush end to
end without a screenshot (the used-ring ack is the device confirming it consumed the frame).
## V4 — The native backend + hot-attach
## V4 — The native backend + hot-attach ✅
- [ ] `VirtioGpuBackend` in the compositor: `surface()` = the shared scanout resource,
`present(damage)` = `resource_flush` of the damaged rect.
- [ ] The driver **announces** to `.display` (looks it up, sends *attach-scanout* with its
`scanout` endpoint + the shared surface as capabilities). The compositor switches
backends and re-presents the current frame full-screen.
- [ ] Boot still starts on `GopBackend`; the upgrade happens on announce.
- [x] `backend.VirtioGpu` in the compositor: `surface()` = the shared `shm` scanout surface
(the compositor composes straight into the device's resource backing; x86 DMA is
coherent, so the cacheable shared pages need no flush), `present(damage)` = a `present`
request over the driver's `.scanout` endpoint (→ transfer-to-host + resource flush).
- [x] The driver **announces** to `.display` after bring-up (looks it up with a bounded retry,
sends `attach_scanout` with the geometry + the shared surface as an `ipc_call` send_cap).
The compositor maps it, looks up `.scanout` itself (no need to pass the endpoint — the
driver registered it), switches backend, and re-composites the current frame full-screen.
The present is deferred to a one-shot timer so it runs *after* the reply unblocks the
driver and it serves `.scanout` — presenting inline would deadlock.
- [x] Boot still starts on `backend.Gop`; the upgrade happens on announce. `shm_physical` (a
new syscall) gives the driver the guest-physical of the shared surface for `attach_backing`.
**Gate (automated):** boot with virtio-gpu + `display-demo`; the compositor logs
`display: scanout upgraded to virtio-gpu`, drives frames through the native backend, and
**reads a pixel back** from the shared scanout resource after a present to confirm the
composited frame landed (`display: native present verified`), while `display-demo: ok`
still fires. Without `-device virtio-gpu`, no `scanout` is announced and it stays on GOP —
the v1 `display-service`/`display-demo` gates still pass unchanged.
**Gate (met):** the `display-native` case (QEMU `-device virtio-gpu-pci`, `mem` bumped since it
boots the whole system) starts the compositor + `display-demo` + device-manager; the driver
announces, the compositor logs `display: scanout upgraded to virtio-gpu`, drives frames through
the native backend, and **reads a pixel back** from the shared surface after a present to
confirm the composited frame landed (`display: native present verified`), while `display-demo:
ok` still fires — checked order-independently. Without `-device virtio-gpu-pci` nothing is
announced and it stays on GOP: the v1 `display-service`/`display-demo` gates pass unchanged.
## V5 — Mode-setting, EDID, and vsync
## V5 — Mode-setting, EDID, and vsync ✅
- [ ] virtio-gpu `GET_EDID` → a mode list; `set_scanout` at a chosen mode = runtime
resolution change. `runtime.display` gains `modes()` / `setMode(m)`.
- [ ] A vsync/fenced `resource_flush` present path → genuinely tear-free.
- [ ] The compositor reports the native backend's `canModeSet`/`hasVsync` = true.
- [x] The driver negotiates `VIRTIO_GPU_F_EDID` (when offered) and reads the monitor's EDID,
logging its preferred mode; it offers a small mode list over `.scanout` `get_modes`. The
resource + shared surface are sized to the largest mode, so `set_mode` just re-points the
scanout rectangle (no resource/surface churn) — a runtime resolution change. `runtime.display`
gains `modes()` / `setMode()` (display-protocol `get_modes`/`set_mode`, forwarded to the backend).
- [x] Every `resource_flush` is issued fenced (`VIRTIO_GPU_FLAG_FENCE`); the device signals the
fence when the frame is on screen, which the used-ring ack the synchronous present waits on
already gates — a tear-free present.
- [x] `backend.VirtioGpu` reports `canModeSet` / `hasVsync` = true.
**Gate (automated):** a `display-modeset` case reads the EDID mode list, calls `setMode`
to a different resolution, and confirms the change by reading the driver's scanout geometry
back (`display: mode set to {w}x{h}, verified`); the vsync/fenced present path is exercised
and confirmed by the flush **fence completing** (`display: vsync present ok`) — both from
serial, no eyeballing.
**Gate (met):** the `display-modeset` case (reusing the display-native boot) upgrades to
virtio-gpu, queries the driver's modes, `setMode`s to a different resolution, and confirms the
change by reading the backend's geometry back (`display: mode set to {w}x{h}, verified`); the
fenced present path is exercised and confirmed (`display: vsync present ok`) — both from serial,
passing 3/3. The driver also logs the EDID preferred mode (`virtio-gpu: EDID preferred mode …`).
## V6 — Resilience (restart + re-attach) + tests + docs
## V6 — Resilience (restart + re-attach) + tests + docs ✅
- [ ] The virtio-gpu driver is supervised (device-manager / init) and restartable; on
driver loss the compositor freezes the last frame and **re-attaches** when the driver
re-announces. Only a permanent give-up (crash-loop cap) attempts GOP again.
- [ ] `test/qemu_test.py` cases: `virtio-gpu`, hot-attach, `display-modeset`, and a
driver-kill/re-attach case. Update [display-v2.md](display-v2.md) status; README index.
- [x] The virtio-gpu driver now **hellos** the device manager (role: bus) so it is properly
supervised — no longer stopped at the hello deadline — and is restarted on death. On
driver loss the compositor keeps the last frame (its `.scanout` calls now return
`-EPEER` instead of hanging — a kernel fix: an endpoint is marked dead when its owner
dies) and **re-attaches** when the restarted driver re-announces. A permanent give-up
(crash-loop cap) leaves the frozen frame; GOP is not re-taken.
- [x] `test/qemu_test.py`: the `virtio-gpu`, `display-native` (hot-attach), `display-modeset`,
and `display-reattach` (driver-kill/re-attach) cases. display-v2.md status updated.
**Gate:** kill the virtio-gpu driver mid-run; the compositor survives and re-attaches on
restart (`display: scanout re-attached`); all v1 + v2 cases pass; default `zig build` clean.
**Gate (met):** the `display-reattach` case — device-manager (in `test-scanout-restart` mode)
kills the virtio-gpu driver once after it hellos; the restart policy respawns it, it
re-announces, and the compositor logs `display: scanout re-attached` after the initial
`display: scanout upgraded to virtio-gpu`, with no CPU exception / panic (the compositor
survives) — passing 3/3. All v1 + v2 cases (host tests, `ipc`/`ipc-call`/`ipc-cap`,
`supervision`, `shm`, `display-service`, `display-demo`, `virtio-gpu`, `display-native`,
`display-modeset`) pass; default `zig build` is clean.
---
+5
View File
@@ -1,5 +1,10 @@
# The display service v2: a pluggable scanout backend
**Status: complete (V1–V6).** The compositor boots on the GOP framebuffer and, when a
virtio-gpu driver announces itself, hot-attaches a native backend over the shared `shm`
scanout surface — with runtime mode-setting, EDID, and fenced (vsync) presents, and it
re-attaches across driver restarts. All serial-gated (see [display-v2-plan.md](display-v2-plan.md)).
v1 ([display.md](display.md)) is a compositor that owns the **GOP framebuffer** — it
composites a layer stack into a cacheable back buffer and streams damage to the linear
framebuffer the firmware handed over. That path is portable and good: it drives any GPU,
+27
View File
@@ -59,6 +59,33 @@ pub fn present() bool {
return transact(.{ .operation = @intFromEnum(protocol.Operation.present) }, &reply);
}
/// One selectable display mode.
pub const Mode = protocol.Mode;
/// Fill `out` with the resolutions the display can switch to; returns how many were written
/// (zero on the GOP floor, or if the service never came up).
pub fn modes(out: []Mode) usize {
const h = service() orelse return 0;
var request = protocol.Request{ .operation = @intFromEnum(protocol.Operation.get_modes) };
var reply: [protocol.modes_reply_size]u8 = undefined;
const len = ipc.call(h, std.mem.asBytes(&request), &reply) catch return 0;
if (len < protocol.modes_reply_size) return 0;
const answer = std.mem.bytesToValue(protocol.ModesReply, reply[0..protocol.modes_reply_size]);
if (answer.status != 0) return 0;
const count = @min(@min(answer.count, protocol.max_modes), out.len);
for (0..count) |i| out[i] = answer.modes[i];
return count;
}
/// Change the display resolution. Only a native backend that supports mode-setting honours it
/// (on the GOP floor it returns false); on success the display's `info()` reports the new mode.
pub fn setMode(width: u32, height: u32) bool {
var reply: protocol.Reply = undefined;
const changed = transact(.{ .operation = @intFromEnum(protocol.Operation.set_mode), .width = width, .height = height }, &reply);
if (changed) mode = null; // the cached mode is stale now
return changed;
}
/// The mode, cached after the first `info()` so `color()` doesn't round-trip per pixel.
var mode: ?Info = null;
+8
View File
@@ -38,6 +38,11 @@ pub const device = @import("device.zig");
/// DMA-capable memory for drivers: contiguous, pinned, uncacheable buffers.
pub const dma = @import("dma.zig");
/// Shared cacheable memory: create a region + capability, pass the capability to another
/// process (an `ipc_call` send_cap), map the same pages there. See library/runtime/shm.zig
/// and docs/display-v2.md.
pub const shm = @import("shm.zig");
/// USB class-driver client: open a device on the xHCI bus and drive it
/// (control / interrupt / bulk transfers). See library/runtime/usb.zig.
pub const usb = @import("usb.zig");
@@ -51,6 +56,9 @@ pub const block = @import("block.zig");
pub const display = @import("display.zig");
/// The display wire protocol (shared with the display service and its clients).
pub const display_protocol = @import("display-protocol");
/// The scanout wire protocol: the compositor's present channel to a native scanout driver
/// (virtio-gpu). See system/services/display/scanout-protocol.zig and docs/display-v2.md.
pub const scanout_protocol = @import("scanout-protocol");
/// The danos-native file API (open/read/write/list over the user-space VFS) — the
/// layer danos programs use directly, and where the operations that later become
+57
View File
@@ -0,0 +1,57 @@
//! User-space shared memory: `shm_create` / `shm_map`. A process creates a shareable,
//! zeroed, cacheable RAM region and gets back a pointer plus a **capability handle**; it
//! passes that handle to another process as an `ipc_call` send_cap, and the receiver
//! `shm_map`s it to map the same physical pages. The kernel primitive under the display
//! compositor↔native-driver and app↔compositor surface paths (docs/display-v2.md). The
//! generalization of capability passing from endpoints to memory objects.
const abi = @import("abi");
const sc = @import("system-call.zig");
const ipc = @import("ipc.zig");
inline fn failed(r: usize) bool {
return r > ~@as(usize, 0) - 4095; // a wrapped -errno lands in the top page
}
/// A shared region: the `ptr` the CPU touches, and the `handle` (a capability) to hand to
/// another process as an `ipc_call` send_cap.
pub const Region = struct {
ptr: [*]u8,
handle: ipc.Handle,
len: usize,
};
/// Grant `len` bytes (rounded up to whole pages) of shareable, zeroed, cacheable RAM.
/// Returns the region or null on failure. Two return values — vaddr in rax, handle in rdx —
/// so this is a hand-written stub like `dma.alloc`.
pub fn create(len: usize) ?Region {
var rax: usize = undefined;
var rdx: usize = undefined; // out: the capability handle
asm volatile ("syscall"
: [rax] "={rax}" (rax),
[rdx] "={rdx}" (rdx),
: [n] "{rax}" (@intFromEnum(abi.SystemCall.shm_create)),
[a0] "{rdi}" (len),
: .{ .rcx = true, .r11 = true, .memory = true });
if (failed(rax)) return null;
return .{ .ptr = @ptrFromInt(rax), .handle = rdx, .len = len };
}
/// Map the shared region named by a capability `handle` this process received (via an
/// `ipc_call` send_cap) into its address space — the same physical pages the creator sees.
/// Returns the pointer, or null on failure.
pub fn map(handle: ipc.Handle) ?[*]u8 {
const r = sc.systemCall1(.shm_map, handle);
if (failed(r)) return null;
return @ptrFromInt(r);
}
/// The guest-physical base of the shared region named by `handle` (which this process must
/// hold a capability for). The region's frames are contiguous, so this single address plus
/// the region length is all a device needs — e.g. a virtio-gpu driver programming an
/// `attach_backing`. Returns null on failure.
pub fn physical(handle: ipc.Handle) ?usize {
const r = sc.systemCall1(.shm_physical, handle);
if (failed(r)) return null;
return r;
}
+5
View File
@@ -60,6 +60,9 @@ pub const SystemCall = enum(u64) {
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
klog_read = 32, // klog_read(offset, ptr, len) -> bytes copied: copy the kernel RAM log buffer out to a user buffer (for persisting the boot log to disk)
wall_clock = 33, // wall_clock() -> Unix epoch seconds (UTC): the RTC wall-clock time, for filesystem timestamps (mtime). Monotonic time is `clock`.
shm_create = 34, // shm_create(len) -> vaddr (rax), handle (rdx): a shareable, zeroed, cacheable RAM region mapped into this AS; the handle is a capability passed to another process as an ipc_call send_cap (docs/display-v2.md)
shm_map = 35, // shm_map(cap) -> vaddr: map the shared region named by a received capability into this AS (the same physical pages the creator sees)
shm_physical = 36, // shm_physical(cap) -> paddr: the guest-physical base of a shared region held by capability, so a driver can program it into a device (e.g. virtio-gpu attach_backing); the pages are contiguous (docs/display-v2.md)
_,
};
@@ -184,6 +187,8 @@ pub const ServiceId = enum(u32) {
block = 7, // a block-device driver (USB mass storage today): read/write of fixed-size blocks, the storage a filesystem sits on
fat = 8, // the FAT filesystem server; the VFS mounts it and forwards paths under its mount point (/mnt/usb) to it
display = 9, // the display service: owns the framebuffer, composites a layer stack, presents frames (docs/display.md)
shm_test = 10, // the shm test server (V2): a client passes it a shared-memory capability, it maps + verifies (docs/display-v2.md)
scanout = 11, // a native scanout driver (virtio-gpu): the compositor finds it here to upgrade off the GOP framebuffer (docs/display-v2.md)
_,
};
@@ -0,0 +1,139 @@
//! The virtio-gpu control protocol — the command/response structs the driver exchanges with
//! the device over its control virtqueue (virtio spec, "GPU Device"). `extern` structs, so
//! the layout matches the little-endian wire format exactly. Host-tested for size. See
//! docs/display-v2.md.
const std = @import("std");
/// Control command / response types (virtio_gpu_ctrl_type). Commands are 0x01xx, responses
/// 0x11xx (ok) / 0x12xx (error).
pub const CmdType = enum(u32) {
get_display_info = 0x0100,
resource_create_2d = 0x0101,
resource_unref = 0x0102,
set_scanout = 0x0103,
resource_flush = 0x0104,
transfer_to_host_2d = 0x0105,
resource_attach_backing = 0x0106,
resource_detach_backing = 0x0107,
get_edid = 0x010a,
resp_ok_nodata = 0x1100,
resp_ok_display_info = 0x1101,
resp_ok_edid = 0x1104,
resp_err_unspec = 0x1200,
_,
};
/// Set in a command's `flags` to request a fence; the device echoes `fence_id` in the
/// response and does not report completion until the command's effects are visible.
pub const flag_fence: u32 = 1 << 0;
/// VIRTIO_GPU_F_EDID — device feature bit 1 (the low feature word): the device answers the
/// `get_edid` command. Negotiate it only when the device offers it.
pub const feature_edid: u32 = 1 << 1;
/// virtio_gpu_ctrl_hdr — the header on every command and response.
pub const CtrlHdr = extern struct {
type: u32,
flags: u32 = 0,
fence_id: u64 = 0,
ctx_id: u32 = 0,
ring_idx: u8 = 0,
padding: [3]u8 = .{ 0, 0, 0 },
};
pub const Rect = extern struct {
x: u32,
y: u32,
width: u32,
height: u32,
};
/// 2D pixel formats. QEMU's virtio-gpu host default is B8G8R8X8 (matches our bgrx).
pub const format_b8g8r8x8_unorm: u32 = 2;
pub const format_r8g8b8x8_unorm: u32 = 134;
pub const ResourceCreate2d = extern struct {
hdr: CtrlHdr,
resource_id: u32,
format: u32,
width: u32,
height: u32,
};
/// One scatter-gather entry of a resource's guest backing (a physical span).
pub const MemEntry = extern struct {
addr: u64,
length: u32,
padding: u32 = 0,
};
/// Header for RESOURCE_ATTACH_BACKING; `nr_entries` `MemEntry` follow it inline.
pub const ResourceAttachBacking = extern struct {
hdr: CtrlHdr,
resource_id: u32,
nr_entries: u32,
};
pub const SetScanout = extern struct {
hdr: CtrlHdr,
rect: Rect,
scanout_id: u32,
resource_id: u32,
};
pub const ResourceFlush = extern struct {
hdr: CtrlHdr,
rect: Rect,
resource_id: u32,
padding: u32 = 0,
};
/// Copy the guest backing into the host resource for `rect` (2D resources must transfer
/// before a flush shows the update).
pub const TransferToHost2d = extern struct {
hdr: CtrlHdr,
rect: Rect,
offset: u64,
resource_id: u32,
padding: u32 = 0,
};
pub const max_scanouts = 16;
pub const DisplayOne = extern struct {
rect: Rect,
enabled: u32,
flags: u32,
};
pub const RespDisplayInfo = extern struct {
hdr: CtrlHdr,
pmodes: [max_scanouts]DisplayOne,
};
pub const GetEdid = extern struct {
hdr: CtrlHdr,
scanout: u32,
padding: u32 = 0,
};
pub const RespEdid = extern struct {
hdr: CtrlHdr,
size: u32,
padding: u32 = 0,
edid: [1024]u8,
};
test "virtio-gpu struct sizes match the wire layout" {
try std.testing.expectEqual(@as(usize, 24), @sizeOf(CtrlHdr));
try std.testing.expectEqual(@as(usize, 16), @sizeOf(Rect));
try std.testing.expectEqual(@as(usize, 40), @sizeOf(ResourceCreate2d));
try std.testing.expectEqual(@as(usize, 16), @sizeOf(MemEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(ResourceAttachBacking));
try std.testing.expectEqual(@as(usize, 48), @sizeOf(SetScanout));
try std.testing.expectEqual(@as(usize, 48), @sizeOf(ResourceFlush));
try std.testing.expectEqual(@as(usize, 56), @sizeOf(TransferToHost2d));
try std.testing.expectEqual(@as(usize, 24 + 4 + 4 + 1024), @sizeOf(RespEdid));
}
+622
View File
@@ -0,0 +1,622 @@
//! /system/drivers/virtio-gpu — the virtio-gpu (virtio 1.0, modern PCI) display driver.
//! The device manager spawns it for the display/other PCI function (class 0x0380) whose
//! config space says vendor 0x1AF4 / device 0x1050; this instance claims that device and
//! brings up a single 2D scanout.
//!
//! V3 (this increment): the whole path end to end, proven from serial without a screenshot.
//! Claim the function, map its config space (resource 0) and the BAR that carries the
//! virtio structures, walk the vendor capabilities to find common-config / notify, reset
//! and negotiate VERSION_1, stand up the control virtqueue in DMA memory, then drive the
//! GPU: RESOURCE_CREATE_2D → ATTACH_BACKING (a coherent DMA buffer) → SET_SCANOUT, paint a
//! known test pattern, TRANSFER_TO_HOST_2D → RESOURCE_FLUSH, and **wait for the device's
//! used-ring ack**. Reading the backing back confirms it is CPU-visible; the ack confirms
//! the device consumed the frame. The compositor backend, hot-attach, mode-set/EDID, and
//! restart/re-attach are V4–V6. See docs/display-v2.md.
const std = @import("std");
const runtime = @import("runtime");
const mmio = @import("mmio");
const device = runtime.device;
const dma = runtime.dma;
const shm = runtime.shm;
const system = runtime.system;
const ipc = runtime.ipc;
const dp = runtime.display_protocol;
const sp = runtime.scanout_protocol;
const dm = runtime.device_manager_protocol;
const vp = @import("virtio-pci.zig");
const vg = @import("virtio-gpu-protocol.zig");
/// The DisplayFormat (device-abi) our B8G8R8X8 scanout resource presents: bgrx = 1. Handed to
/// the compositor in the announce so it packs colours in the surface's byte order.
const display_format_bgrx: u32 = 1;
/// The PCI vendor/device ids of a modern virtio-gpu (Red Hat / virtio; GPU is a
/// virtio-1.0-only device, so the id is always the modern 0x1050 — no legacy variant).
const virtio_vendor: u16 = 0x1AF4;
const virtio_gpu_device: u16 = 0x1050;
/// The scanout resource + shared surface are sized to the *largest* mode we offer; a mode
/// change (V5) re-points the scanout rectangle within it, so the resource, its backing, and
/// the shared surface never churn — and the surface's row stride is always `max_width`, which
/// the compositor is told in the announce. Kept modest so the backing is an easy contiguous run.
const max_width: u32 = 800;
const max_height: u32 = 600;
const scanout_bytes: usize = @as(usize, max_width) * max_height * 4;
const resource_id: u32 = 1;
/// The modes this scanout offers (all ≤ max). The first is the mode it comes up in.
const Mode = struct { width: u32, height: u32 };
const offered_modes = [_]Mode{ .{ .width = 640, .height = 480 }, .{ .width = 800, .height = 600 } };
/// The active mode — the scanout rectangle within the max-sized surface. Changed by `set_mode`.
var current_width: u32 = offered_modes[0].width;
var current_height: u32 = offered_modes[0].height;
/// Monotonic fence id for fenced (vsync) flushes; the device signals the fence when the flush
/// is complete, which its used-ring ack already gates our synchronous present on.
var fence_next: u64 = 1;
/// Whether the device offered VIRTIO_GPU_F_EDID, so `get_edid` is worth issuing.
var edid_available = false;
/// The control virtqueue. We drive it synchronously — one command, notify, poll the used
/// ring — so a depth of 16 is ample; we ask the device to shrink to it (virtio 1.0 lets the
/// driver reduce queue_size), keeping the whole ring inside one page.
const queue_size: u16 = 16;
const desc_offset: usize = 0; // 16 * 16 = 256 bytes
const avail_offset: usize = 256; // flags + idx + ring[16] + used_event = 38 bytes
const used_offset: usize = 1024; // flags + idx + ring[16] + avail_event = 134 bytes
/// The command scratch: the request the device reads, then its response, in one DMA page.
const request_offset: usize = 0;
const response_offset: usize = 2048;
var device_id: u64 = 0;
// Mapped virtio structures (virtual addresses into the device's BAR).
var common_base: usize = 0;
var notify_base: usize = 0;
var notify_multiplier: u32 = 0;
var notify_addr: usize = 0;
// Per-BAR mapping cache: several capabilities usually share one BAR, and mmio_map must not
// be asked to map the same resource twice.
var bar_virtual: [6]usize = .{ 0, 0, 0, 0, 0, 0 };
// DMA memory: the virtqueue rings and the command scratch.
var ring: dma.Region = undefined;
var command: dma.Region = undefined;
// The scanout backing is a **shared** (shm) region, not DMA: cacheable so the compositor
// composites into it cheaply (x86 DMA is coherent, so the device still sees the writes), and
// shareable so the same physical pages the device scans out of are the ones the compositor
// paints. The driver keeps the capability to hand to the compositor in the announce.
var surface: shm.Region = undefined;
// Split-virtqueue producer/consumer shadows.
var avail_shadow: u16 = 0;
var used_shadow: u16 = 0;
/// Format one whole log line and emit it in a single `write`, so this driver's output can
/// never interleave mid-line with the other drivers the manager runs concurrently.
fn log(comptime fmt: []const u8, arguments: anytype) void {
var line: [160]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// --- common-config register access (little-endian MMIO at `common_base`) ---------------
fn cfgRead(comptime T: type, comptime field: []const u8) T {
return mmio.read(T, common_base + @offsetOf(vp.CommonCfg, field));
}
fn cfgWrite(comptime T: type, comptime field: []const u8, value: T) void {
mmio.write(T, common_base + @offsetOf(vp.CommonCfg, field), value);
}
/// Write a 64-bit common-config register as two 32-bit halves (low then high) — the widest
/// access every virtio-pci host is required to accept for the queue-address registers.
fn cfgWrite64(comptime field: []const u8, value: u64) void {
const at = common_base + @offsetOf(vp.CommonCfg, field);
mmio.write(u32, at, @truncate(value));
mmio.write(u32, at + 4, @truncate(value >> 32));
}
fn orStatus(bit: u8) void {
cfgWrite(u8, "device_status", cfgRead(u8, "device_status") | bit);
}
// --- PCI config-space capability walk (config space is resource 0) ---------------------
/// Map the BAR numbered `bar` (0..5) and return its virtual base, correlating the BAR's
/// physical address (read from config space) with one of our device resources — because a
/// virtio capability names a BAR *number*, while `mmio_map` takes a *resource index* (and
/// resource 0 is config space, so BAR resources are re-numbered and gaps skipped).
fn mapBar(config: usize, descriptor: *const device.DeviceDescriptor, bar: u8) ?usize {
if (bar >= 6) return null;
if (bar_virtual[bar] != 0) return bar_virtual[bar];
const low = mmio.read(u32, config + 0x10 + @as(usize, bar) * 4);
if (low & 0x1 != 0) return null; // an I/O-space BAR — virtio structures are in memory BARs
var base: u64 = low & 0xFFFF_FFF0;
if ((low & 0x6) == 0x4) { // 64-bit memory BAR: the high half is the next dword
const high = mmio.read(u32, config + 0x10 + (@as(usize, bar) + 1) * 4);
base |= @as(u64, high) << 32;
}
for (descriptor.resources[0..@intCast(descriptor.resource_count)], 0..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory) and resource.start == base) {
const v = device.mmioMap(device_id, index) orelse return null;
bar_virtual[bar] = v;
return v;
}
}
log("virtio-gpu: BAR {d} (physical 0x{x}) is not a mapped resource\n", .{ bar, base });
return null;
}
/// Walk the PCI capability list from mapped config space, recording the common-config and
/// notify structures (the only two V3 needs). Returns false if either is missing.
fn walkCapabilities(config: usize, descriptor: *const device.DeviceDescriptor) bool {
if (mmio.read(u16, config + 0x06) & 0x10 == 0) { // Status bit 4: capabilities list present
log("virtio-gpu: device has no PCI capability list\n", .{});
return false;
}
var cap: u8 = @as(u8, @truncate(mmio.read(u8, config + 0x34))) & 0xFC;
var guard: u32 = 0;
while (cap != 0 and guard < 48) : (guard += 1) {
const at = config + cap;
const id = mmio.read(u8, at + 0);
const next = mmio.read(u8, at + 1) & 0xFC;
// Only map BARs for the structures V3 uses (common + notify). The other virtio
// capabilities (isr, device, and especially the cfg_pci back-door, which carries a
// placeholder bar=0/offset=0) reference BARs we never touch, so mapping them would
// just log spurious "not a mapped resource" noise.
if (id == vp.pci_cap_vendor) {
const cfg_type = mmio.read(u8, at + 3);
if (cfg_type == vp.cfg_common or cfg_type == vp.cfg_notify) {
const bar = mmio.read(u8, at + 4);
const offset = mmio.read(u32, at + 8);
if (mapBar(config, descriptor, bar)) |bar_base| {
if (cfg_type == vp.cfg_common) {
common_base = bar_base + offset;
} else {
notify_base = bar_base + offset;
notify_multiplier = mmio.read(u32, at + 16); // virtio_pci_notify_cap tail
}
}
}
}
cap = next;
}
if (common_base == 0 or notify_base == 0) {
log("virtio-gpu: missing common-config or notify capability\n", .{});
return false;
}
return true;
}
// --- the control virtqueue -------------------------------------------------------------
/// Publish the two-descriptor chain (request read by the device, response written by it),
/// notify the control queue, and wait for the device to return the buffer on the used ring.
fn submit(request_len: usize, response_len: usize) bool {
const desc: [*]vp.Desc = @ptrFromInt(ring.virtual + desc_offset);
desc[0] = .{
.addr = command.physical + request_offset,
.len = @intCast(request_len),
.flags = vp.desc_flag_next,
.next = 1,
};
desc[1] = .{
.addr = command.physical + response_offset,
.len = @intCast(response_len),
.flags = vp.desc_flag_write,
.next = 0,
};
const avail_ring: [*]u16 = @ptrFromInt(ring.virtual + avail_offset + 4);
avail_ring[avail_shadow % queue_size] = 0; // head of the chain is descriptor 0
mmio.wmb();
avail_shadow +%= 1;
mmio.write(u16, ring.virtual + avail_offset + 2, avail_shadow); // avail.idx
mmio.wmb();
mmio.write(u16, notify_addr, 0); // ring the control queue's doorbell
return waitUsed();
}
/// Spin, then sleep-poll, on the used-ring index until the device advances it. QEMU
/// processes the notify on its own thread, so the ack usually lands immediately; the sleep
/// fallback covers a device that defers it without burning the CPU.
fn waitUsed() bool {
var tries: u32 = 0;
while (tries < 2000) : (tries += 1) {
mmio.rmb();
const idx = mmio.read(u16, ring.virtual + used_offset + 2); // used.idx
if (idx != used_shadow) {
used_shadow = idx;
return true;
}
if (tries > 8) system.sleep(1);
}
return false;
}
/// The type field of the response the device wrote — `resp_ok_nodata` on success.
fn responseType() u32 {
const response: *vg.CtrlHdr = @ptrFromInt(command.virtual + response_offset);
return response.type;
}
/// Submit a command whose response is a bare header, returning its response type (0 if the
/// device never acked).
fn command_nodata(request_len: usize) u32 {
if (!submit(request_len, @sizeOf(vg.CtrlHdr))) return 0;
return responseType();
}
const ok_nodata: u32 = @intFromEnum(vg.CmdType.resp_ok_nodata);
fn requestAt(comptime T: type) *T {
return @ptrFromInt(command.virtual + request_offset);
}
/// A deterministic, recognisable pixel so a read-back is a real check, not a tautology.
fn testPixel(index: u32) u32 {
return 0xFF00_0000 | (index *% 0x9E37_79B1);
}
// --- bring-up --------------------------------------------------------------------------
fn initialise(endpoint: ipc.Handle) bool {
_ = endpoint;
if (!device.claim(device_id)) {
log("virtio-gpu: unable to claim device {d}\n", .{device_id});
return false;
}
var descriptors: [64]device.DeviceDescriptor = undefined;
const total = device.enumerate(&descriptors);
const descriptor = for (descriptors[0..@min(total, descriptors.len)]) |*d| {
if (d.id == device_id) break d;
} else {
log("virtio-gpu: device {d} not in the device tree\n", .{device_id});
return false;
};
// Config space is resource 0. Confirm it really is a virtio-gpu, then enable memory-space
// decode + bus mastering (the device DMAs the ring and backing out of RAM); pci-bus only
// preserves whatever the firmware left, and a secondary display is often left disabled.
const config = device.mmioMap(device_id, 0) orelse {
log("virtio-gpu: config-space map failed\n", .{});
return false;
};
const vendor = mmio.read(u16, config + 0x00);
const dev = mmio.read(u16, config + 0x02);
if (vendor != virtio_vendor or dev != virtio_gpu_device) {
log("virtio-gpu: not a virtio-gpu (vendor 0x{x} device 0x{x})\n", .{ vendor, dev });
return false;
}
mmio.write(u16, config + 0x04, mmio.read(u16, config + 0x04) | 0x06); // MEM + bus master
if (!walkCapabilities(config, descriptor)) return false;
// Reset, then the modern feature handshake: acknowledge, take driver ownership, require
// VERSION_1 and offer nothing else, and confirm the device accepts that.
cfgWrite(u8, "device_status", 0);
orStatus(vp.status_acknowledge);
orStatus(vp.status_driver);
// Low feature word (device-specific): note whether the device offers EDID (bit 1).
cfgWrite(u32, "device_feature_select", 0);
edid_available = cfgRead(u32, "device_feature") & vg.feature_edid != 0;
// High feature word: VERSION_1 (bit 32) is required for a modern device.
cfgWrite(u32, "device_feature_select", vp.feature_version_1_word);
if (cfgRead(u32, "device_feature") & vp.feature_version_1_bit == 0) {
log("virtio-gpu: device does not offer VERSION_1 (not a modern device)\n", .{});
return false;
}
// Accept exactly VERSION_1, plus EDID when the device offered it (never a feature it didn't).
cfgWrite(u32, "driver_feature_select", 0);
cfgWrite(u32, "driver_feature", if (edid_available) vg.feature_edid else 0);
cfgWrite(u32, "driver_feature_select", vp.feature_version_1_word);
cfgWrite(u32, "driver_feature", vp.feature_version_1_bit);
orStatus(vp.status_features_ok);
if (cfgRead(u8, "device_status") & vp.status_features_ok == 0) {
log("virtio-gpu: device rejected the negotiated features\n", .{});
return false;
}
// Stand up the control virtqueue (queue 0) in coherent DMA memory.
cfgWrite(u16, "queue_select", 0);
const device_qsize = cfgRead(u16, "queue_size");
if (device_qsize < queue_size) {
log("virtio-gpu: control queue too small ({d})\n", .{device_qsize});
return false;
}
ring = dma.alloc(4096, dma.coherent) orelse {
log("virtio-gpu: virtqueue allocation failed\n", .{});
return false;
};
command = dma.alloc(4096, dma.coherent) orelse {
log("virtio-gpu: command-buffer allocation failed\n", .{});
return false;
};
mmio.write(u16, ring.virtual + avail_offset, 1); // VIRTQ_AVAIL_F_NO_INTERRUPT: we poll
cfgWrite(u16, "queue_size", queue_size);
cfgWrite64("queue_desc", ring.physical + desc_offset);
cfgWrite64("queue_driver", ring.physical + avail_offset);
cfgWrite64("queue_device", ring.physical + used_offset);
cfgWrite(u16, "queue_msix_vector", 0xFFFF); // VIRTIO_MSI_NO_VECTOR
cfgWrite(u16, "queue_enable", 1);
cfgWrite(u16, "queue_select", 0);
notify_addr = notify_base + @as(usize, cfgRead(u16, "queue_notify_off")) * notify_multiplier;
orStatus(vp.status_driver_ok);
// Drive the GPU: create a 2D resource at the *max* mode, back it with a shared surface, and
// scan out the current-mode rectangle within it.
{
const request = requestAt(vg.ResourceCreate2d);
request.* = .{
.hdr = .{ .type = @intFromEnum(vg.CmdType.resource_create_2d) },
.resource_id = resource_id,
.format = vg.format_b8g8r8x8_unorm,
.width = max_width,
.height = max_height,
};
if (command_nodata(@sizeOf(vg.ResourceCreate2d)) != ok_nodata) {
log("virtio-gpu: resource_create_2d failed\n", .{});
return false;
}
}
// Back the resource with a shared (shm) surface, so the compositor and the device work
// the same physical pages. The device needs the guest-physical base for attach_backing.
surface = shm.create(scanout_bytes) orelse {
log("virtio-gpu: scanout surface allocation failed\n", .{});
return false;
};
const surface_physical = shm.physical(surface.handle) orelse {
log("virtio-gpu: could not resolve the scanout surface physical address\n", .{});
return false;
};
{
const request = requestAt(vg.ResourceAttachBacking);
request.* = .{
.hdr = .{ .type = @intFromEnum(vg.CmdType.resource_attach_backing) },
.resource_id = resource_id,
.nr_entries = 1,
};
const entry: *vg.MemEntry = @ptrFromInt(command.virtual + request_offset + @sizeOf(vg.ResourceAttachBacking));
entry.* = .{ .addr = surface_physical, .length = @intCast(scanout_bytes) };
if (command_nodata(@sizeOf(vg.ResourceAttachBacking) + @sizeOf(vg.MemEntry)) != ok_nodata) {
log("virtio-gpu: resource_attach_backing failed\n", .{});
return false;
}
}
if (!setScanoutRect()) {
log("virtio-gpu: set_scanout failed\n", .{});
return false;
}
log("virtio-gpu: scanout {d}x{d} online\n", .{ current_width, current_height });
// Hello the device manager so it counts us as up (and does not stop us at the hello
// deadline). A restarted instance re-hellos here and re-announces below — the compositor
// re-attaches to the fresh scanout (V6).
helloManager();
// Read the monitor's EDID (best-effort, when the device offers it) — the mode list a real
// driver derives from it; we log the preferred mode and keep our fixed offered list.
readEdid();
// Paint a known pattern, present it, and read it back — the V3 self-test that proves the
// whole path (virtqueue, resource, shared backing, transfer, flush) before a client attaches.
const pixels: [*]u32 = @ptrCast(@alignCast(surface.ptr));
const pixel_count: usize = @as(usize, max_width) * max_height;
for (0..pixel_count) |i| pixels[i] = testPixel(@intCast(i));
if (!presentFull()) {
log("virtio-gpu: initial present failed\n", .{});
return false;
}
// The scanout surface is CPU-visible RAM: read the pattern back to prove the mapping,
// which together with the flush ack above is the automated stand-in for "it's on screen".
mmio.rmb();
if (pixels[0] != testPixel(0) or pixels[pixel_count / 2] != testPixel(@intCast(pixel_count / 2))) {
log("virtio-gpu: pixel read-back mismatch\n", .{});
return false;
}
log("virtio-gpu: flush acked, pixel check ok\n", .{});
// Offer the shared surface to the compositor so it upgrades off the GOP floor (V4).
announce();
return true;
}
/// Point scanout 0 at the current-mode rectangle of the resource. Reused by initial bring-up
/// and by `set_mode`.
fn setScanoutRect() bool {
const request = requestAt(vg.SetScanout);
request.* = .{
.hdr = .{ .type = @intFromEnum(vg.CmdType.set_scanout) },
.rect = .{ .x = 0, .y = 0, .width = current_width, .height = current_height },
.scanout_id = 0,
.resource_id = resource_id,
};
return command_nodata(@sizeOf(vg.SetScanout)) == ok_nodata;
}
/// Read and log the monitor's preferred mode from its EDID (VIRTIO_GPU_F_EDID). Best-effort:
/// a device that doesn't offer EDID, or a missing/short block, is logged and ignored.
fn readEdid() void {
if (!edid_available) {
log("virtio-gpu: EDID not offered by device\n", .{});
return;
}
const request = requestAt(vg.GetEdid);
request.* = .{ .hdr = .{ .type = @intFromEnum(vg.CmdType.get_edid) }, .scanout = 0 };
if (!submit(@sizeOf(vg.GetEdid), @sizeOf(vg.RespEdid))) {
log("virtio-gpu: EDID request not acked\n", .{});
return;
}
const response: *vg.RespEdid = @ptrFromInt(command.virtual + response_offset);
if (response.hdr.type != @intFromEnum(vg.CmdType.resp_ok_edid) or response.size < 64) {
log("virtio-gpu: EDID unavailable\n", .{});
return;
}
// The first detailed timing descriptor (EDID base-block offset 54) is the preferred mode:
// active pixels are 12-bit, low byte + high nibble (bytes 2/4 horizontal, 5/7 vertical).
const e = &response.edid;
const h_active = @as(u32, e[56]) | (@as(u32, e[58] & 0xF0) << 4);
const v_active = @as(u32, e[59]) | (@as(u32, e[61] & 0xF0) << 4);
log("virtio-gpu: EDID preferred mode {d}x{d}\n", .{ h_active, v_active });
}
/// Present the whole surface: copy the guest backing into the host resource, then flush it to
/// the panel. Reused by the V3 self-test and by every compositor present over `.scanout`. V4
/// presents the full surface; the damage-rect fast path is a later refinement.
fn presentFull() bool {
mmio.wmb(); // the surface writes must be visible before the device transfers them
{
// Transfer the current-mode rectangle from the guest backing to the host resource. The
// device uses the resource's (max) width as the row stride, so the top-left rect at
// offset 0 is exactly the visible area — the compositor composes at that same stride.
const request = requestAt(vg.TransferToHost2d);
request.* = .{
.hdr = .{ .type = @intFromEnum(vg.CmdType.transfer_to_host_2d) },
.rect = .{ .x = 0, .y = 0, .width = current_width, .height = current_height },
.offset = 0,
.resource_id = resource_id,
};
if (command_nodata(@sizeOf(vg.TransferToHost2d)) != ok_nodata) return false;
}
{
// A fenced flush (vsync): the device signals the fence when the frame is actually on
// screen — which its used-ring ack, what our synchronous submit waits on, already gates.
const request = requestAt(vg.ResourceFlush);
request.* = .{
.hdr = .{ .type = @intFromEnum(vg.CmdType.resource_flush), .flags = vg.flag_fence, .fence_id = fence_next },
.rect = .{ .x = 0, .y = 0, .width = current_width, .height = current_height },
.resource_id = resource_id,
};
fence_next += 1;
if (command_nodata(@sizeOf(vg.ResourceFlush)) != ok_nodata) return false;
}
return true;
}
/// Hello the device manager (role: bus — we own a PCI function, though we report no children):
/// the handshake that marks us up so the manager doesn't stop us at the hello deadline, and
/// (as a supervised driver) restarts us if we die. Best-effort: without a manager we still run.
fn helloManager() void {
var tries: u32 = 0;
const manager = while (tries < 100) : (tries += 1) {
if (ipc.lookup(.device_manager)) |h| break h;
system.sleep(20);
} else {
log("virtio-gpu: no device manager to hello\n", .{});
return;
};
const hello = dm.Hello{ .role = @intFromEnum(dm.Role.bus), .device_id = device_id };
var reply: [dm.reply_size]u8 = undefined;
const n = ipc.call(manager, std.mem.asBytes(&hello), &reply) catch {
log("virtio-gpu: hello call failed\n", .{});
return;
};
if (n < dm.reply_size or std.mem.bytesToValue(dm.HelloReply, reply[0..dm.reply_size]).status != 0) {
log("virtio-gpu: hello refused\n", .{});
return;
}
log("virtio-gpu: hello acknowledged\n", .{});
}
/// Announce the scanout to the display service so it upgrades off the GOP framebuffer: hand it
/// the shared surface as a capability plus the geometry. Best-effort and non-fatal — without a
/// display service (the standalone virtio-gpu bring-up test) the driver is still a valid
/// scanout service; it just serves no one. The display replies immediately (it defers its
/// first present to a timer), so this returns before we start serving `.scanout` — no deadlock.
fn announce() void {
var tries: u32 = 0;
const display = while (tries < 50) : (tries += 1) {
if (ipc.lookup(.display)) |h| break h;
system.sleep(20);
} else {
log("virtio-gpu: no display service to announce to (scanout-only)\n", .{});
return;
};
var request = dp.Request{
.operation = @intFromEnum(dp.Operation.attach_scanout),
.x = max_width, // the shared surface's row stride in pixels (it is sized to the max mode)
.width = current_width,
.height = current_height,
.colour = display_format_bgrx,
};
var reply: [dp.reply_size]u8 = undefined;
_ = ipc.callCap(display, std.mem.asBytes(&request), &reply, surface.handle) catch {
log("virtio-gpu: announce to display failed\n", .{});
return;
};
log("virtio-gpu: announced scanout to display\n", .{});
}
/// A `sp.Reply{status}` written into `reply`.
fn scanoutStatus(reply: []u8, ok: bool) usize {
const response = sp.Reply{ .status = if (ok) 0 else -1 };
@memcpy(reply[0..sp.reply_size], std.mem.asBytes(&response));
return sp.reply_size;
}
/// The `.scanout` service: the compositor drives present / mode queries here. The pixels are
/// already in the shared surface, so a present is a transfer-to-host + fenced flush; a mode
/// change just re-points the scanout rectangle (the surface is sized to the largest mode).
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize {
_ = sender;
_ = capability;
if (message.len < sp.request_size) return 0;
const request = std.mem.bytesToValue(sp.Request, message[0..sp.request_size]);
switch (request.operation) {
@intFromEnum(sp.Operation.present) => return scanoutStatus(reply, presentFull()),
@intFromEnum(sp.Operation.get_modes) => {
var response = sp.ModesReply{ .status = 0, .count = offered_modes.len, .modes = undefined };
for (0..sp.max_modes) |i| {
response.modes[i] = if (i < offered_modes.len)
.{ .width = offered_modes[i].width, .height = offered_modes[i].height }
else
.{ .width = 0, .height = 0 };
}
@memcpy(reply[0..sp.modes_reply_size], std.mem.asBytes(&response));
return sp.modes_reply_size;
},
@intFromEnum(sp.Operation.set_mode) => {
const w = request.width;
const h = request.height;
if (w == 0 or h == 0 or w > max_width or h > max_height) return scanoutStatus(reply, false);
current_width = w;
current_height = h;
return scanoutStatus(reply, setScanoutRect());
},
else => return 0,
}
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse {
_ = system.write("virtio-gpu: missing device id (argv[1])\n");
return;
};
device_id = std.fmt.parseInt(u64, argument, 10) catch {
log("virtio-gpu: malformed device id '{s}'\n", .{argument});
return;
};
runtime.service.run(256, .{
.service = .scanout,
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+105
View File
@@ -0,0 +1,105 @@
//! virtio 1.0 PCI transport — the vendor capabilities in PCI config space that point at the
//! device's structures (common config, notify, ISR) in a BAR, the common-config register
//! block, and the split-virtqueue layout. `extern` structs matching the spec. Host-tested
//! for size. See docs/display-v2.md.
const std = @import("std");
/// PCI vendor-specific capability id (0x09) — virtio 1.0 structures are advertised as these.
pub const pci_cap_vendor: u8 = 0x09;
/// virtio_pci_cap `cfg_type`: which structure a vendor capability points at.
pub const cfg_common: u8 = 1;
pub const cfg_notify: u8 = 2;
pub const cfg_isr: u8 = 3;
pub const cfg_device: u8 = 4;
pub const cfg_pci: u8 = 5;
/// virtio_pci_cap — a vendor capability naming a structure at (bar, offset, length) within
/// a PCI BAR. Read straight out of config space.
pub const PciCap = extern struct {
cap_vndr: u8, // 0x09
cap_next: u8, // next capability's offset in config space (0 = end)
cap_len: u8,
cfg_type: u8, // cfg_common / cfg_notify / ...
bar: u8, // which BAR the structure lives in
padding: [3]u8,
offset: u32, // offset within the BAR
length: u32, // length of the structure
};
/// virtio_pci_notify_cap: a notify capability carries a multiplier after the base cap; the
/// per-queue notify address is `notify_base + queue_notify_off * notify_off_multiplier`.
pub const NotifyCap = extern struct {
cap: PciCap,
notify_off_multiplier: u32,
};
/// virtio_pci_common_cfg — the common configuration register block (little-endian MMIO).
pub const CommonCfg = extern struct {
device_feature_select: u32,
device_feature: u32,
driver_feature_select: u32,
driver_feature: u32,
msix_config: u16,
num_queues: u16,
device_status: u8,
config_generation: u8,
queue_select: u16,
queue_size: u16,
queue_msix_vector: u16,
queue_enable: u16,
queue_notify_off: u16,
queue_desc: u64,
queue_driver: u64,
queue_device: u64,
};
/// device_status bits (written to `CommonCfg.device_status` during bring-up).
pub const status_acknowledge: u8 = 1;
pub const status_driver: u8 = 2;
pub const status_driver_ok: u8 = 4;
pub const status_features_ok: u8 = 8;
/// VIRTIO_F_VERSION_1 — feature bit 32 (in the second 32-bit feature word). Required for a
/// modern device; we negotiate exactly this bit and nothing else.
pub const feature_version_1_word: u32 = 1; // device_feature_select value for bits 32..63
pub const feature_version_1_bit: u32 = 1 << 0; // bit 32 within that word
// --- split virtqueue -------------------------------------------------------
pub const Desc = extern struct {
addr: u64, // guest-physical
len: u32,
flags: u16,
next: u16,
};
pub const desc_flag_next: u16 = 1; // buffer continues in `next`
pub const desc_flag_write: u16 = 2; // device-writable (else driver-writable/device-readable)
/// The available ring's fixed header; a `[queue_size]u16` ring and a trailing `used_event`
/// u16 follow it in memory (laid out by the driver).
pub const AvailHdr = extern struct {
flags: u16,
idx: u16,
};
/// One entry of the used ring.
pub const UsedElem = extern struct {
id: u32,
len: u32,
};
/// The used ring's fixed header; a `[queue_size]UsedElem` ring and a trailing `avail_event`
/// u16 follow it.
pub const UsedHdr = extern struct {
flags: u16,
idx: u16,
};
test "virtio-pci struct sizes match the spec" {
try std.testing.expectEqual(@as(usize, 16), @sizeOf(PciCap));
try std.testing.expectEqual(@as(usize, 56), @sizeOf(CommonCfg));
try std.testing.expectEqual(@as(usize, 16), @sizeOf(Desc));
try std.testing.expectEqual(@as(usize, 8), @sizeOf(UsedElem));
}
@@ -188,6 +188,13 @@ pub fn mapUserDmaInto(root: u64, virtual: u64, physical: u64, len: u64) void {
paging.mapUserDmaInto(root, virtual, physical, len);
}
/// Map shared cacheable RAM into address space `root`: write-back cacheable, RW+NX, and
/// marked so teardown won't free the frames (they're owned by a refcounted shm object,
/// freed when its last capability drops). For shm_create/shm_map.
pub fn mapUserSharedInto(root: u64, virtual: u64, physical: u64, len: u64) void {
paging.mapUserSharedInto(root, virtual, physical, len);
}
/// Map a page into the kernel address space (non-executable). For the heap, etc.
pub fn mapPage(virtual: u64, physical: u64, writable: bool) void {
paging.map(virtual, physical, writable);
@@ -393,6 +393,30 @@ pub fn leafIsWriteCombining(pml4: u64, virtual: u64) ?bool {
return (e & pte_pat != 0) and (e & pcd == 0) and (e & pwt == 0);
}
/// Map `[physical, physical+len)` into the user half rooted at `pml4` as **shared cacheable
/// RAM**: write-back cacheable (RW + NX) for CPU compositing, and carrying `device_grant`
/// so teardown (`freeSubtree`) does **not** return the frames to the allocator. The frames
/// are owned by a refcounted shared-memory object (system/kernel/ipc-synchronous.zig) and
/// freed only when its last capability drops — not when one sharer's address space dies, or
/// the other sharers would be left mapping freed RAM. The caller aligns `virtual`/`physical`.
pub fn mapUserSharedInto(pml4: u64, virtual: u64, physical: u64, len: u64) void {
const flags: u64 = present | user | writable | no_execute | device_grant; // WB cacheable
const first = physical & ~@as(u64, page_size - 1);
const last = (physical + (if (len == 0) 1 else len) - 1) & ~@as(u64, page_size - 1);
var off: u64 = 0;
while (first + off <= last) : (off += page_size) {
const v = virtual + off;
const pml4e = &tableAt(pml4)[(v >> 39) & 0x1FF];
const pdpt = descendUser(pml4e);
const pdpte = &tableAt(pdpt)[(v >> 30) & 0x1FF];
const pd = descendUser(pdpte);
const pde = &tableAt(pd)[(v >> 21) & 0x1FF];
const pt = descendUser(pde);
tableAt(pt)[(v >> 12) & 0x1FF] = ((first + off) & address_mask) | flags;
invalidate(v);
}
}
/// Create a new address space: a fresh PML4 with an empty user half and the
/// kernel's higher half shared in (copying PML4[256..512), whose entries point
/// at the kernel's PDPTs — pre-created at init and never restaled, so growth in
+127 -18
View File
@@ -28,6 +28,7 @@ const architecture = @import("architecture");
const scheduler = @import("scheduler.zig");
const sync = @import("sync.zig");
const heap = @import("heap.zig");
const pmm = @import("pmm.zig");
const page_size = abi.page_size;
const Task = scheduler.Task;
@@ -89,6 +90,10 @@ const user_half_end: u64 = 0x0000_8000_0000_0000;
/// (per process) and/or by a registry slot, counted by `refcount`.
pub const Endpoint = struct {
refcount: u32 = 1,
// The task that created it. When that task dies, the endpoint is marked `dead` so a caller
// gets -EPEER instead of blocking forever on a service that will never reply again (V6).
owner: u32 = 0,
dead: bool = false,
// Callers blocked in `call`, awaiting receive, in FIFO order (threaded via
// Task.next; each such task is .blocked and in no scheduler queue).
sender_head: ?*Task = null,
@@ -109,10 +114,30 @@ pub const Endpoint = struct {
pub fn createIpcEndpoint() ?*Endpoint {
const endpoint = heap.allocator().create(Endpoint) catch return null;
endpoint.* = .{};
endpoint.* = .{ .owner = scheduler.currentId() };
return endpoint;
}
/// A task is dying: kill the endpoints it registered as services. Mark each `dead` (so a later
/// `call` returns -EPEER rather than blocking on a reply that will never come), wake anyone
/// already parked sending to it with that error, and vacate its registry slot. Only *registered*
/// endpoints are reachable from here; unregistered ones drop with the task's handle table. The
/// caller holds the big kernel lock (this runs on the death path). See docs/display-v2.md (V6).
pub fn killOwnedEndpointsLocked(task_id: u32) void {
for (&registry) |*slot| {
const endpoint = slot.* orelse continue;
if (endpoint.owner != task_id) continue;
endpoint.dead = true;
while (dequeueSender(endpoint)) |sender| {
sender.ipc_status = -EPEER;
sender.ipc_received_cap = abi.no_cap;
scheduler.readyLocked(sender);
}
slot.* = null;
dropRef(endpoint);
}
}
/// Drop a reference; free the endpoint when the last one goes. (Frames are leaked
/// today like other kernel objects — but the refcount bookkeeping lands now.)
pub fn dropRef(endpoint: *Endpoint) void {
@@ -123,6 +148,43 @@ pub fn dropRef(endpoint: *Endpoint) void {
}
}
// --- capability objects: what a handle-table entry can name ------------------
/// The `kind` tag on a `scheduler.HandleObject` — which capability object a handle names.
/// Defined here (not in scheduler) because the meaning is the IPC/capability layer's.
pub const handle_kind_endpoint: u8 = 0;
pub const handle_kind_shm: u8 = 1;
/// A page-aligned block of **shared cacheable RAM** (docs/display-v2.md), referenced by
/// capability handles across processes and freed when the last one drops. `phys` is its
/// contiguous physical base, `pages` its length. A sharer's address-space teardown never
/// reclaims these frames (the mapping carries `device_grant`); this object owns them.
pub const ShmObject = struct {
refcount: u32 = 1,
phys: u64,
pages: usize,
};
/// Wrap `pages` contiguous frames at `phys` (already allocated + zeroed by the caller) in a
/// refcounted shm object, or null if the heap is out of room.
pub fn createShm(phys: u64, pages: usize) ?*ShmObject {
const shm = heap.allocator().create(ShmObject) catch return null;
shm.* = .{ .phys = phys, .pages = pages };
return shm;
}
/// Drop a shared-memory reference; when the last one goes, return its frames to the
/// allocator and free the object. (The mappings themselves are torn down with each
/// sharer's address space; `device_grant` keeps that from freeing the frames early.)
pub fn dropShmRef(shm: *ShmObject) void {
if (shm.refcount > 1) {
shm.refcount -= 1;
} else {
for (0..shm.pages) |i| pmm.free(shm.phys + i * page_size);
heap.allocator().destroy(shm);
}
}
// --- sender FIFO (endpoint-local, via Task.next) ----------------------------
fn enqueueSender(endpoint: *Endpoint, t: *Task) void {
@@ -222,11 +284,25 @@ pub fn copyFromUser(user_as: u64, user_va: u64, destination: []u8) bool {
/// no live handle, or `-ENOSPC` if `to`'s table is full. Callers only invoke this when
/// `cap != no_cap`. Used by both IPC directions to carry an endpoint with a message.
fn shareCapability(from: *Task, to: *Task, cap: u64) i64 {
const endpoint = resolveHandle(from, cap) orelse return -EBADF;
endpoint.refcount += 1;
const handle = installHandle(to, endpoint);
if (cap >= from.handles.len) return -EBADF;
const entry = from.handles[@intCast(cap)] orelse return -EBADF;
// Bump the named object's refcount (a copy, not a move — the sender keeps its handle),
// dispatching by kind so both endpoints and shared-memory regions can travel with a
// message.
switch (entry.kind) {
handle_kind_endpoint => {
const e: *Endpoint = @ptrCast(@alignCast(entry.ptr));
e.refcount += 1;
},
handle_kind_shm => {
const s: *ShmObject = @ptrCast(@alignCast(entry.ptr));
s.refcount += 1;
},
else => return -EBADF,
}
const handle = installEntry(to, entry);
if (handle < 0) {
dropRef(endpoint); // undo the bump; the receiver had no room
dropEntry(entry); // undo the bump; the receiver had no room
return -ENOSPC;
}
return handle;
@@ -241,6 +317,7 @@ pub fn call(endpoint: *Endpoint, message_ptr: u64, message_len: u64, reply_ptr:
if (message_len > MESSAGE_MAXIMUM or reply_cap > MESSAGE_MAXIMUM) return -E2BIG;
const flags = sync.enter();
defer sync.leave(flags);
if (endpoint.dead) return -EPEER; // the service that owned this endpoint is gone — don't block
const me = scheduler.current();
me.ipc_send_ptr = message_ptr;
@@ -416,36 +493,68 @@ pub fn notifyFromIsr(endpoint: *Endpoint, badge: u64) void {
// --- per-process handle table + name registry -------------------------------
/// Install `endpoint` in task `t`'s handle table; returns the small-int handle or
/// -ENOSPC. The caller has already taken/holds the reference the slot represents.
pub fn installHandle(t: *Task, endpoint: *Endpoint) i64 {
/// Install a capability object (kind + pointer) in task `t`'s handle table; returns the
/// small-int handle or -ENOSPC. The caller has already taken/holds the reference the slot
/// represents.
fn installEntry(t: *Task, entry: scheduler.HandleObject) i64 {
for (&t.handles, 0..) |*slot, i| {
if (slot.* == null) {
slot.* = @ptrCast(endpoint);
slot.* = entry;
return @intCast(i);
}
}
return -ENOSPC;
}
/// Resolve a handle to its endpoint, or null if out of range / unused.
pub fn resolveHandle(t: *Task, h: u64) ?*Endpoint {
if (h >= t.handles.len) return null;
const slot = t.handles[@intCast(h)] orelse return null;
return @ptrCast(@alignCast(slot));
/// Install an endpoint handle. The common case; keeps the endpoint callers' signature.
pub fn installHandle(t: *Task, endpoint: *Endpoint) i64 {
return installEntry(t, .{ .kind = handle_kind_endpoint, .ptr = @ptrCast(endpoint) });
}
/// Drop every endpoint reference an exiting task holds. Called from the scheduler
/// exit path so a dead server's endpoints don't linger referenced.
/// Install a shared-memory handle.
pub fn installShmHandle(t: *Task, shm: *ShmObject) i64 {
return installEntry(t, .{ .kind = handle_kind_shm, .ptr = @ptrCast(shm) });
}
/// Resolve a handle to its endpoint, or null if out of range, unused, or a different kind
/// (e.g. an shm handle used where an endpoint is expected).
pub fn resolveHandle(t: *Task, h: u64) ?*Endpoint {
if (h >= t.handles.len) return null;
const entry = t.handles[@intCast(h)] orelse return null;
if (entry.kind != handle_kind_endpoint) return null;
return @ptrCast(@alignCast(entry.ptr));
}
/// Resolve a handle to its shared-memory object, or null if out of range, unused, or not
/// an shm handle.
pub fn resolveShm(t: *Task, h: u64) ?*ShmObject {
if (h >= t.handles.len) return null;
const entry = t.handles[@intCast(h)] orelse return null;
if (entry.kind != handle_kind_shm) return null;
return @ptrCast(@alignCast(entry.ptr));
}
/// Drop every capability reference an exiting task holds, dispatching by kind so a dead
/// task's endpoints *and* shared-memory regions are released correctly. Called from the
/// scheduler exit path.
pub fn closeHandles(t: *Task) void {
for (&t.handles) |*slot| {
if (slot.*) |p| {
dropRef(@ptrCast(@alignCast(p)));
if (slot.*) |entry| {
dropEntry(entry);
slot.* = null;
}
}
}
/// Drop the reference a handle-table entry represents, by kind.
fn dropEntry(entry: scheduler.HandleObject) void {
switch (entry.kind) {
handle_kind_endpoint => dropRef(@ptrCast(@alignCast(entry.ptr))),
handle_kind_shm => dropShmRef(@ptrCast(@alignCast(entry.ptr))),
else => {},
}
}
var registry: [maximum_services]?*Endpoint = .{null} ** maximum_services;
/// Publish `endpoint` under well-known `id` (takes a reference). Returns 0 or -errno.
+91
View File
@@ -81,6 +81,18 @@ pub const device_arena_end: u64 = device_arena_base + (4 << 30);
pub const dma_arena_base: u64 = 0x0000_7200_0000_0000;
pub const dma_arena_end: u64 = dma_arena_base + (256 << 20); // 256 MiB per process
/// The shared-memory arena: where `shm_create`/`shm_map` place shared cacheable regions, in
/// PML4[230] — a user-exclusive region distinct from the DMA arena. The frames are owned by
/// a refcounted shm object and freed when its last capability drops, not on teardown, so the
/// mapping carries `device_grant`. Per-process cursor in `Task.shm_map_next` (docs/display-v2.md).
pub const shm_arena_base: u64 = 0x0000_7300_0000_0000;
pub const shm_arena_end: u64 = shm_arena_base + (256 << 20); // 256 MiB per process
/// Largest single `shm_create`, in pages (32 MiB) — enough for a 4K framebuffer surface;
/// also an overflow guard on the page count. shm frames are contiguous (like DMA), so this
/// bounds the contiguous allocation asked of the frame allocator.
const maximum_shm_pages = 8192;
/// Largest single `mmap` grant, in pages (32 MiB). Big enough for a display service's
/// back buffer at up to 4K (3840x2160x4 ≈ 8100 pages); the user heap otherwise grows in
/// small chunks. `systemMmap` maps page by page with rollback, so this is only a sanity
@@ -212,6 +224,9 @@ fn system_call(state: *architecture.CpuState) void {
.timer_bind => systemTimerBind(state),
.klog_read => systemKlogRead(state),
.wall_clock => systemWallClock(state),
.shm_create => systemShmCreate(state),
.shm_map => systemShmMap(state),
.shm_physical => systemShmPhysical(state),
_ => fail(state),
}
}
@@ -460,6 +475,81 @@ fn systemDmaFree(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, 0);
}
/// shm_create(len) -> vaddr (rax), handle (rdx): grant `len` bytes (rounded up to whole
/// pages) of **shareable, zeroed, cacheable** RAM — contiguous frames mapped into the
/// caller's shm arena — and hand back the virtual address plus a capability handle. Unlike
/// `dma_alloc` the memory is write-back cacheable (for CPU compositing, not device DMA) and
/// its frames are owned by a refcounted object: the handle is passed to another process as
/// an `ipc_call` send_cap, that process `shm_map`s it, and the frames free only when the
/// last capability drops (docs/display-v2.md — the compositor↔native-driver and
/// app↔compositor surface path).
fn systemShmCreate(state: *architecture.CpuState) void {
const len = architecture.systemCallArg(state, 0);
const t = scheduler.current();
if (t.aspace == 0 or len == 0) return fail(state);
const pages: usize = @intCast((len + page_size - 1) / page_size);
if (pages == 0 or pages > maximum_shm_pages) return fail(state);
// Reserve arena virtual space up front, so a mapping failure needs no rollback.
if (t.shm_map_next == 0) t.shm_map_next = shm_arena_base;
const base_v = t.shm_map_next;
if (base_v + pages * page_size > shm_arena_end) return fail(state); // arena exhausted
const phys = pmm.allocContiguous(pages, ~@as(u64, 0)) orelse return fail(state);
// Zero through the physmap (the frames aren't mapped in the caller yet).
const kernel_view: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(phys));
@memset(kernel_view[0 .. pages * page_size], 0);
const shm = ipc.createShm(phys, pages) orelse {
for (0..pages) |i| pmm.free(phys + i * page_size);
return fail(state);
};
const handle = ipc.installShmHandle(t, shm);
if (handle < 0) {
ipc.dropShmRef(shm); // last ref: frees the object and its frames
return fail(state);
}
architecture.mapUserSharedInto(t.aspace, base_v, phys, pages * page_size);
t.shm_map_next = base_v + pages * page_size;
architecture.setSystemCallResult(state, base_v); // vaddr for the CPU
architecture.setSystemCallResult2(state, @intCast(handle)); // capability handle to pass on
}
/// shm_map(cap) -> vaddr: map the shared region named by a capability handle the caller
/// received (via an `ipc_call` send_cap) into its shm arena — the same physical frames the
/// creator sees — returning the virtual address. The handle already holds a reference (taken
/// when the capability was shared), so this only adds a mapping; it never bumps the refcount.
fn systemShmMap(state: *architecture.CpuState) void {
const cap = architecture.systemCallArg(state, 0);
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const shm = ipc.resolveShm(t, cap) orelse return fail(state); // not an shm handle we hold
if (t.shm_map_next == 0) t.shm_map_next = shm_arena_base;
const base_v = t.shm_map_next;
const size = shm.pages * page_size;
if (base_v + size > shm_arena_end) return fail(state);
architecture.mapUserSharedInto(t.aspace, base_v, shm.phys, size);
t.shm_map_next = base_v + size;
architecture.setSystemCallResult(state, base_v);
}
/// shm_physical(cap) -> paddr: the guest-physical base of a shared region the caller holds a
/// capability for. The frames are contiguous (allocated by `allocContiguous`), so a single
/// physical base + length describes the whole region — which is exactly what a driver needs
/// to hand a shm surface to a device (virtio-gpu `attach_backing`). Only a holder of the
/// capability can ask; there is no ambient way to turn a virtual address into a physical one.
fn systemShmPhysical(state: *architecture.CpuState) void {
const cap = architecture.systemCallArg(state, 0);
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const shm = ipc.resolveShm(t, cap) orelse return fail(state); // not an shm handle we hold
architecture.setSystemCallResult(state, shm.phys);
}
/// device_register(parent_id, descriptor_ptr) -> id: publish a child device below a device
/// this process has claimed. The bus-driver primitive: a process that owns a bus
/// enumerates it and hands each device it finds to the table, where a class driver
@@ -647,6 +737,7 @@ fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
scheduler.readyLocked(client); // its blocked `call` now returns the error
}
ipc.abandonSenderLocked(t);
ipc.killOwnedEndpointsLocked(t.id); // its registered services are gone: callers get -EPEER, not a hang
scheduler.removeFromWaitQueueLocked(t);
scheduler.forgetIpcClientLocked(t);
ipc.closeHandles(t);
+13 -3
View File
@@ -86,9 +86,11 @@ pub const Task = struct {
// uninitialised, process.zig seeds it on the first mmio_map). User task only.
device_map_next: u64 = 0,
// --- synchronous IPC (ipc_sync.zig) ---
// Per-process handle table: small-int handle -> *ipc_sync.Endpoint, kept
// opaque here so the scheduler and IPC modules don't import each other.
handles: [ipc_maximum_handles]?*anyopaque = .{null} ** ipc_maximum_handles,
// Per-process handle table: a small-int handle names a kernel capability object.
// Each entry tags its `kind` (an IPC endpoint or a shared-memory object) so the
// close/exit and cap-passing paths reclaim the right type. Kept opaque here so the
// scheduler and IPC modules don't import each other (ipc_sync.zig owns the kinds).
handles: [ipc_maximum_handles]?HandleObject = .{null} ** ipc_maximum_handles,
// A server holds the caller it currently owes a reply to (set by ReplyWait's
// receive, cleared when it replies). A client, while blocked in Call, records
// its message + reply buffers here and its result lands in `ipc_status`.
@@ -99,6 +101,7 @@ pub const Task = struct {
ipc_reply_cap: u64 = 0,
ipc_status: i64 = 0, // client: reply length / -errno, written by the replier
dma_map_next: u64 = 0, // bump pointer into this task's DMA arena (0 = unseeded)
shm_map_next: u64 = 0, // bump pointer into this task's shared-memory arena (0 = unseeded)
ipc_send_cap: u64 = ~@as(u64, 0), // handle to transfer with this message (abi.no_cap = none)
ipc_received_cap: u64 = ~@as(u64, 0), // client: handle the reply's transferred cap landed at (abi.no_cap = none)
next: ?*Task = null, // ready-queue link (also the endpoint sender-FIFO link)
@@ -125,6 +128,13 @@ pub const maximum_task_name = abi.maximum_process_name;
/// it dimensions a field of `Task`; ipc_sync.zig re-exports it.
pub const ipc_maximum_handles = 16;
/// One handle-table entry: a capability object plus a `kind` tag saying what `ptr` points
/// at (an ipc endpoint or a shared-memory object), so a task's exit path and the
/// capability-passing path reclaim/share the right type. The `kind` values are defined by
/// ipc_sync.zig (`handle_kind_*`); kept an opaque `u8` here so the scheduler doesn't import
/// the IPC module.
pub const HandleObject = struct { kind: u8, ptr: *anyopaque };
var tasks = [_]Task{.{}} ** maximum_tasks;
var next_id: u32 = 1;
+230 -67
View File
@@ -101,6 +101,14 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
displayServiceTest(boot_information);
} else if (eql(case, "display-demo")) {
displayDemoTest(boot_information);
} else if (eql(case, "shm")) {
shmTest(boot_information);
} else if (eql(case, "virtio-gpu")) {
virtioGpuTest(boot_information);
} else if (eql(case, "display-native")) {
displayNativeTest(boot_information);
} else if (eql(case, "display-reattach")) {
displayReattachTest(boot_information);
} else if (eql(case, "clock")) {
clockTest();
} else if (eql(case, "smp")) {
@@ -230,6 +238,19 @@ fn bufferHas(needle: []const u8) bool {
return std.mem.indexOf(u8, process.write_buffer[0..process.write_len], needle) != null;
}
/// How many `pci_device` functions the devices broker currently holds. A durable
/// snapshot, unlike a `bufferHas` poll of the single-latest write_buffer line, so a
/// test can wait on it without racing transient log output. `scratch` is
/// caller-owned to keep the (large) descriptor array off this helper's own frame.
fn brokerPciCount(scratch: []device_abi.DeviceDescriptor) u32 {
const k = @min(devices_broker.enumerate(scratch), scratch.len);
var count: u32 = 0;
for (scratch[0..k]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) count += 1;
}
return count;
}
/// Non-destructive checks of the memory map and frame allocator.
fn smoke(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: smoke\n", .{});
@@ -1895,18 +1916,14 @@ fn pciScanTest(boot_information: *const BootInformation) void {
};
// Post-flip (M19.3) ground truth: the kernel no longer enumerates PCI
// functions, so equivalence inverts — the broker's function count after
// the scan must equal what the driver itself reported finding.
// functions, so equivalence inverts — every PCI function the broker holds was
// put there by the ring-3 driver, so before the driver runs the broker holds
// none. One reusable descriptor buffer (each snapshot is ~20 KiB; three live
// at once would overflow the 64 KiB bootstrap stack this test runs on).
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var boot_pci: u32 = 0;
for (buffer[0..n]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) boot_pci += 1;
}
check("the kernel seeded no PCI functions (the walk retired)", boot_pci == 0);
check("the kernel seeded no PCI functions (the walk retired)", brokerPciCount(&buffer) == 0);
process.setInitialRamdisk(image);
process.write_count = 0;
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
@@ -1917,68 +1934,49 @@ fn pciScanTest(boot_information: *const BootInformation) void {
}
check("device-manager spawned (test-pci-restart mode)", manager != 0);
// First scan: wait for the driver's count line and parse the number.
const count_prefix = "pci-bus: ";
const count_suffix = " functions found";
var reported: u32 = 0;
scheduler.setPriority(1);
// The manager spawns pci-bus, which scans the ECAM window and registers every
// function it finds, so the broker's PCI count climbs from zero and plateaus.
// Wait for it to *settle*: latch N only once the count has held steady for a
// stretch, so a mid-scan sample can't latch a low N that the rest of the same
// scan then appears to exceed. The count is monotonic (registrations only add;
// the table has no unregister) so the plateau is permanent — stability is
// reached the moment the scan finishes and holds indefinitely. Sleep between
// samples rather than busy-yield: the boot context outranks the drivers, and a
// busy spin would starve the very processes it waits on; a sleeping task is
// woken by the timer, so the drivers run in between.
const poll_ms = 5;
var registered: u32 = 0;
var steady: u32 = 0;
var deadline = architecture.millis() + 15000;
while (architecture.millis() < deadline and reported == 0) {
const line = process.write_buffer[0..process.write_len];
if (std.mem.indexOf(u8, line, count_prefix)) |start| {
if (std.mem.indexOf(u8, line, count_suffix)) |digits_end| {
reported = std.fmt.parseInt(u32, line[start + count_prefix.len .. digits_end], 10) catch 0;
}
}
scheduler.yield();
while (architecture.millis() < deadline and steady < 60) { // 60 * 5ms = 300ms steady
const now = brokerPciCount(&buffer);
if (now != 0 and now == registered) steady += 1 else steady = 0;
registered = now;
scheduler.sleep(poll_ms);
}
scheduler.setPriority(4);
check("the ring-3 scan reported a function count", reported >= 1);
check("the ring-3 scan registered its PCI functions in the broker", registered >= 1);
// Every reported function was registered: the broker holds exactly them.
var registered: [64]device_abi.DeviceDescriptor = undefined;
const r = @min(devices_broker.enumerate(&registered), registered.len);
var registered_pci: u32 = 0;
for (registered[0..r]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) registered_pci += 1;
// The restart drill — the manager kills pci-bus ~1 s after its scan, prunes its
// own child tree, and respawns it to re-claim, re-scan, and re-register — is
// asserted by the harness's ordered regex over the whole serial log, the way
// every restart drill is (see usbReportTest / driverRestartTest): the manager's
// kill/prune/respawn lines are transient and would race a write_buffer poll, and
// the *broker* count can't witness the restart at all — the table has no
// unregister and register is idempotent (devices-broker.zig), so the kill leaves
// the nodes in place and the respawn's re-registration dedupes against them.
//
// That idempotence is exactly this test's kernel-side claim: watch, across the
// whole drill, that the count never grows past N. A broken dedup would append
// the re-scanned functions as duplicates (N -> 2N), and with no unregister that
// overshoot would persist — so a single late sample would catch it; the loop is
// belt-and-braces over the ~2 s the kill + backoff + respawn takes.
var duplicated = false;
deadline = architecture.millis() + 5000;
while (architecture.millis() < deadline and !duplicated) {
if (brokerPciCount(&buffer) > registered) duplicated = true;
scheduler.sleep(20);
}
check("the broker holds exactly the reported functions", registered_pci == reported);
const kernel_count = reported; // the no-duplicate check below reuses it
// The restart drill: the manager kills pci-bus after its reports; the
// respawn re-claims, re-scans, and re-registers.
const restart_marker = "device-manager: restarting pci-bus";
scheduler.setPriority(1);
deadline = architecture.millis() + 15000;
var restarted = false;
while (architecture.millis() < deadline and !restarted) {
if (bufferHas(restart_marker)) restarted = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the manager restarted pci-bus", restarted);
var marker_buffer: [48]u8 = undefined;
const marker = std.fmt.bufPrint(&marker_buffer, "pci-bus: {d} functions found", .{reported}) catch "";
scheduler.setPriority(1);
deadline = architecture.millis() + 15000;
var seen = false;
while (architecture.millis() < deadline and !seen) {
if (bufferHas(marker)) seen = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the respawned scan reported the same count", seen);
// No duplicates: the registrations deduped against the kernel's own nodes
// on the first pass, and against themselves on the second.
var after: [64]device_abi.DeviceDescriptor = undefined;
const m = @min(devices_broker.enumerate(&after), after.len);
var after_count: u32 = 0;
for (after[0..m]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) after_count += 1;
}
check("no duplicate PCI nodes after register + restart + re-register", after_count == kernel_count);
check("no duplicate PCI nodes after the restart drill", !duplicated);
result();
}
@@ -2377,6 +2375,171 @@ fn displayDemoTest(boot_information: *const BootInformation) void {
while (true) scheduler.yield();
}
/// V2 — cross-process shared memory (docs/display-v2.md). Spawn shm-server and shm-client:
/// the client shm_creates a region, writes a pattern, and passes the region's capability to
/// the server as an ipc_call send_cap; the server shm_maps it and confirms the pattern is
/// visible — proving the two processes share the same physical pages, and that the extended
/// capability-passing (endpoints → memory objects) works. Its `shm: shared 4096 bytes ok`
/// heartbeat is the marker.
fn shmTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: shm\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
if (!spawnNamed(rd, "shm-server")) {
log("shm: could not spawn shm-server\n", .{});
result();
return;
}
_ = spawnNamed(rd, "shm-client");
scheduler.setPriority(1); // below the two, so they run
while (true) scheduler.yield();
}
/// V3 — the virtio-gpu driver, end to end (docs/display-v2.md). Boot the device-manager
/// stack (in its normal mode) so it discovers the virtio-gpu PCI function — present because
/// the harness boots this case with QEMU's `-device virtio-gpu-pci` — and spawns the driver.
/// The driver claims the device, brings up the control virtqueue, creates a 2D scanout,
/// paints a known pattern, flushes it, and waits for the device's used-ring ack. Its serial
/// heartbeats — `virtio-gpu: scanout WxH online` and `virtio-gpu: flush acked, pixel check
/// ok` — are the harness's markers (it reads serial directly, like the display cases). The
/// used-ring ack is the device confirming it consumed the frame; the pixel read-back proves
/// the backing is CPU-visible RAM — together the automated stand-in for "it's on screen".
fn virtioGpuTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: virtio-gpu\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
// Spawn device-manager in its normal mode: its initialise discovers the PCI host bridge
// from the kernel device tree, spawns pci-bus, and matches the virtio-gpu class triple to
// spawn our driver with the function's device id as argv[1].
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0;
break;
}
if (manager == 0) {
log("virtio-gpu: could not spawn device-manager\n", .{});
result();
return;
}
scheduler.setPriority(1); // below the manager and the driver it spawns, so they run
while (true) scheduler.yield();
}
/// V4 — the native backend + hot-attach (docs/display-v2.md). Boot the compositor and the
/// hardware-free `display-demo` client (as displayDemoTest does), then the device-manager
/// stack so it discovers the virtio-gpu function — present via QEMU's `-device
/// virtio-gpu-pci` — and spawns the driver. The driver brings up its scanout, then announces
/// the shared surface to the already-running compositor, which maps it, upgrades off the GOP
/// floor, and presents the composited frame through the native backend. Its serial heartbeats
/// — `display: scanout upgraded to virtio-gpu` and `display: native present verified` — plus
/// the demo's own `display-demo: ok` are the harness's markers. Display is spawned first so
/// it is registered on `.display` before the driver announces.
fn displayNativeTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: display-native\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0;
break;
}
if (manager == 0) {
log("display-native: could not spawn device-manager\n", .{});
result();
return;
}
if (!spawnNamed(rd, "display")) {
log("display-native: could not spawn the display service\n", .{});
result();
return;
}
_ = spawnNamed(rd, "display-demo");
scheduler.setPriority(1); // below the compositor, the demo, and the driver, so they run
while (true) scheduler.yield();
}
/// V6 — resilience: the compositor survives the virtio-gpu driver dying and re-attaches when
/// device-manager restarts it (docs/display-v2.md). Same boot as display-native, but the
/// manager runs in "test-scanout-restart" mode: a moment after the driver hellos, it kills it
/// once; the normal restart policy respawns it, the restarted driver re-announces, and the
/// compositor re-attaches to the fresh scanout — logging `display: scanout re-attached` after
/// the initial `display: scanout upgraded to virtio-gpu`. The compositor must not crash.
fn displayReattachTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: display-reattach\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-scanout-restart" }, scheduler.currentId(), null) catch 0;
break;
}
if (manager == 0) {
log("display-reattach: could not spawn device-manager\n", .{});
result();
return;
}
if (!spawnNamed(rd, "display")) {
log("display-reattach: could not spawn the display service\n", .{});
result();
return;
}
_ = spawnNamed(rd, "display-demo");
scheduler.setPriority(1); // below the compositor, the demo, and the driver, so they run
while (true) scheduler.yield();
}
/// Process arguments, end to end: spawn args-echo bare (its argv[0] is the
/// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through
/// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument
@@ -41,6 +41,15 @@ const xhci_pci_class: u64 = pci_class.ClassCode.pack(.{
.prog_if = @intFromEnum(pci_class.serial_bus.usb.ProgIf.xhci),
});
/// The PCI class triple of a virtio-gpu — Display Controller / Other (0x80) / 0. The class
/// alone cannot tell it from any other display/other function, so the driver re-confirms
/// vendor 0x1AF4 / device 0x1050 from config space once spawned; this only gets it spawned.
const virtio_gpu_pci_class: u64 = pci_class.ClassCode.pack(.{
.base = @intFromEnum(pci_class.BaseClass.display),
.subclass = 0x80, // "Other" — no named SubClass member (PCI convention)
.prog_if = 0,
});
/// The driver that serves a *reported* PCI function (M19.3: matching moved
/// from the boot snapshot to the bus reports), or null. A machine can carry
/// several identical controllers — one driver instance per reported device,
@@ -48,6 +57,7 @@ const xhci_pci_class: u64 = pci_class.ClassCode.pack(.{
fn pciDriverForIdentity(identity: u64) ?[]const u8 {
return switch (identity) {
xhci_pci_class => "usb-xhci-bus",
virtio_gpu_pci_class => "virtio-gpu",
else => null,
};
}
@@ -149,6 +159,8 @@ var test_restart_mode = false;
var test_usb_restart_mode = false;
var test_usb_killed = false;
var test_pci_restart_mode = false;
var test_scanout_restart_mode = false;
var test_scanout_killed = false;
var test_kill_pid: u32 = 0;
var test_kill_due_ns: u64 = 0;
@@ -415,6 +427,14 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime
} else if (driverByProcess(sender)) |driver| {
driver.state = .running;
writeLine("/system/services/device-manager: hello from {s} (device {d})\n", .{ driver.name(), hello.device_id });
// Resilience drill (V6): once, kill the virtio-gpu driver a moment after it hellos, so
// the normal restart policy respawns it — the compositor must survive and re-attach.
if (test_scanout_restart_mode and !test_scanout_killed and std.mem.eql(u8, driver.name(), "virtio-gpu")) {
test_scanout_killed = true;
test_kill_pid = sender;
test_kill_due_ns = system.clock() + 1_500_000_000;
_ = system.timerOnce(manager_endpoint, 1600);
}
} else {
status = -1;
writeLine("/system/services/device-manager: hello from unknown process {d}\n", .{sender});
@@ -557,6 +577,7 @@ pub fn main(init: runtime.process.Init) void {
test_restart_mode = std.mem.eql(u8, mode, "test-restart");
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart");
test_scanout_restart_mode = std.mem.eql(u8, mode, "test-scanout-restart");
}
runtime.service.run(protocol.message_maximum, .{
.service = .device_manager,
+273
View File
@@ -0,0 +1,273 @@
//! The compositor's **scanout backend** — how a finished frame reaches the panel
//! (docs/display-v2.md). The compositor composes its layer stack into the backend's
//! cacheable `surface()` and calls `present(damage)`; everything device-specific lives
//! here. Today there is one backend, `Gop` — the firmware framebuffer: a cacheable back
//! buffer streamed write-combining to the linear framebuffer. A native virtio-gpu backend
//! slots in beside it later (V4); the compositor never learns which is active.
const std = @import("std");
const runtime = @import("runtime");
const compositor = @import("compositor.zig");
const system = runtime.system;
const device = runtime.device;
const ipc = runtime.ipc;
const scanout_protocol = runtime.scanout_protocol;
const Rect = compositor.Rect;
const Surface = compositor.Surface;
/// The current display mode, as a backend reports it.
pub const Info = struct { width: u32, height: u32, pitch: u32, format: u32 };
/// Enumeration scratch — a `DeviceDescriptor` is large, and only one scan is ever needed.
var device_table: [64]device.DeviceDescriptor = undefined;
/// The GOP framebuffer backend: claims the kernel-seeded `display` device, maps the linear
/// framebuffer write-combining as the front buffer, and keeps a cacheable back buffer of
/// the same geometry as the compose target. `present` streams the damaged rectangle from
/// the back buffer to the LFB (sequential WC writes; the LFB is never read). No mode-set,
/// no vsync — the portable floor (docs/display-v2.md).
pub const Gop = struct {
device_id: u64,
front: [*]volatile u8, // the LFB (write-combining)
back: [*]u8, // cacheable compose target, same geometry
width: u32,
height: u32,
pitch: u32,
format: u32,
/// The framebuffer's id and geometry, captured together. `findDisplay` reads these out of
/// the enumeration table and returns them by value, so the caller never re-reads the table
/// across later syscalls (`device_enumerate` writes the whole table straight into this
/// process's memory; reading a descriptor's tail again after other syscalls have run is a
/// window we simply avoid by copying the few fields we need up front).
const Found = struct { id: u64, width: u32, height: u32, pitch: u32, format: u32 };
/// The first `display`-class device with a *valid* (non-zero) geometry, or null. A zero
/// geometry is treated as "not ready yet" so the caller retries — a real framebuffer always
/// has a non-zero width, height, and pitch.
fn findDisplay() ?Found {
const total = device.enumerate(&device_table);
const n = @min(total, device_table.len);
for (device_table[0..n]) |*d| {
if (d.class != @intFromEnum(device.DeviceClass.display)) continue;
if (d.display.width == 0 or d.display.height == 0 or d.display.pitch == 0) continue;
return .{ .id = d.id, .width = d.display.width, .height = d.display.height, .pitch = d.display.pitch, .format = d.display.format };
}
return null;
}
/// Claim the framebuffer (retrying while discovery catches up), map the LFB, and
/// allocate the back buffer. Null if there is no framebuffer or a mapping fails.
pub fn init() ?Gop {
var tries: u32 = 0;
const found = while (tries < 100) : (tries += 1) {
if (findDisplay()) |f| break f;
system.sleep(50);
} else {
_ = system.write("display: no framebuffer device (headless?)\n");
return null;
};
if (!device.claim(found.id)) {
_ = system.write("display: could not claim the framebuffer\n");
return null;
}
// Resource 0 is the framebuffer memory window; the kernel maps it write-combining
// because the resource carries that flag (docs/display-plan.md D1).
const front_base = device.mmioMap(found.id, 0) orelse {
_ = system.write("display: could not map the framebuffer\n");
return null;
};
const size = @as(usize, found.height) * found.pitch;
const back_base = system.mmap(size, system.PROT_READ | system.PROT_WRITE);
if (system.mmapFailed(back_base)) {
_ = system.write("display: could not allocate the back buffer\n");
return null;
}
return .{
.device_id = found.id,
.front = @ptrFromInt(front_base),
.back = @ptrFromInt(back_base),
.width = found.width,
.height = found.height,
.pitch = found.pitch,
.format = found.format,
};
}
pub fn info(self: *const Gop) Info {
return .{ .width = self.width, .height = self.height, .pitch = self.pitch, .format = self.format };
}
/// The cacheable compose target (the back buffer).
pub fn surface(self: *const Gop) Surface {
return .{
.pixels = @ptrCast(@alignCast(self.back)),
.stride = self.pitch / 4, // pitch is bytes; a 32-bpp row is pitch/4 pixels
.width = self.width,
.height = self.height,
};
}
/// Stream the damaged rectangle from the back buffer to the write-combining LFB, row by
/// row (sequential writes — what WC memory wants; the LFB is never read).
pub fn present(self: *const Gop, damage: Rect) void {
const c = damage.intersect(.{ .x = 0, .y = 0, .w = @intCast(self.width), .h = @intCast(self.height) });
if (c.isEmpty()) return;
var y: i32 = c.y;
while (y < c.bottom()) : (y += 1) {
const off = @as(usize, @intCast(y)) * self.pitch;
const src: [*]const u32 = @ptrCast(@alignCast(self.back + off));
const dst: [*]volatile u32 = @ptrCast(@alignCast(self.front + off));
var x: i32 = c.x;
while (x < c.right()) : (x += 1) dst[@intCast(x)] = src[@intCast(x)];
}
}
};
/// A display mode the native backend can switch to.
pub const Mode = scanout_protocol.Mode;
/// The native virtio-gpu backend: the compositor composes into a **shared** scanout surface
/// (an `shm` region the driver created and handed over) and `present` asks the driver to put
/// a frame on the panel over its `.scanout` endpoint. Unlike GOP there is no local copy — the
/// surface *is* the device's resource backing, so compositing writes land straight where the
/// driver transfers-and-flushes from (x86 DMA is cache-coherent, so the cacheable shared pages
/// need no explicit flush). Built by the display service when a driver announces (V4). The
/// surface is sized to the driver's largest mode, so `stride` (its row width) is fixed while
/// `width`/`height` — the active mode — change under `setMode` (V5).
pub const VirtioGpu = struct {
pixels: [*]u32, // the shared scanout surface, mapped into the compositor
stride: u32, // the surface's row stride in pixels (the driver's max mode width) — fixed
width: u32, // the active mode
height: u32,
format: u32,
scanout: ipc.Handle, // the driver's present + mode channel (looked up on `.scanout`)
pub fn info(self: *const VirtioGpu) Info {
return .{ .width = self.width, .height = self.height, .pitch = self.stride * 4, .format = self.format };
}
pub fn surface(self: *const VirtioGpu) Surface {
return .{ .pixels = self.pixels, .stride = self.stride, .width = self.width, .height = self.height };
}
/// Ask the driver to present. The composited pixels are already in the shared surface, so
/// this is a single request over `.scanout`; the driver transfers + fenced-flushes.
pub fn present(self: *const VirtioGpu, damage: Rect) void {
_ = damage;
var request = scanout_protocol.Request{
.operation = @intFromEnum(scanout_protocol.Operation.present),
.width = self.width,
.height = self.height,
};
var reply: [scanout_protocol.reply_size]u8 = undefined;
_ = ipc.call(self.scanout, std.mem.asBytes(&request), &reply) catch {};
}
/// Fill `out` with the driver's offered modes; returns how many were written.
pub fn modes(self: *const VirtioGpu, out: []Mode) usize {
var request = scanout_protocol.Request{ .operation = @intFromEnum(scanout_protocol.Operation.get_modes) };
var reply: [scanout_protocol.modes_reply_size]u8 = undefined;
const n = ipc.call(self.scanout, std.mem.asBytes(&request), &reply) catch return 0;
if (n < scanout_protocol.modes_reply_size) return 0;
const answer = std.mem.bytesToValue(scanout_protocol.ModesReply, reply[0..scanout_protocol.modes_reply_size]);
if (answer.status != 0) return 0;
const count = @min(@min(answer.count, scanout_protocol.max_modes), out.len);
for (0..count) |i| out[i] = answer.modes[i];
return count;
}
/// Change the scanout resolution. On success the active `width`/`height` update (the shared
/// surface — sized to the max mode — is unchanged, so `stride` stays put).
pub fn setMode(self: *VirtioGpu, w: u32, h: u32) bool {
if (w == 0 or h == 0 or w > self.stride) return false;
var request = scanout_protocol.Request{
.operation = @intFromEnum(scanout_protocol.Operation.set_mode),
.width = w,
.height = h,
};
var reply: [scanout_protocol.reply_size]u8 = undefined;
const n = ipc.call(self.scanout, std.mem.asBytes(&request), &reply) catch return false;
if (n < scanout_protocol.reply_size) return false;
if (std.mem.bytesToValue(scanout_protocol.Reply, reply[0..scanout_protocol.reply_size]).status != 0) return false;
self.width = w;
self.height = h;
return true;
}
};
/// The pluggable scanout backend. A tagged union so the compositor holds one value and
/// dispatches without caring which is active; the `virtio` native backend joins `gop` at V4.
pub const Backend = union(enum) {
gop: Gop,
virtio: VirtioGpu,
pub fn info(self: *const Backend) Info {
return switch (self.*) {
inline else => |*b| b.info(),
};
}
pub fn surface(self: *const Backend) Surface {
return switch (self.*) {
inline else => |*b| b.surface(),
};
}
pub fn present(self: *const Backend, damage: Rect) void {
switch (self.*) {
inline else => |*b| b.present(damage),
}
}
/// The modes this backend can switch to (none for GOP); returns how many were written.
pub fn modes(self: *const Backend, out: []Mode) usize {
return switch (self.*) {
.virtio => |*v| v.modes(out),
.gop => 0,
};
}
/// Change the resolution; false if this backend can't mode-set or the mode was refused.
pub fn setMode(self: *Backend, w: u32, h: u32) bool {
return switch (self.*) {
.virtio => |*v| v.setMode(w, h),
.gop => false,
};
}
/// Whether this backend supports runtime mode-setting (GOP: no; virtio-gpu: yes, V5).
pub fn canModeSet(self: *const Backend) bool {
return switch (self.*) {
.gop => false,
.virtio => true,
};
}
/// Whether this backend has a vblank/fence for tear-free present (virtio-gpu: yes, V5 — every
/// flush is fenced, so the device signals completion when the frame is actually on screen).
pub fn hasVsync(self: *const Backend) bool {
return switch (self.*) {
.gop => false,
.virtio => true,
};
}
};
/// Which backend to use. The pure selection *decision* is `chooseKind`; `select` below
/// binds it to the (syscall-bound) bring-up.
pub const Kind = enum { gop, virtio };
/// The selection decision, factored out of bring-up so it stays pure and host-testable:
/// prefer a native driver when one has announced itself (docs/display-v2.md V4), else the
/// GOP floor. Trivial today; it grows real inputs when native detection lands.
pub fn chooseKind(native_available: bool) Kind {
return if (native_available) .virtio else .gop;
}
/// Pick and bring up the best available backend. Today the GOP framebuffer is the only one
/// (`chooseKind(false)` → `.gop`), so this is `Gop.init()`. V4 adds the native-if-present
/// branch, with GOP as the floor.
pub fn select() ?Backend {
return switch (chooseKind(false)) {
.gop => .{ .gop = Gop.init() orelse return null },
.virtio => unreachable, // no native detection yet (V4)
};
}
test "selection prefers native when present, else the gop floor" {
try std.testing.expectEqual(Kind.gop, chooseKind(false));
try std.testing.expectEqual(Kind.virtio, chooseKind(true));
}
+199 -134
View File
@@ -1,43 +1,48 @@
//! /system/services/display — the display service (docs/display.md). A ring-3 process
//! that claims the framebuffer the kernel seeded (docs/display-plan.md D1), owns it as a
//! **write-combining front buffer**, composites an ordered stack of **layers** into a
//! **cacheable back buffer**, and presents finished frames — the GUI track's compositor,
//! the sibling of the input service. Reached by name over `ServiceId.display`.
//! /system/services/display — the display service (docs/display.md, docs/display-v2.md).
//! A ring-3 compositor: it composes an ordered stack of **layers** into a cacheable
//! surface and presents finished frames. Scanout — how a frame reaches the panel — is a
//! pluggable **backend** ([backend.zig](backend.zig)): the GOP framebuffer today, a native
//! virtio-gpu driver later; this file never learns which is active. It owns the layer stack
//! and damage tracking; the pixel math is the pure, host-tested
//! [compositor.zig](compositor.zig).
//!
//! A layer is a server-owned surface (its own cacheable buffer) with a screen position,
//! z-order, and visibility. Clients create layers and draw into them by command
//! (`fill_rect`, `blit_tile`), mark `damage`, and ask for a `present`; the compositor
//! repaints only the damaged region — clear it, paint the visible layers bottom-to-top,
//! flush it to the screen. The pixel math lives in the pure, host-tested
//! [compositor.zig](compositor.zig); this file wires real surfaces and the framebuffer to
//! it. Shared-memory client surfaces are a later milestone (docs/display.md).
//! z-order, and visibility. Clients create layers, draw into them by command (`fill_rect`,
//! `blit_tile`), mark `damage`, and ask for a `present`; the compositor repaints only the
//! damaged region — clear it, paint the visible layers bottom-to-top into the backend's
//! surface, then `backend.present(damage)`. Shared-memory client surfaces are later
//! (docs/display-v2.md).
const std = @import("std");
const runtime = @import("runtime");
const compositor = @import("compositor.zig");
const backend_mod = @import("backend.zig");
const protocol = runtime.display_protocol;
const ipc = runtime.ipc;
const system = runtime.system;
const device = runtime.device;
const Rect = compositor.Rect;
const Surface = compositor.Surface;
/// The claimed framebuffer and its off-screen twin. The front buffer is the LFB —
/// write-combining, so it is **only ever written**, never read; all compositing happens
/// in the cacheable back buffer, which is then streamed to the front (docs/display.md).
const Display = struct {
device_id: u64,
front: [*]volatile u8, // the LFB (write-combining)
back: [*]u8, // cacheable, same geometry
width: u32,
height: u32,
pitch: u32, // bytes per row (shared by both buffers)
format: u32, // a device-abi DisplayFormat value
frames: u64 = 0,
};
/// The active scanout backend — the GOP framebuffer at boot, upgraded to a native driver
/// (virtio-gpu) when one announces itself (V4).
var backend: backend_mod.Backend = undefined;
var frames: u64 = 0;
var display: Display = undefined;
/// This service's endpoint, kept so `attach_scanout` can arm a one-shot timer: the very first
/// native present must happen in a *later* loop iteration, after the reply to the driver's
/// announce has unblocked it and it is serving its `.scanout` channel — presenting inline
/// would deadlock (we'd call the driver while it waits on our reply).
var service_endpoint: ipc.Handle = 0;
/// Set when the backend has just been upgraded to virtio-gpu: the next present repaints the
/// whole screen into the shared surface and reads a pixel back to confirm the frame landed.
var pending_native_verify: bool = false;
/// Set alongside it: after the native present is verified, run the mode-set self-check once
/// (query the driver's modes, switch to a different one, confirm the geometry changed) — the
/// serial proof the runtime-resolution-change + fenced-present paths work (V5).
var pending_modeset_check: bool = false;
/// The wallpaper the compositor clears damaged regions to before painting layers.
var background: u32 = 0;
@@ -60,23 +65,11 @@ const Layer = struct {
var layers: [maximum_layers]Layer = [_]Layer{.{}} ** maximum_layers;
var damage: Rect = Rect.empty;
/// Enumeration buffer kept off the stack — a `DeviceDescriptor` is large, and this
/// service only ever needs one scan.
var device_table: [64]device.DeviceDescriptor = undefined;
// --- geometry helpers -------------------------------------------------------
fn screenRect() Rect {
return .{ .x = 0, .y = 0, .w = @intCast(display.width), .h = @intCast(display.height) };
}
fn backSurface() Surface {
return .{
.pixels = @ptrCast(@alignCast(display.back)),
.stride = display.pitch / 4, // pitch is bytes; a 32-bpp row is pitch/4 pixels
.width = display.width,
.height = display.height,
};
const m = backend.info();
return .{ .x = 0, .y = 0, .w = @intCast(m.width), .h = @intCast(m.height) };
}
fn layerScreenRect(l: *const Layer) Rect {
@@ -159,11 +152,11 @@ fn destroyLayer(id: u32) bool {
// --- compositing + present --------------------------------------------------
/// Repaint the damaged region `clip` of the back buffer: clear it to the background, then
/// paint every visible layer that overlaps it, bottom to top (ascending z).
/// Repaint the damaged region `clip` of the backend's compose surface: clear it to the
/// background, then paint every visible layer that overlaps it, bottom to top (ascending z).
fn compositeInto(clip: Rect) void {
const back = backSurface();
compositor.fillRect(back, clip, background);
const target = backend.surface();
compositor.fillRect(target, clip, background);
// z-order the used, visible layers (n ≤ 16; a plain insertion sort of indices).
var order: [maximum_layers]u32 = undefined;
@@ -184,56 +177,141 @@ fn compositeInto(clip: Rect) void {
for (order[0..n]) |i| {
const l = layers[i];
compositor.composite(back, l.x, l.y, l.surface, clip);
compositor.composite(target, l.x, l.y, l.surface, clip);
}
}
/// Stream the damaged rectangle from the cacheable back buffer to the write-combining
/// front buffer, row by row (sequential writes — what WC memory wants; we never read the
/// front buffer). Only the visible width of each row is touched.
fn flushRect(rect: Rect) void {
const c = rect.intersect(screenRect());
if (c.isEmpty()) return;
var y: i32 = c.y;
while (y < c.bottom()) : (y += 1) {
const off = @as(usize, @intCast(y)) * display.pitch;
const src: [*]const u32 = @ptrCast(@alignCast(display.back + off));
const dst: [*]volatile u32 = @ptrCast(@alignCast(display.front + off));
var x: i32 = c.x;
while (x < c.right()) : (x += 1) dst[@intCast(x)] = src[@intCast(x)];
}
}
/// Composite and flush the accumulated damage, then clear it. A no-op when nothing is
/// dirty. The frame counter advances regardless, so callers can name frames.
/// Composite the accumulated damage into the backend's surface, hand it to the backend to
/// put on screen, then clear the damage. A no-op when nothing is dirty. The frame counter
/// advances regardless, so callers can name frames.
fn present() void {
const dirty = damage.intersect(screenRect());
if (!dirty.isEmpty()) {
compositeInto(dirty);
flushRect(dirty);
backend.present(dirty);
}
damage = Rect.empty;
display.frames += 1;
frames += 1;
// The first present after a native upgrade confirms the composited frame actually reached
// the shared scanout surface (the automated stand-in for "it's on screen").
if (pending_native_verify and !dirty.isEmpty()) {
pending_native_verify = false;
verifyNativePresent();
}
}
/// Read a pixel straight back from the shared scanout surface after a native present. The
/// surface starts zeroed, so a non-zero centre pixel means the compositor wrote the frame into
/// the pages the driver transfers-and-flushes from — that, plus the driver acking the present
/// over `.scanout`, is the serial proof the native path works.
fn verifyNativePresent() void {
const s = backend.surface();
const sample = s.pixels[@as(usize, s.height / 2) * s.stride + s.width / 2];
if (sample != 0) {
_ = system.write("display: native present verified\n");
} else {
_ = system.write("display: native present FAILED (blank surface)\n");
}
}
/// A native scanout driver announced itself: map the shared surface it handed over, find its
/// present channel, switch the backend to virtio-gpu, and queue a full-screen repaint. The
/// present is deferred to a timer (see `service_endpoint`) so it happens after this reply
/// unblocks the driver and it starts serving `.scanout`.
fn attachScanout(stride: u32, width: u32, height: u32, format: u32, capability: ?ipc.Handle, reply: []u8) usize {
const cap = capability orelse return fail(reply);
if (width == 0 or height == 0 or stride < width) return fail(reply);
const mapped = runtime.shm.map(cap) orelse return fail(reply);
const scanout = ipc.lookup(.scanout) orelse return fail(reply);
// A second announce means the driver died and was restarted (V6): re-attach to its fresh
// scanout. (The previous shared mapping leaks — there is no shm_unmap syscall yet — but the
// frames are the dead driver's, reclaimed on its exit; a handful across a crash is benign.)
const reattach = switch (backend) {
.virtio => true,
else => false,
};
backend = .{ .virtio = .{
.pixels = @ptrCast(@alignCast(mapped)),
.stride = stride,
.width = width,
.height = height,
.format = format,
.scanout = scanout,
} };
background = protocol.pack(format, 0x20, 0x30, 0x48); // re-pack the wallpaper for the mode
addDamage(screenRect()); // the whole new surface must be painted
pending_native_verify = true;
if (!reattach) pending_modeset_check = true; // the mode-set self-check runs once, on first upgrade
_ = system.timerOnce(service_endpoint, 50); // present once the driver is serving .scanout
_ = system.write(if (reattach)
"display: scanout re-attached\n"
else
"display: scanout upgraded to virtio-gpu\n");
return ok(reply);
}
/// After the native upgrade is verified, prove the runtime-resolution-change and fenced-present
/// paths: query the driver's modes, switch to one that differs from the current, re-composite
/// the whole screen at the new size, and confirm the backend now reports that geometry. The
/// present goes through the driver's fenced flush, so a clean present is a vsync present.
fn modesetSelfCheck() void {
if (!backend.canModeSet()) return;
var mode_list: [4]backend_mod.Mode = undefined;
const count = backend.modes(&mode_list);
if (count == 0) {
_ = system.write("display: mode-set self-check: no modes reported\n");
return;
}
const current = backend.info();
var target: ?backend_mod.Mode = null;
for (mode_list[0..count]) |m| {
if (m.width != current.width or m.height != current.height) {
target = m;
break;
}
}
const wanted = target orelse {
_ = system.write("display: mode-set self-check: no alternate mode offered\n");
return;
};
if (!backend.setMode(wanted.width, wanted.height)) {
_ = system.write("display: mode set FAILED\n");
return;
}
addDamage(screenRect()); // repaint the whole screen at the new resolution, then present it
present();
const now = backend.info();
if (now.width == wanted.width and now.height == wanted.height) {
var line: [80]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "display: mode set to {d}x{d}, verified\n", .{ now.width, now.height }) catch "display: mode set, verified\n");
if (backend.hasVsync()) _ = system.write("display: vsync present ok\n");
} else {
_ = system.write("display: mode set FAILED (geometry unchanged)\n");
}
}
// --- startup self-check -----------------------------------------------------
/// Prove the compositor wiring on the real framebuffer: two overlapping opaque layers,
/// Prove the compositor wiring on the real backend: two overlapping opaque layers,
/// composited, must show the top layer in the overlap and the bottom layer outside it.
/// Exercises the whole path — mmap surfaces, the z-sort, damage, composite into the back
/// buffer — and reads the composited result back. Cleans up after itself.
/// Exercises the whole path — mmap surfaces, the z-sort, damage, composite into the
/// backend surface — and reads the composited result back. Cleans up after itself.
fn selfCheck() void {
const red = protocol.pack(display.format, 0xC0, 0x20, 0x20);
const green = protocol.pack(display.format, 0x20, 0xC0, 0x20);
const format = backend.info().format;
const red = protocol.pack(format, 0xC0, 0x20, 0x20);
const green = protocol.pack(format, 0x20, 0xC0, 0x20);
const bottom = createLayer(100, 100, 80, 80, 0, true) orelse return fail_check("create");
const top = createLayer(140, 140, 80, 80, 1, true) orelse return fail_check("create");
_ = fillLayer(bottom, Rect.init(0, 0, 80, 80), red);
_ = fillLayer(top, Rect.init(0, 0, 80, 80), green);
present();
const back = backSurface();
const overlap = back.pixels[@as(usize, 150) * back.stride + 150]; // in both layers → top
const bottom_only = back.pixels[@as(usize, 110) * back.stride + 110]; // bottom only
const surface = backend.surface();
const overlap = surface.pixels[@as(usize, 150) * surface.stride + 150]; // in both → top
const bottom_only = surface.pixels[@as(usize, 110) * surface.stride + 110]; // bottom only
_ = destroyLayer(top);
_ = destroyLayer(bottom);
@@ -252,67 +330,22 @@ fn fail_check(_: []const u8) void {
// --- service ----------------------------------------------------------------
/// The framebuffer node the kernel seeded (`DeviceClass.display`), or null if none.
fn findDisplay() ?device.DeviceDescriptor {
const total = device.enumerate(&device_table);
const n = @min(total, device_table.len);
for (device_table[0..n]) |d| {
if (d.class == @intFromEnum(device.DeviceClass.display)) return d;
}
return null;
}
fn initialise(endpoint: ipc.Handle) bool {
_ = endpoint;
service_endpoint = endpoint;
// Find the framebuffer, retrying while device discovery catches up with our spawn.
var tries: u32 = 0;
const found = while (tries < 100) : (tries += 1) {
if (findDisplay()) |d| break d;
system.sleep(50);
} else {
_ = system.write("display: no framebuffer device (headless?)\n");
return false; // clean exit: nothing to drive
};
// Pick the scanout backend (GOP today). It logs the reason on failure.
backend = backend_mod.select() orelse return false;
const mode = backend.info();
background = protocol.pack(mode.format, 0x20, 0x30, 0x48); // a dark slate wallpaper
if (!device.claim(found.id)) {
_ = system.write("display: could not claim the framebuffer\n");
return false;
}
// Resource 0 is the framebuffer memory window; the kernel maps it write-combining
// because the resource carries that flag (docs/display-plan.md D1).
const front_base = device.mmioMap(found.id, 0) orelse {
_ = system.write("display: could not map the framebuffer\n");
return false;
};
const geometry = found.display;
const size = @as(usize, geometry.height) * geometry.pitch;
const back_base = system.mmap(size, system.PROT_READ | system.PROT_WRITE);
if (system.mmapFailed(back_base)) {
_ = system.write("display: could not allocate the back buffer\n");
return false;
}
display = .{
.device_id = found.id,
.front = @ptrFromInt(front_base),
.back = @ptrFromInt(back_base),
.width = geometry.width,
.height = geometry.height,
.pitch = geometry.pitch,
.format = geometry.format,
};
background = protocol.pack(display.format, 0x20, 0x30, 0x48); // a dark slate wallpaper
// Clear the whole screen through the back buffer → present path (double buffering:
// no direct-to-LFB drawing).
// Clear the whole screen through the compose surface → present path (double buffering:
// no direct-to-scanout drawing).
addDamage(screenRect());
present();
var line: [96]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "display: online {d}x{d} pitch {d} format {d}\n", .{
display.width, display.height, display.pitch, display.format,
mode.width, mode.height, mode.pitch, mode.format,
}) catch "display: online\n");
_ = system.write("display: presented frame 0\n");
@@ -336,20 +369,16 @@ fn fail(reply: []u8) usize {
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize {
_ = sender;
_ = capability;
if (message.len < protocol.request_size) return fail(reply);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..];
// Switch on the raw operation value — an out-of-range one must fail cleanly, not
// panic an `@enumFromInt`.
switch (request.operation) {
@intFromEnum(protocol.Operation.info) => return writeReply(reply, .{
.status = 0,
.width = display.width,
.height = display.height,
.pitch = display.pitch,
.format = display.format,
}),
@intFromEnum(protocol.Operation.info) => {
const m = backend.info();
return writeReply(reply, .{ .status = 0, .width = m.width, .height = m.height, .pitch = m.pitch, .format = m.format });
},
@intFromEnum(protocol.Operation.create_layer) => {
// x/y are signed coordinates carried in the u32 wire fields — reinterpret the
// bits (@bitCast), don't range-check (@intCast) which a negative would fail.
@@ -379,15 +408,51 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Han
present();
return ok(reply);
},
@intFromEnum(protocol.Operation.attach_scanout) => {
return attachScanout(request.x, request.width, request.height, request.colour, capability, reply);
},
@intFromEnum(protocol.Operation.set_mode) => {
if (!backend.setMode(request.width, request.height)) return fail(reply);
addDamage(screenRect()); // repaint the whole screen at the new resolution
present();
return ok(reply);
},
@intFromEnum(protocol.Operation.get_modes) => {
var list: [4]backend_mod.Mode = undefined;
const count = backend.modes(&list);
var response = protocol.ModesReply{ .status = 0, .count = @intCast(count), .modes = undefined };
for (0..protocol.max_modes) |i| {
response.modes[i] = if (i < count)
.{ .width = list[i].width, .height = list[i].height }
else
.{ .width = 0, .height = 0 };
}
const bytes = std.mem.asBytes(&response);
@memcpy(reply[0..bytes.len], bytes);
return bytes.len;
},
else => return fail(reply),
}
}
/// The only notification the compositor arms is the post-attach present timer: repaint the
/// screen into the freshly attached native surface, verify the frame landed, then run the
/// one-shot mode-set self-check (V5).
fn onNotification(badge: u64) void {
_ = badge;
present(); // native present + verify (first timer fire after the upgrade)
if (pending_modeset_check) {
pending_modeset_check = false;
modesetSelfCheck();
}
}
pub fn main() void {
runtime.service.run(protocol.message_maximum, .{
.service = .display,
.init = initialise,
.on_message = onMessage,
.on_notification = onNotification,
});
}
+23
View File
@@ -24,6 +24,17 @@ pub const Operation = enum(u32) {
damage = 6,
/// present(): composite the dirty layers and flush to the screen.
present = 7,
/// attach_scanout(x=stride, width, height, colour=format) + <surface capability>: a native
/// scanout driver announces itself, handing over the shared scanout surface as an `ipc_call`
/// send_cap. The compositor maps it, looks up the driver's `.scanout` present channel, and
/// upgrades off the GOP floor (docs/display-v2.md V4). `x` is the surface's row stride in
/// pixels, `colour` the DisplayFormat.
attach_scanout = 8,
/// set_mode(width, height): change the display resolution — only a native backend that
/// reports `canModeSet` honours it; on the GOP floor it fails (docs/display-v2.md V5).
set_mode = 9,
/// get_modes() -> ModesReply: the resolutions the display can switch to (empty on GOP).
get_modes = 10,
};
/// The fixed request header. A `blit_tile`'s pixel payload (width*height 32-bit pixels)
@@ -54,6 +65,18 @@ pub const Reply = extern struct {
reserved2: u32 = 0,
};
/// One selectable display mode.
pub const Mode = extern struct { width: u32, height: u32 };
pub const max_modes = 4;
/// The reply to `get_modes`: a small fixed list of resolutions the display can switch to.
pub const ModesReply = extern struct {
status: i32,
count: u32,
modes: [max_modes]Mode,
};
pub const modes_reply_size: usize = @sizeOf(ModesReply);
/// The IPC message size — the kernel caps every message at `MESSAGE_MAXIMUM` (256 bytes,
/// system/kernel/ipc-synchronous.zig), so this matches it (a larger receive/reply buffer
/// is rejected with -E2BIG). A `blit_tile` therefore carries only a *small* tile inline —
@@ -0,0 +1,49 @@
//! The scanout wire protocol — what the compositor says to a native scanout driver (e.g.
//! virtio-gpu) over its well-known `.scanout` endpoint to put a composited frame on screen.
//! The driver owns the panel and the shared scanout surface it handed the compositor (via the
//! display service's `attach_scanout`); the compositor composites into that surface, then asks
//! the driver to present a damaged rectangle. Tiny by design — one present request. Separate
//! from the display protocol because the directions differ: clients call the compositor over
//! `.display`; the compositor calls the driver over `.scanout`. See docs/display-v2.md.
const std = @import("std");
pub const Operation = enum(u32) {
/// present(x, y, width, height): put the given rectangle of the shared scanout surface on
/// the panel (on virtio-gpu: transfer-to-host of the region, then a fenced resource flush).
present = 0,
/// get_modes() -> ModesReply: the display modes this scanout can switch to (V5).
get_modes = 1,
/// set_mode(width, height): change the scanout resolution — the shared surface is sized to
/// the largest mode, so this just re-points the scanout rectangle; the surface is unchanged.
set_mode = 2,
};
pub const Request = extern struct {
operation: u32,
x: u32 = 0,
y: u32 = 0,
width: u32 = 0,
height: u32 = 0,
};
pub const Reply = extern struct {
status: i32, // 0 on success, negative on failure
reserved: u32 = 0,
};
/// One offered display mode.
pub const Mode = extern struct { width: u32, height: u32 };
pub const max_modes = 4;
/// The reply to `get_modes`: a small fixed list of modes.
pub const ModesReply = extern struct {
status: i32,
count: u32,
modes: [max_modes]Mode,
};
pub const message_maximum: usize = 64;
pub const request_size: usize = @sizeOf(Request);
pub const reply_size: usize = @sizeOf(Reply);
pub const modes_reply_size: usize = @sizeOf(ModesReply);
+51
View File
@@ -0,0 +1,51 @@
//! system/services/shm-client — the creating half of the shm test (docs/display-v2.md V2).
//! It `shm_create`s a shared region, writes a known pattern into it, and hands the region's
//! capability to `shm-server` as an `ipc_call` send_cap. The server maps that capability and
//! confirms the pattern is visible — proving cross-process shared memory over the extended
//! capability-passing path.
const runtime = @import("runtime");
const system = runtime.system;
const shm = runtime.shm;
const ipc = runtime.ipc;
const pattern_len = 4096;
/// The pattern the server checks — must match shm-server.zig.
fn expected(i: usize) u8 {
return @truncate(i *% 7 +% 3);
}
fn lookupServer() ?ipc.Handle {
var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.shm_test)) |h| return h;
system.sleep(50);
}
return null;
}
pub fn main() void {
const region = shm.create(pattern_len) orelse {
_ = system.write("shm: create failed\n");
return;
};
var i: usize = 0;
while (i < pattern_len) : (i += 1) region.ptr[i] = expected(i);
const server = lookupServer() orelse {
_ = system.write("shm: no server\n");
return;
};
// A non-empty message (so it reaches on_message, not the ping path), carrying the shm
// region's capability. The reply is empty; we just need the round trip.
var reply: [64]u8 = undefined;
_ = ipc.callCap(server, "shm", &reply, region.handle) catch {
_ = system.write("shm: call failed\n");
};
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+49
View File
@@ -0,0 +1,49 @@
//! system/services/shm-server — the receiving half of the shm test (docs/display-v2.md V2).
//! It registers under `ServiceId.shm_test`; when `shm-client` calls it carrying a
//! shared-memory capability, it `shm_map`s that capability and checks the client's pattern
//! is visible through the mapping — proving the two processes share the same physical pages
//! (not a copy). On success it prints `shm: shared 4096 bytes ok`, the test's marker.
const runtime = @import("runtime");
const system = runtime.system;
const shm = runtime.shm;
const ipc = runtime.ipc;
const pattern_len = 4096;
/// The pattern the client writes — must match shm-client.zig.
fn expected(i: usize) u8 {
return @truncate(i *% 7 +% 3);
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
const cap = capability orelse {
_ = system.write("shm: shared FAILED (no capability)\n");
return 0;
};
const ptr = shm.map(cap) orelse {
_ = system.write("shm: shared FAILED (map)\n");
return 0;
};
var i: usize = 0;
while (i < pattern_len) : (i += 1) {
if (ptr[i] != expected(i)) {
_ = system.write("shm: shared FAILED (mismatch)\n");
return 0;
}
}
_ = system.write("shm: shared 4096 bytes ok\n");
return 0; // empty reply — the client only needs the round trip to unblock
}
pub fn main() void {
runtime.service.run(64, .{ .service = .shm_test, .on_message = onMessage });
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+61 -3
View File
@@ -185,6 +185,51 @@ CASES = [
{"name": "display-demo",
"expect": r"display-demo: scene up[\s\S]*display-demo: ok",
"fail": r"display-demo: (no display|create failed)|display: could not|CPU EXCEPTION|KERNEL PANIC"},
# Shared memory (v2 V2): shm-client creates a region, writes a pattern, and passes its
# capability to shm-server, which maps it and confirms the same bytes — proving
# cross-process shared pages over the extended capability passing.
{"name": "shm",
"expect": r"shm: shared 4096 bytes ok",
"fail": r"shm: (shared FAILED|create failed|no server|call failed|map)|CPU EXCEPTION|KERNEL PANIC"},
# virtio-gpu driver (v2 V3): boot with an emulated virtio-gpu. The device-manager stack
# discovers the PCI function and spawns the driver, which brings up the control virtqueue,
# creates a 2D scanout resource backed by DMA memory, set_scanouts it, paints a test
# pattern, transfers + flushes it, and waits for the device's used-ring ack, then reads
# the backing back. `scanout WxH online` + `flush acked, pixel check ok` are the markers.
{"name": "virtio-gpu",
"qemu_extra": ["-device", "virtio-gpu-pci"],
"expect": r"virtio-gpu: scanout \d+x\d+ online[\s\S]*virtio-gpu: flush acked, pixel check ok",
"fail": r"virtio-gpu:.*(failed|not acked|mismatch|unable to claim|not a virtio-gpu|too small|no PCI capability|does not offer|rejected|missing common-config|not a mapped resource|could not spawn)|CPU EXCEPTION|KERNEL PANIC"},
# Native backend + hot-attach (v2 V4): boot the compositor + display-demo with an emulated
# virtio-gpu. The driver announces its shared scanout surface to the compositor, which maps
# it, upgrades off the GOP floor, and drives frames through the native backend — reading a
# pixel back to confirm the composited frame reached the shared surface, while the demo runs.
{"name": "display-native",
"qemu_extra": ["-device", "virtio-gpu-pci"],
"mem": "512M", # boots the compositor + demo + the whole device-manager driver stack at once
# Order-independent: the demo's `ok` may print before or after the driver announces, so
# require all three markers to appear somewhere rather than in a fixed order.
"expect": r"(?s)(?=.*display: scanout upgraded to virtio-gpu)(?=.*display: native present verified)(?=.*display-demo: ok)",
"fail": r"display: native present FAILED|display: could not|display-demo: (no display|create failed)|CPU EXCEPTION|KERNEL PANIC"},
# Mode-set + EDID + vsync (v2 V5): same boot as display-native. After upgrading, the
# compositor queries the driver's modes, switches to a different resolution, and confirms the
# backend now reports it; the fenced present path makes it a vsync present. (The driver also
# logs the EDID preferred mode during bring-up.) Reuses the display-native kernel scenario.
{"name": "display-modeset",
"build_case": "display-native",
"qemu_extra": ["-device", "virtio-gpu-pci"],
"mem": "512M",
"expect": r"(?s)(?=.*display: mode set to \d+x\d+, verified)(?=.*display: vsync present ok)",
"fail": r"display: mode set FAILED|display: mode-set self-check: |display: native present FAILED|CPU EXCEPTION|KERNEL PANIC"},
# Resilience: driver restart + re-attach (v2 V6). device-manager (in test-scanout-restart
# mode) kills the virtio-gpu driver once after it hellos; the restart policy respawns it, it
# re-announces, and the compositor re-attaches — surviving the loss. Expect the initial
# upgrade AND the re-attach; any CPU exception / panic (the compositor crashing) is a fail.
{"name": "display-reattach",
"qemu_extra": ["-device", "virtio-gpu-pci"],
"mem": "512M",
"expect": r"(?s)(?=.*display: scanout upgraded to virtio-gpu)(?=.*display: scanout re-attached)",
"fail": r"CPU EXCEPTION|KERNEL PANIC|display: could not"},
# Monotonic clock (clock() syscall source): calibrated, advancing, never backwards.
{"name": "clock",
"expect": r"DANOS-TEST-RESULT: PASS",
@@ -441,12 +486,23 @@ CASES = [
"expect": r"acpi: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*"
r"acpi: reported PNP0F13 \(device \d+, 1 resources\)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M19.1: the ring-3 PCI scan (pci-bus walks the ECAM through its mmio_map
# grant) finds exactly the functions the kernel's own walk recorded.
# M19.1/M19.3: the ring-3 PCI scan. pci-bus walks the ECAM through its mmio_map
# grant and registers every function it finds; the kernel's own walk retired, so
# the broker starts empty and the driver populates it. The manager then runs the
# restart drill: ~1 s after the scan it kills pci-bus, prunes its child tree, and
# respawns it to re-claim, re-scan, and re-register the same functions. The kernel
# test asserts the broker equivalence (empty before, populated after, no
# duplicates); this ordered regex asserts the drill itself over the whole serial
# log — the backreference requires the respawn to re-scan the same count, and the
# full-capture match is immune to the transient-line races an in-kernel poll hits.
{"name": "pci-scan",
"smp": 4,
"timeout": 60,
"expect": r"DANOS-TEST-RESULT: PASS",
"expect": r"pci-bus: (\d+) functions found[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-manager: restarting pci-bus[\s\S]*"
r"pci-bus: \1 functions found[\s\S]*"
r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.3: the application surface — device-list enumerates the tree over IPC,
# subscribes (endpoint as capability), and observes the removed/added events
@@ -594,6 +650,8 @@ def run_case(arch, case):
cmd = [arch["qemu"]] + arch["qemu_args"](arch, boot_volume, vars_fd, serial)
if case.get("smp"): # some cases need more than one core (e.g. parallelism)
cmd += ["-smp", str(case["smp"])]
if case.get("mem"): # a case that boots the whole system at once needs more than the 128M floor
cmd[cmd.index("-m") + 1] = case["mem"]
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
cmd += case["qemu_extra"]
# A QMP control socket, always present (additive): how a case's `qmp_after`