display: resilient scanout — supervised driver, survive loss, re-attach (v2 V6)
The compositor now survives the virtio-gpu driver dying and re-attaches when device-manager restarts it — the last piece of display v2. Three parts: - The driver hellos the device manager (role: bus). It never did, so the manager — which spawns it supervised and expects a hello — was stopping it at the 3s hello deadline every run (the gate markers just printed first). Now it is properly supervised: not stopped for silence, and restarted on death. - A kernel IPC fix so a call to a dead service errors instead of hanging. An endpoint records its owner; when that task dies, its registered endpoints are marked dead (and any parked senders woken with -EPEER), so ipc_call returns -EPEER rather than blocking on a reply that will never come. Without this the compositor's first present after the driver died blocked forever. General robustness — any client of any service benefits. - The compositor re-attaches. Its .scanout calls now fail cleanly (caught), freezing the last frame; when the restarted driver re-announces, attach_scanout detects the backend is already virtio and logs "scanout re-attached", mapping the fresh shared surface and re-looking-up .scanout. (The previous shm mapping leaks — no shm_unmap syscall yet — but its frames are the dead driver's, reclaimed on exit.) - device-manager gains a "test-scanout-restart" mode (like test-usb-restart) that kills the virtio-gpu driver once after it hellos; the displayReattachTest kernel scenario drives it. Gate: python3 test/qemu_test.py display-reattach — "scanout upgraded to virtio-gpu" then "scanout re-attached", no CPU exception, passing 3/3. host tests, ipc/ipc-call/ipc-cap, supervision, shm, display-service, display-demo, virtio-gpu, display-native, and display-modeset all pass; default zig build clean. v2 (V1-V6) complete.
This commit is contained in:
@@ -737,6 +737,7 @@ fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
|
||||
scheduler.readyLocked(client); // its blocked `call` now returns the error
|
||||
}
|
||||
ipc.abandonSenderLocked(t);
|
||||
ipc.killOwnedEndpointsLocked(t.id); // its registered services are gone: callers get -EPEER, not a hang
|
||||
scheduler.removeFromWaitQueueLocked(t);
|
||||
scheduler.forgetIpcClientLocked(t);
|
||||
ipc.closeHandles(t);
|
||||
|
||||
Reference in New Issue
Block a user