Author SHA1 Message Date
Daniel Samson f3342118f5 docs: make display-v2 gates fully automated (unattended-safe)
Rewrite V3/V4/V5 gates so none need a screenshot: the driver/compositor read
their own pixels back (the scanout resource is shm-backed CPU-visible RAM), and a
virtio resource_flush is confirmed by the device's used-ring ack / fence. Pixel
readback + flush-ack is the serial stand-in for "it's on screen," so the whole
V1→V6 plan can run and self-verify without a human eyeballing QEMU.
2026-07-14 09:33:00 +01:00
Daniel Samson 2723b6f778 docs: display v2 design + plan (pluggable scanout backend)
Keep the GOP framebuffer as the floor, make scanout a pluggable backend, and
upgrade to a native virtio-gpu driver when it announces itself (dynamic
hot-attach; GOP stays the fallback for "no driver ever"). v2 builds the shm
cross-process memory capability, shared with the future client-surface path.
Milestones V1 (backend seam) → V6 (resilience + tests), each with a gate.
2026-07-14 09:28:31 +01:00
Daniel Samson a01a4f3b3d init: restart a crashed boot service
init supervises its boot services (spawned against an exit endpoint) but the
child-exit handler was `if (got.isNotification()) continue;` — it silently
dropped a dead service. Now init restarts it: on a child-exit notification it
finds the service, and unless it exited cleanly (chose to stop) or has hit the
crash-loop cap (maximum_restarts), respawns it and logs the death + reason. The
reincarnation half of resilience (docs/resilience.md) at the service level, the
counterpart to the device manager's driver restarts. A `shutting_down` flag
skips restarts during the orderly stop sequence, whose child deaths are expected.
2026-07-14 09:05:10 +01:00
Daniel Samson e1605e3235 kernel: contain a fatal fault — name the task, release the BKL before halt
Two gaps a real-hardware crash exposed, both in onException's terminal path (and
the panic path):

- The report was anonymous. Add the faulting task's id + name and whether it
  trapped in ring 3 (a user process) or ring 0 (the trusted base) — so a fatal
  fault says WHAT crashed and WHERE, not just the vector. scheduler gains
  currentIdSafe/currentNameSafe (early-boot-guarded, like currentCpuIndex, so the
  reporter can't fault a second time).

- halt() never released the big kernel lock, so a core that died holding it
  deadlocked every other core spinning in acquire() — the whole machine hangs,
  not just the one core the design promises. The BKL now records its owner
  (architecture.cpuLocal(), a unique per-core token); sync.releaseIfHeldHere()
  frees the lock only if this core holds it, called before halt on both fatal
  paths. Caveat: if we held it mid-mutation the shared state may be inconsistent,
  but letting the other cores + the supervisor keep running is strictly more
  recoverable than a guaranteed total hang.
2026-07-14 09:05:10 +01:00
Daniel Samson 1e80c57484 init: launch display-demo at boot
Spawn the display-demo client alongside the compositor so a normal boot shows
the moving scene (wallpaper + sliding rectangle + cursor) — the GUI track's
visible payoff. A demo fixture: drop it from boot_services to boot to a bare
compositor.
2026-07-14 08:26:55 +01:00
Daniel Samson c3e9c59086 kernel: route routine status to the log, not the framebuffer console
Now that the user-space display service owns the framebuffer, the bootstrap
console shrinks to fatal-only. status/statusPrint go to the diagnostic log
alone, so routine boot output no longer scribbles on a screen the compositor is
about to paint — and a driver's recoverable fault report stays off it too. A new
fatal/fatalPrint keeps panics and kernel-mode faults on screen, forcing the
console back on so a dying machine's last words show even over a live display.
console.zig now documents its early-boot + fatal-fallback role.
2026-07-14 08:26:55 +01:00
Daniel Samson 88ed3c5417 Merge claude/display-service-arch-a42261: display service v1 — framebuffer compositor (D1-D5)
A user-space display service: owns the write-combining framebuffer, composites a
z-ordered layer stack into a cacheable back buffer, presents only the damaged
region, and is driven over IPC by runtime.display. Proven end to end by the
display-demo client. Kernel changes: a display device node + write-combining
mmio_map, and a page-by-page mmap that lifts the 1 MiB per-call cap. Deferred by
design: shared-memory client surfaces, a native mode-setting backend, and vsync.
2026-07-14 07:57:43 +01:00
8 changed files with 417 additions and 41 deletions
+3 -1
View File
@@ -73,7 +73,9 @@ rather than restate it. Roughly in the order things happen at runtime:
track: a user-space compositor that owns the framebuffer, composes a layer stack into
a double buffer, and presents it. Why GOP and the PCI display device are two views of
one controller, the device-node + write-combining handoff, and what flicker-free buys
that tear-free doesn't. Plan: [display-plan.md](display-plan.md).
that tear-free doesn't. Plan: [display-plan.md](display-plan.md). **v2** makes scanout
a pluggable backend (GOP floor + a native virtio-gpu driver, hot-attached):
[display-v2.md](display-v2.md), plan [display-v2-plan.md](display-v2-plan.md).
20. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
+139
View File
@@ -0,0 +1,139 @@
# Display v2 — build plan (pluggable scanout: GOP floor + virtio-gpu native)
The ordered, checkpointable build-out for [display-v2.md](display-v2.md). Each milestone
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
[display-plan.md](display-plan.md). Read display-v2.md first for the *why*.
## Locked decisions (do not relitigate)
- **First native backend = virtio-gpu** (VM standard: mode-set + present/flush + vsync).
- **Dynamic hot-attach**: boot on GOP, upgrade to native when the driver **announces**
(push, not polling); re-attach across driver restarts; GOP is the floor for "no driver
ever," not a live fall-back after a reprogram.
- **v2 builds the `shm` capability** (endpoints → memory objects), shared with the future
client-surface path.
- The compositor's layers/back-buffer/damage are **unchanged**; only scanout is pluggable.
## Conventions
Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations,
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
`addUserBinary` and get packed into the initial-ramdisk; protocols are
`b.addModule("…-protocol", …)` imported into `runtime`; new syscalls extend
[abi.zig](../system/abi.zig) `SystemCall` + a `library/runtime` wrapper.
## How to verify along the way
**Every gate is serial-checkable — no screenshots** (this plan is built to run unattended).
Where "does it actually display" would otherwise need a human eyeball, the code **reads its
own pixels back**: the scanout resource is CPU-visible RAM (shm-backed) and the back buffer
is cacheable, so a driver/compositor can write a known value, read it back, and log a
pass/fail — and a virtio `resource_flush` is confirmed by the device **acking it on the
used ring**. Those two together (pixel-readback + flush-ack) are the automated stand-in for
"it's on screen."
- `zig build test` — host unit tests (backend selection, virtio struct sizes/encodings,
pixel-check helpers).
- `python3 test/qemu_test.py <case>` — boots the kernel in QEMU; asserts on serial markers.
The virtio cases boot with `-device virtio-gpu` (a per-case `qemu_extra`).
- `run-x86-64` renders to a window — for the human's own satisfaction, **not** a gate.
---
## V1 — The scanout backend seam (refactor, no behaviour change)
Extract scanout from the compositor so today's path becomes one backend among future ones.
- [ ] A `Backend` interface in `system/services/display/`: `surface() -> {ptr, pitch,
format, width, height}`, `present(damage: Rect)`, and capability flags
(`canModeSet`, `hasVsync`, both false for now).
- [ ] Wrap the v1 GOP path as `GopBackend` (claim the `display` node, WC-map the LFB,
`present` = the current damage-rect WC copy). The compositor composes into
`backend.surface()` and calls `backend.present(damage)` — no direct LFB references
left in the compositor core.
- [ ] Pure backend-selection logic factored so it's host-testable.
**Gate:** `display-service` + `display-demo` still pass unchanged (pure refactor; GOP is
the only backend), and `zig build test` stays green.
## V2 — The `shm` cross-process memory capability (kernel)
- [ ] [abi.zig](../system/abi.zig): `shm_create`, `shm_map` syscalls (+ a `ServiceId`/cap
convention if needed). Kernel handlers: `shm_create(len)` allocates page-aligned RAM,
returns a handle + maps it; passing the handle as an `ipc_call` `send_cap` shares it;
`shm_map(cap)` maps the same physical pages into the receiver. Reclaimed on death.
- [ ] `library/runtime/shm.zig` (+ barrel export): `create(len) -> Region{handle, ptr}`,
`map(cap) -> ptr`.
- [ ] Reuse the M13 capability-passing machinery (endpoints → memory objects).
**Gate:** a kernel/qemu `shm` test — process A `shm_create`s a region, writes a pattern,
passes the cap to process B, which `shm_map`s it and reads the same bytes back (proving
shared physical pages, not a copy). Heartbeat `shm: shared N bytes ok`.
## V3 — The virtio-gpu driver: bring-up + a frame on screen
- [ ] `system/drivers/virtio-gpu/`: claim the virtio-gpu PCI function (device-manager
match on its PCI/virtio id), map BARs, negotiate features, set up the control
virtqueue. `virtio-gpu-protocol.zig` for the control structs (host-tested sizes).
- [ ] Create a 2D scanout resource backed by an `shm` region, `attach_backing`,
`set_scanout` to CRTC 0, and `resource_flush` a test pattern.
- [ ] Register a `scanout` service (new `ServiceId`).
**Gate (automated):** a `virtio-gpu` case (QEMU `-device virtio-gpu`) where the driver
writes a known test pattern into the shm scanout resource, `resource_flush`es it, and
**waits for the device's used-ring ack**, then reads the resource back and checks the
pattern — logging `virtio-gpu: scanout {w}x{h} online` and `virtio-gpu: flush acked, pixel
check ok`. That proves virtqueue + resource + attach + set_scanout + flush end to end
without a screenshot (the used-ring ack is the device confirming it consumed the frame).
## V4 — The native backend + hot-attach
- [ ] `VirtioGpuBackend` in the compositor: `surface()` = the shared scanout resource,
`present(damage)` = `resource_flush` of the damaged rect.
- [ ] The driver **announces** to `.display` (looks it up, sends *attach-scanout* with its
`scanout` endpoint + the shared surface as capabilities). The compositor switches
backends and re-presents the current frame full-screen.
- [ ] Boot still starts on `GopBackend`; the upgrade happens on announce.
**Gate (automated):** boot with virtio-gpu + `display-demo`; the compositor logs
`display: scanout upgraded to virtio-gpu`, drives frames through the native backend, and
**reads a pixel back** from the shared scanout resource after a present to confirm the
composited frame landed (`display: native present verified`), while `display-demo: ok`
still fires. Without `-device virtio-gpu`, no `scanout` is announced and it stays on GOP —
the v1 `display-service`/`display-demo` gates still pass unchanged.
## V5 — Mode-setting, EDID, and vsync
- [ ] virtio-gpu `GET_EDID` → a mode list; `set_scanout` at a chosen mode = runtime
resolution change. `runtime.display` gains `modes()` / `setMode(m)`.
- [ ] A vsync/fenced `resource_flush` present path → genuinely tear-free.
- [ ] The compositor reports the native backend's `canModeSet`/`hasVsync` = true.
**Gate (automated):** a `display-modeset` case reads the EDID mode list, calls `setMode`
to a different resolution, and confirms the change by reading the driver's scanout geometry
back (`display: mode set to {w}x{h}, verified`); the vsync/fenced present path is exercised
and confirmed by the flush **fence completing** (`display: vsync present ok`) — both from
serial, no eyeballing.
## V6 — Resilience (restart + re-attach) + tests + docs
- [ ] The virtio-gpu driver is supervised (device-manager / init) and restartable; on
driver loss the compositor freezes the last frame and **re-attaches** when the driver
re-announces. Only a permanent give-up (crash-loop cap) attempts GOP again.
- [ ] `test/qemu_test.py` cases: `virtio-gpu`, hot-attach, `display-modeset`, and a
driver-kill/re-attach case. Update [display-v2.md](display-v2.md) status; README index.
**Gate:** kill the virtio-gpu driver mid-run; the compositor survives and re-attaches on
restart (`display: scanout re-attached`); all v1 + v2 cases pass; default `zig build` clean.
---
## Deferred (explicitly not in this plan)
- **Client-rendered surfaces** — now unblocked by the `shm` capability (V2): an app renders
its own bitmap and hands the compositor a reference. A natural follow-on.
- **Bochs DISPI backend** — a simpler second native backend (mode-set only, dumb scanout);
slots behind the same interface if wanted.
- **Real-GPU (NVIDIA/AMD/Intel) drivers** — out of scope; those devices stay on the GOP
floor by design.
- **Hardware-accelerated compositing / multiple heads** — future.
+126
View File
@@ -0,0 +1,126 @@
# The display service v2: a pluggable scanout backend
v1 ([display.md](display.md)) is a compositor that owns the **GOP framebuffer** — it
composites a layer stack into a cacheable back buffer and streams damage to the linear
framebuffer the firmware handed over. That path is portable and good: it drives any GPU,
including a real NVIDIA card at an ultrawide's native resolution, with zero GPU-specific
code. v2 keeps it as the **floor** and makes *scanout* — how a finished frame reaches the
panel — a **pluggable backend**, so the compositor can **upgrade to a real GPU driver when
one is present** and fall back to the framebuffer when it isn't.
The compositor itself (layers, back buffer, damage) does not change. Only the last step —
"put this frame on screen" — becomes swappable.
## The shape
```
compositor (display service) ── layer stack + back buffer + damage (unchanged)
│ composites a frame, then: backend.present(damage)
▼
scanout backend (selected at runtime — GOP by default, native when it appears)
│
├─ GopBackend the v1 path: WC copy back→front to the firmware LFB.
│ Always available. No mode-set, no vsync. THE FLOOR.
│
└─ VirtioGpuBackend talks to a virtio-gpu driver process over a `scanout`
service: present via a shared resource + flush (real vsync),
EDID mode list, runtime mode-set.
```
A **backend** is a small interface the compositor calls:
- `surface()` → the pixels to compose into and their geometry `{ptr, pitch, format, w, h}`
(the LFB for GOP; a shared scanout resource for virtio-gpu),
- `present(damage: Rect)` → make the damaged region visible (a no-op-ish WC copy for GOP;
a virtio flush, optionally vsync-fenced, for the native path),
- capability queries — `canModeSet`, `hasVsync` — and, when supported, `modes()` /
`setMode(m)`.
The compositor composes into `surface()` and calls `present(damage)` exactly as it does
today; everything device-specific lives behind the interface.
## Selection and hot-attach
The choice is **dynamic**, because a GPU driver is spawned asynchronously (the device
manager brings it up after boot), and because danos is meant to be resilient:
1. **Boot on GOP.** The compositor starts on `GopBackend` immediately, so there is never a
blank screen while drivers load — the exact v1 behaviour.
2. **Upgrade on announce.** When the virtio-gpu driver has claimed its device and set up a
scanout, it **announces itself to the display service** (a `push`: the driver looks up
`.display` and sends an *attach-scanout* message carrying its `scanout` endpoint as a
capability). The compositor switches to `VirtioGpuBackend` and re-presents the current
frame full-screen. Push beats polling — the compositor doesn't know a priori which
driver, if any, exists, and danos has no service-registration pub/sub.
3. **Native is restartable, not fallback-on-crash.** Once a native driver has reprogrammed
the device, the firmware's GOP framebuffer is **stale** — "native → GOP" is not a clean
fall-back. So a native driver that **crashes** is *restarted* by its supervisor (the
resilience work already merged), re-announces, and the compositor **re-attaches**
(native → native). The screen freezes on the last frame during the gap — acceptable.
4. **GOP is the floor for "no driver was ever there."** On a real GPU (NVIDIA/AMD/Intel)
the class-0x03 device matches nothing in the driver table, no `scanout` is ever
announced, and the compositor stays on GOP forever — no special-casing. Only if a
native driver *permanently* gives up (crash-loop cap) does the compositor attempt GOP
again, and even then only if the LFB is still mappable.
## The shared-memory primitive this needs
virtio-gpu's scanout resource is **guest RAM** — the driver allocates it and attaches it
to a virtio resource, and the compositor composes into it. That means the compositor
writing into the driver's buffer is **cross-process memory sharing**, the primitive v1
deferred (docs/display.md, "What v1 does not do"). v2 builds it: the natural generalization
of M13 capability-passing from *endpoints* to *memory objects* —
```
shm_create(len) -> {handle, vaddr} // a shareable, page-aligned RAM region
… pass `handle` as the send_cap on an ipc_call …
shm_map(cap) -> vaddr // the receiver maps the same physical pages
```
The payoff is leverage: the **same** primitive unlocks **both** native GPU drivers *and*
client-rendered surfaces (an app composing its own bitmap and handing the compositor a
reference instead of drawing by command). One piece of kernel work, two features.
## The virtio-gpu driver
A new ring-3 driver process (the topology v1 anticipated — "split the driver from the
compositor when a second backend arrives"). It claims the virtio-gpu PCI function, and:
- sets up the **virtqueues** (control + cursor) and the device's config space,
- creates a **2D scanout resource** backed by an `shm` region, `attach_backing`s it,
`set_scanout`s it to a CRTC, and `resource_flush`es damaged rectangles,
- reads **EDID** (the `GET_EDID` control command) for the mode list, and `set_scanout`
at a chosen mode for **runtime mode-setting**,
- registers a `scanout` service and announces to the display service.
Its `resource_flush` is the real **present** — and gives a genuine **vsync/tear-free**
path a dumb GOP framebuffer can't.
## What v2 unlocks — and its honest scope
Behind the abstraction, a native backend gives runtime **mode-setting** (resolution /
refresh / bpp), **EDID** enumeration, and **vsync**. But only on devices we have a driver
for — realistically **VMs** (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs
need per-vendor KMS-class drivers that aren't getting written, so they **stay on GOP** —
which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold:
the **pluggable architecture** (a driver slots in when one exists) and a **rich, vsync'd
path in VMs**, where danos development happens. The framebuffer floor never goes away.
## Locked decisions
- **First native backend: virtio-gpu** — the VM standard; gives mode-set + a real
present/flush (and vsync), and exercises the whole pluggable design. Tested with QEMU
`-device virtio-gpu`.
- **Dynamic hot-attach** — boot on GOP, upgrade to native on the driver's announce,
re-attach across driver restarts; GOP is the floor for "no driver ever," not a live
fall-back after a reprogram.
- **Detection = push** (the driver announces to `.display`), not compositor polling.
- **v2 builds the `shm` capability** (endpoints → memory objects), shared with the future
client-surface path.
## See also
- [display.md](display.md) — v1: the compositor, the GOP-vs-device split, the WC discipline.
- [display-v2-plan.md](display-v2-plan.md) — the ordered build-out.
- [driver-model.md](driver-model.md) — claim / `mmio_map` / MSI / capability passing (M13).
- [resilience.md](resilience.md) — the restart machinery the hot-attach leans on.
+9 -6
View File
@@ -2,12 +2,15 @@
//! into the linear framebuffer the bootloader handed us. No firmware, no driver
//! — just pixels.
//!
//! This is a **bootstrap** console — a stop-gap so early boot has something on
//! screen. The framebuffer is a general graphics surface, *not* inherently a text
//! terminal; once the driver machinery exists it becomes a proper graphics device
//! driver and this text-grid crutch goes away. It is therefore kept **separate
//! from the diagnostic [log](log.zig)** — the log fans out to serial/debugcon/file,
//! while this only paints the handful of user-facing status lines and panics.
//! This is a **bootstrap / fatal-fallback** console. The driver machinery now exists — the
//! user-space **display service** ([../services/display](../services/display/display.zig),
//! docs/display.md) owns the framebuffer in normal operation — so this no longer paints
//! routine status. It exists for the two cases the display service can't cover: **early
//! boot**, before the service has claimed the framebuffer, and **fatal errors** (a kernel
//! panic or a kernel-mode fault), which force it back on (`setSuppressed`) so a dying
//! machine's last words reach the screen even over a live display. It is kept **separate
//! from the diagnostic [log](log.zig)** — the log fans out to serial/debugcon/file and
//! carries all routine kernel output; this only paints those fatal cases.
//!
//! The module owns a single console and a `present` flag; `write` is a no-op when
//! the firmware handed over no framebuffer (a headless machine), so the kernel
+46 -21
View File
@@ -9,6 +9,7 @@ const wall_clock = @import("wall-clock.zig");
const pmm = @import("pmm.zig");
const heap = @import("heap.zig");
const scheduler = @import("scheduler.zig");
const sync = @import("sync.zig");
const process = @import("process.zig");
const devices_broker = @import("devices-broker.zig");
const irq = @import("irq.zig");
@@ -156,12 +157,13 @@ fn kmain(boot_information: *const BootInformation) noreturn {
log.print(" page tables: root = 0x{x:0>16}\n", .{architecture.activePageTable()});
log.print(" kernel segs: {d} (mapped with W^X permissions)\n", .{boot_information.kernel_segment_count});
// Now on our own tables, the framebuffer window is write-combining: bring up
// the on-screen console and clear it (a fast burst here, not the loader's
// uncached crawl). From here `status` reaches the screen as well as the log.
// Now on our own tables, the framebuffer window is write-combining: bring up the
// on-screen console and clear it to a blank canvas (a fast burst here, not the loader's
// uncached crawl). Routine boot output goes only to the log; this console now exists for
// early-boot and fatal (`fatal`/panic) output, until the display service takes over.
console.init(fb);
log.write(if (console.present())
"/system/kernel: framebuffer console online (bootstrap; graphics driver later)\n"
"/system/kernel: framebuffer ready (early-boot + fatal fallback; the display service drives it in normal operation)\n"
else
"/system/kernel: no framebuffer (headless) -> logging to serial/debugcon only\n");
@@ -414,11 +416,22 @@ fn bringUpSecondaries() void {
log.print("/system/kernel: {d}/{d} cores online\n", .{ scheduler.onlineCount(), cores.len });
}
/// A user-facing status line: to the diagnostic `log` *and* the on-screen console
/// (if a framebuffer is present). The verbose log uses `log.*` directly and never
/// touches the framebuffer.
/// A user-facing status line. Now that the user-space **display service** owns the
/// framebuffer in normal operation (docs/display.md), routine kernel output goes to the
/// diagnostic `log` (serial/debugcon/RAM) *only* — never to the on-screen console, which
/// the compositor is about to paint over. For a message that must reach the screen even so
/// — a panic or a fatal fault, when the machine is going down — use `fatal`.
fn status(message: []const u8) void {
log.write(message);
}
/// A fatal, user-facing message: to the diagnostic log *and* the on-screen console, forcing
/// the console back on (`setSuppressed(false)`) first — a dying machine's last words outrank
/// any display service holding the framebuffer. The console is otherwise silent in normal
/// operation (see `status`); it exists now only for early-boot and fatal output.
fn fatal(message: []const u8) void {
log.write(message);
console.setSuppressed(false);
console.write(message);
}
@@ -427,6 +440,11 @@ fn statusPrint(comptime fmt: []const u8, args: anytype) void {
status(std.fmt.bufPrint(&buffer, fmt, args) catch return);
}
fn fatalPrint(comptime fmt: []const u8, args: anytype) void {
var buffer: [256]u8 = undefined;
fatal(std.fmt.bufPrint(&buffer, fmt, args) catch return);
}
/// Frames (4 KiB pages) to whole MiB.
fn mib(pages: u64) u64 {
return pages * abi.page_size / (1024 * 1024);
@@ -486,20 +504,26 @@ fn onException(state: *const architecture.CpuState) noreturn {
}
log.checkpoint(cp_exception);
// The machine is going down: force the console back on even if a display service
// was holding the framebuffer, so the exception actually reaches the screen.
console.setSuppressed(false);
const core = scheduler.currentCpuIndex();
// A fault is user-facing enough to paint on screen too (via statusPrint), on
// top of the diagnostic log.
statusPrint("\nCPU EXCEPTION on core {d}: {s} (vector {d})\n", .{ core, architecture.exceptionName(state.vector), state.vector });
statusPrint(" error code : 0x{x}\n", .{state.error_code});
statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
statusPrint(" SP : 0x{x:0>16}\n", .{architecture.stackPointer(state)});
if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address});
// The machine is going down: paint the exception on screen too — `fatalPrint` forces the
// console back on even if a display service was holding the framebuffer — on top of the
// diagnostic log.
fatalPrint("\nCPU EXCEPTION on core {d}: {s} (vector {d})\n", .{ core, architecture.exceptionName(state.vector), state.vector });
// Name the culprit: which task, and whether it faulted in ring 3 (a process the
// kernel would normally kill — landing here means it had no address space) or ring 0
// (the trusted base itself). Without this the fatal report is anonymous.
fatalPrint(" task : {d} ({s}), {s}\n", .{ scheduler.currentIdSafe(), scheduler.currentNameSafe(), if (architecture.fromUser(state)) "ring 3 (user)" else "ring 0 (kernel)" });
fatalPrint(" error code : 0x{x}\n", .{state.error_code});
fatalPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
fatalPrint(" SP : 0x{x:0>16}\n", .{architecture.stackPointer(state)});
if (architecture.faultAddress(state)) |address| fatalPrint(" fault addr : 0x{x:0>16}\n", .{address});
var buffer: [128]u8 = undefined;
log.recordPanic(std.fmt.bufPrint(&buffer, "CPU exception {s} (vector {d}) on core {d} at IP 0x{x}", .{ architecture.exceptionName(state.vector), state.vector, core, architecture.instructionPointer(state) }) catch "cpu exception");
// Free the BKL if this core held it (a kernel-mode fault, or a nested fault in the
// recovery teardown), so halting this one core doesn't deadlock every other core on
// the lock. Only that core stops; the rest — and the supervisor — keep running.
sync.releaseIfHeldHere();
architecture.halt();
}
@@ -511,10 +535,11 @@ pub const panic = std.debug.FullPanic(struct {
_ = first_trace_address;
log.checkpoint(cp_panic);
log.recordPanic(message);
console.setSuppressed(false); // a panic outranks any display service holding the screen
status("\nKERNEL PANIC: ");
status(message);
status("\n");
fatal("\nKERNEL PANIC: "); // a panic outranks any display service holding the screen
fatal(message);
fatal("\n");
fatalPrint(" task : {d} ({s})\n", .{ scheduler.currentIdSafe(), scheduler.currentNameSafe() });
sync.releaseIfHeldHere(); // don't deadlock the other cores on the lock we may hold
architecture.halt();
}
}.panic);
+15
View File
@@ -791,6 +791,21 @@ pub fn currentCpuIndex() u32 {
return thisCpu().index;
}
/// The running task's id, or 0 if this core's scheduler isn't up yet (early boot, no GS
/// base). Safe for a fault reporter to call unconditionally — like `currentCpuIndex`,
/// it never dereferences an unpublished per-CPU pointer and so can't fault a second time.
pub fn currentIdSafe() u32 {
if (architecture.cpuLocal() == 0) return 0;
return thisCpu().current.id;
}
/// The running task's name (argv[0]), or "" if this core's scheduler isn't up yet.
/// The companion to `currentIdSafe` for naming the culprit in a fatal fault report.
pub fn currentNameSafe() []const u8 {
if (architecture.cpuLocal() == 0) return "";
return thisCpu().current.name();
}
/// Change the running task's priority (takes effect next time it's enqueued).
pub fn setPriority(p: Priority) void {
current().priority = p;
+22
View File
@@ -33,6 +33,13 @@ const architecture = @import("architecture");
/// 0 = free, 1 = held. A single global lock for the whole kernel.
var held = std.atomic.Value(u32).init(0);
/// The per-CPU base pointer (`architecture.cpuLocal()`) of the core currently holding
/// the lock, or 0 when free. Metadata only — `held` is what enforces exclusion — read
/// solely by `releaseIfHeldHere` on the fatal-fault path. `cpuLocal()` is a unique,
/// architecture-level token per core (0 before this core's GS base is published, which
/// is fine: that window is single-core early boot, where no other core can deadlock).
var owner = std.atomic.Value(usize).init(0);
/// Enter the kernel: disable interrupts on this core, then spin until we own the
/// lock. Returns the caller's prior interrupt flags for `leave` to restore.
/// Interrupts stay off for the whole critical section so this core's timer tick
@@ -67,14 +74,29 @@ export fn releaseForFreshTask() callconv(.c) void {
release();
}
/// Release the big kernel lock **only if this core is the one holding it** — a no-op
/// otherwise. For the fatal-fault path (a kernel-mode fault, or a nested fault inside the
/// recovery teardown, both of which run under the lock): a core that dies holding the BKL
/// must free it, or every other core spins forever in `acquire` and the whole machine
/// deadlocks instead of just that core stopping. It must NOT free a lock another core
/// owns, hence the owner check. Caveat: if we held it mid-mutation the shared state may be
/// inconsistent — but letting the other cores (and the supervisor) run on possibly-degraded
/// state is strictly more recoverable than a guaranteed total hang.
pub fn releaseIfHeldHere() void {
const me = architecture.cpuLocal();
if (me != 0 and owner.load(.monotonic) == me) release();
}
fn acquire() void {
// Test-and-test-and-set: try once, then spin read-only until the lock looks
// free before retrying the (bus-locked) swap — cheaper on the coherency fabric.
while (held.swap(1, .acquire) != 0) {
while (held.load(.monotonic) != 0) architecture.cpuRelax();
}
owner.store(architecture.cpuLocal(), .monotonic);
}
fn release() void {
owner.store(0, .monotonic);
held.store(0, .release);
}
+57 -13
View File
@@ -30,12 +30,22 @@ const log_path = "/mnt/usb/DANOS.LOG";
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent
/// on purpose: the device manager owns those. (A future init reads this from a
/// manifest under /system/services instead of a hardcoded list.)
const boot_services = [_][]const u8{ "vfs", "input", "device-manager", "fat", "display" };
const boot_services = [_][]const u8{ "vfs", "input", "device-manager", "fat", "display", "display-demo" };
var children: [boot_services.len]u32 = .{0} ** boot_services.len;
var child_count: usize = 0;
/// The live process id of each boot service (0 = not running), indexed by its position
/// in `boot_services`, plus how many times init has restarted it. init supervises these:
/// it spawns them against `supervision_endpoint` and, on a child's death, restarts it (up
/// to `maximum_restarts`) — the reincarnation half of resilience (docs/resilience.md), the
/// service-level counterpart to the device manager's driver restarts.
var child_ids: [boot_services.len]u32 = .{0} ** boot_services.len;
var restart_counts: [boot_services.len]u32 = .{0} ** boot_services.len;
var shutting_down = false;
var supervision_endpoint: runtime.ipc.Handle = 0;
/// Give up restarting a service after this many crashes — a crash-loop cap, so a service
/// that faults immediately on every spawn doesn't respawn forever.
const maximum_restarts = 3;
pub fn main() void {
// Prove the heap end to end: allocate through the runtime allocator (which
// mmaps pages from the kernel and carves them with the free list), write into
@@ -63,11 +73,8 @@ pub fn main() void {
// Bring up the boot services, supervised so init can stop them cleanly.
// Best-effort and silent: each service announces its own readiness, and in
// an isolation test with no initial-ramdisk the spawns simply no-op.
for (boot_services) |service| {
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| {
children[child_count] = id;
child_count += 1;
}
for (boot_services, 0..) |service, i| {
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| child_ids[i] = id;
}
// Once the storage stack is up, a one-shot copies the boot log to the USB
@@ -105,11 +112,47 @@ pub fn main() void {
if (receive[1] == @intFromEnum(power.Event.power_button)) shutDown();
continue;
}
// Child-exit notifications and anything else: keep waiting.
if (got.isChildExit()) {
restartChild(got.childProcessId());
continue;
}
// Anything else: keep waiting.
if (got.isNotification()) continue;
}
}
/// A supervised boot service died. Find which one and restart it — unless it exited
/// cleanly (it chose to stop, e.g. a driver with no hardware) or has hit the crash-loop
/// cap. Reclaiming the dead process is already the kernel's job (docs/process-lifecycle.md
/// iron rule 1); init only decides whether to bring it back.
fn restartChild(id: u32) void {
if (shutting_down) return; // deaths during the stop sequence are expected, not crashes
for (boot_services, 0..) |service, i| {
if (child_ids[i] != id) continue;
child_ids[i] = 0;
// An unknown reason (the record aged out) is treated as a crash worth restarting.
const reason = runtime.process.exitReason(id) orelse .fault;
if (reason == .exited) {
logLine("/system/services/init: {s} exited cleanly; not restarting\n", .{service});
return;
}
restart_counts[i] += 1;
if (restart_counts[i] > maximum_restarts) {
logLine("/system/services/init: {s} keeps crashing; giving up after {d} restarts\n", .{ service, maximum_restarts });
return;
}
logLine("/system/services/init: {s} died ({s}); restarting ({d}/{d})\n", .{ service, @tagName(reason), restart_counts[i], maximum_restarts });
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |new_id| child_ids[i] = new_id;
return;
}
// An untracked child (e.g. the log-flush one-shot): nothing to restart.
}
fn logLine(comptime fmt: []const u8, args: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, args) catch return);
}
/// Look up the power service and subscribe our endpoint (handed over as the
/// call's capability) so events arrive as buffered messages here.
fn subscribePower() void {
@@ -153,15 +196,16 @@ fn flushKernelLog() void {
/// it), waiting up to a deadline for each to exit before killing it, then ask the
/// power service to enter S5.
fn shutDown() void {
shutting_down = true; // the stop loop below kills children — those deaths aren't crashes
_ = runtime.system.write("/system/services/init: shutting down\n");
// Persist the fullest log to the USB volume BEFORE tearing anything down: the
// reverse-order stop loop below kills the fat server (children[3]) first, so
// /mnt/usb must be written while it is still mounted.
// reverse-order stop loop below kills the fat server first, so /mnt/usb must be
// written while it is still mounted.
flushKernelLog();
var i = child_count;
var i = boot_services.len;
while (i > 0) {
i -= 1;
if (children[i] != 0) runtime.process.stop(children[i], 2000, supervision_endpoint);
if (child_ids[i] != 0) runtime.process.stop(child_ids[i], 2000, supervision_endpoint);
}
if (runtime.ipc.lookup(.power)) |h| {
const request = power.Shutdown{};