Compare commits
7
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f3342118f5 | ||
|
|
2723b6f778 | ||
|
|
a01a4f3b3d | ||
|
|
e1605e3235 | ||
|
|
1e80c57484 | ||
|
|
c3e9c59086 | ||
|
|
88ed3c5417 |
+3
-1
@@ -73,7 +73,9 @@ rather than restate it. Roughly in the order things happen at runtime:
|
|||||||
track: a user-space compositor that owns the framebuffer, composes a layer stack into
|
track: a user-space compositor that owns the framebuffer, composes a layer stack into
|
||||||
a double buffer, and presents it. Why GOP and the PCI display device are two views of
|
a double buffer, and presents it. Why GOP and the PCI display device are two views of
|
||||||
one controller, the device-node + write-combining handoff, and what flicker-free buys
|
one controller, the device-node + write-combining handoff, and what flicker-free buys
|
||||||
that tear-free doesn't. Plan: [display-plan.md](display-plan.md).
|
that tear-free doesn't. Plan: [display-plan.md](display-plan.md). **v2** makes scanout
|
||||||
|
a pluggable backend (GOP floor + a native virtio-gpu driver, hot-attached):
|
||||||
|
[display-v2.md](display-v2.md), plan [display-v2-plan.md](display-v2-plan.md).
|
||||||
20. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
20. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,139 @@
|
|||||||
|
# Display v2 — build plan (pluggable scanout: GOP floor + virtio-gpu native)
|
||||||
|
|
||||||
|
The ordered, checkpointable build-out for [display-v2.md](display-v2.md). Each milestone
|
||||||
|
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
|
||||||
|
[display-plan.md](display-plan.md). Read display-v2.md first for the *why*.
|
||||||
|
|
||||||
|
## Locked decisions (do not relitigate)
|
||||||
|
|
||||||
|
- **First native backend = virtio-gpu** (VM standard: mode-set + present/flush + vsync).
|
||||||
|
- **Dynamic hot-attach**: boot on GOP, upgrade to native when the driver **announces**
|
||||||
|
(push, not polling); re-attach across driver restarts; GOP is the floor for "no driver
|
||||||
|
ever," not a live fall-back after a reprogram.
|
||||||
|
- **v2 builds the `shm` capability** (endpoints → memory objects), shared with the future
|
||||||
|
client-surface path.
|
||||||
|
- The compositor's layers/back-buffer/damage are **unchanged**; only scanout is pluggable.
|
||||||
|
|
||||||
|
## Conventions
|
||||||
|
|
||||||
|
Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations,
|
||||||
|
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
|
||||||
|
`addUserBinary` and get packed into the initial-ramdisk; protocols are
|
||||||
|
`b.addModule("…-protocol", …)` imported into `runtime`; new syscalls extend
|
||||||
|
[abi.zig](../system/abi.zig) `SystemCall` + a `library/runtime` wrapper.
|
||||||
|
|
||||||
|
## How to verify along the way
|
||||||
|
|
||||||
|
**Every gate is serial-checkable — no screenshots** (this plan is built to run unattended).
|
||||||
|
Where "does it actually display" would otherwise need a human eyeball, the code **reads its
|
||||||
|
own pixels back**: the scanout resource is CPU-visible RAM (shm-backed) and the back buffer
|
||||||
|
is cacheable, so a driver/compositor can write a known value, read it back, and log a
|
||||||
|
pass/fail — and a virtio `resource_flush` is confirmed by the device **acking it on the
|
||||||
|
used ring**. Those two together (pixel-readback + flush-ack) are the automated stand-in for
|
||||||
|
"it's on screen."
|
||||||
|
|
||||||
|
- `zig build test` — host unit tests (backend selection, virtio struct sizes/encodings,
|
||||||
|
pixel-check helpers).
|
||||||
|
- `python3 test/qemu_test.py <case>` — boots the kernel in QEMU; asserts on serial markers.
|
||||||
|
The virtio cases boot with `-device virtio-gpu` (a per-case `qemu_extra`).
|
||||||
|
- `run-x86-64` renders to a window — for the human's own satisfaction, **not** a gate.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## V1 — The scanout backend seam (refactor, no behaviour change)
|
||||||
|
|
||||||
|
Extract scanout from the compositor so today's path becomes one backend among future ones.
|
||||||
|
|
||||||
|
- [ ] A `Backend` interface in `system/services/display/`: `surface() -> {ptr, pitch,
|
||||||
|
format, width, height}`, `present(damage: Rect)`, and capability flags
|
||||||
|
(`canModeSet`, `hasVsync`, both false for now).
|
||||||
|
- [ ] Wrap the v1 GOP path as `GopBackend` (claim the `display` node, WC-map the LFB,
|
||||||
|
`present` = the current damage-rect WC copy). The compositor composes into
|
||||||
|
`backend.surface()` and calls `backend.present(damage)` — no direct LFB references
|
||||||
|
left in the compositor core.
|
||||||
|
- [ ] Pure backend-selection logic factored so it's host-testable.
|
||||||
|
|
||||||
|
**Gate:** `display-service` + `display-demo` still pass unchanged (pure refactor; GOP is
|
||||||
|
the only backend), and `zig build test` stays green.
|
||||||
|
|
||||||
|
## V2 — The `shm` cross-process memory capability (kernel)
|
||||||
|
|
||||||
|
- [ ] [abi.zig](../system/abi.zig): `shm_create`, `shm_map` syscalls (+ a `ServiceId`/cap
|
||||||
|
convention if needed). Kernel handlers: `shm_create(len)` allocates page-aligned RAM,
|
||||||
|
returns a handle + maps it; passing the handle as an `ipc_call` `send_cap` shares it;
|
||||||
|
`shm_map(cap)` maps the same physical pages into the receiver. Reclaimed on death.
|
||||||
|
- [ ] `library/runtime/shm.zig` (+ barrel export): `create(len) -> Region{handle, ptr}`,
|
||||||
|
`map(cap) -> ptr`.
|
||||||
|
- [ ] Reuse the M13 capability-passing machinery (endpoints → memory objects).
|
||||||
|
|
||||||
|
**Gate:** a kernel/qemu `shm` test — process A `shm_create`s a region, writes a pattern,
|
||||||
|
passes the cap to process B, which `shm_map`s it and reads the same bytes back (proving
|
||||||
|
shared physical pages, not a copy). Heartbeat `shm: shared N bytes ok`.
|
||||||
|
|
||||||
|
## V3 — The virtio-gpu driver: bring-up + a frame on screen
|
||||||
|
|
||||||
|
- [ ] `system/drivers/virtio-gpu/`: claim the virtio-gpu PCI function (device-manager
|
||||||
|
match on its PCI/virtio id), map BARs, negotiate features, set up the control
|
||||||
|
virtqueue. `virtio-gpu-protocol.zig` for the control structs (host-tested sizes).
|
||||||
|
- [ ] Create a 2D scanout resource backed by an `shm` region, `attach_backing`,
|
||||||
|
`set_scanout` to CRTC 0, and `resource_flush` a test pattern.
|
||||||
|
- [ ] Register a `scanout` service (new `ServiceId`).
|
||||||
|
|
||||||
|
**Gate (automated):** a `virtio-gpu` case (QEMU `-device virtio-gpu`) where the driver
|
||||||
|
writes a known test pattern into the shm scanout resource, `resource_flush`es it, and
|
||||||
|
**waits for the device's used-ring ack**, then reads the resource back and checks the
|
||||||
|
pattern — logging `virtio-gpu: scanout {w}x{h} online` and `virtio-gpu: flush acked, pixel
|
||||||
|
check ok`. That proves virtqueue + resource + attach + set_scanout + flush end to end
|
||||||
|
without a screenshot (the used-ring ack is the device confirming it consumed the frame).
|
||||||
|
|
||||||
|
## V4 — The native backend + hot-attach
|
||||||
|
|
||||||
|
- [ ] `VirtioGpuBackend` in the compositor: `surface()` = the shared scanout resource,
|
||||||
|
`present(damage)` = `resource_flush` of the damaged rect.
|
||||||
|
- [ ] The driver **announces** to `.display` (looks it up, sends *attach-scanout* with its
|
||||||
|
`scanout` endpoint + the shared surface as capabilities). The compositor switches
|
||||||
|
backends and re-presents the current frame full-screen.
|
||||||
|
- [ ] Boot still starts on `GopBackend`; the upgrade happens on announce.
|
||||||
|
|
||||||
|
**Gate (automated):** boot with virtio-gpu + `display-demo`; the compositor logs
|
||||||
|
`display: scanout upgraded to virtio-gpu`, drives frames through the native backend, and
|
||||||
|
**reads a pixel back** from the shared scanout resource after a present to confirm the
|
||||||
|
composited frame landed (`display: native present verified`), while `display-demo: ok`
|
||||||
|
still fires. Without `-device virtio-gpu`, no `scanout` is announced and it stays on GOP —
|
||||||
|
the v1 `display-service`/`display-demo` gates still pass unchanged.
|
||||||
|
|
||||||
|
## V5 — Mode-setting, EDID, and vsync
|
||||||
|
|
||||||
|
- [ ] virtio-gpu `GET_EDID` → a mode list; `set_scanout` at a chosen mode = runtime
|
||||||
|
resolution change. `runtime.display` gains `modes()` / `setMode(m)`.
|
||||||
|
- [ ] A vsync/fenced `resource_flush` present path → genuinely tear-free.
|
||||||
|
- [ ] The compositor reports the native backend's `canModeSet`/`hasVsync` = true.
|
||||||
|
|
||||||
|
**Gate (automated):** a `display-modeset` case reads the EDID mode list, calls `setMode`
|
||||||
|
to a different resolution, and confirms the change by reading the driver's scanout geometry
|
||||||
|
back (`display: mode set to {w}x{h}, verified`); the vsync/fenced present path is exercised
|
||||||
|
and confirmed by the flush **fence completing** (`display: vsync present ok`) — both from
|
||||||
|
serial, no eyeballing.
|
||||||
|
|
||||||
|
## V6 — Resilience (restart + re-attach) + tests + docs
|
||||||
|
|
||||||
|
- [ ] The virtio-gpu driver is supervised (device-manager / init) and restartable; on
|
||||||
|
driver loss the compositor freezes the last frame and **re-attaches** when the driver
|
||||||
|
re-announces. Only a permanent give-up (crash-loop cap) attempts GOP again.
|
||||||
|
- [ ] `test/qemu_test.py` cases: `virtio-gpu`, hot-attach, `display-modeset`, and a
|
||||||
|
driver-kill/re-attach case. Update [display-v2.md](display-v2.md) status; README index.
|
||||||
|
|
||||||
|
**Gate:** kill the virtio-gpu driver mid-run; the compositor survives and re-attaches on
|
||||||
|
restart (`display: scanout re-attached`); all v1 + v2 cases pass; default `zig build` clean.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Deferred (explicitly not in this plan)
|
||||||
|
|
||||||
|
- **Client-rendered surfaces** — now unblocked by the `shm` capability (V2): an app renders
|
||||||
|
its own bitmap and hands the compositor a reference. A natural follow-on.
|
||||||
|
- **Bochs DISPI backend** — a simpler second native backend (mode-set only, dumb scanout);
|
||||||
|
slots behind the same interface if wanted.
|
||||||
|
- **Real-GPU (NVIDIA/AMD/Intel) drivers** — out of scope; those devices stay on the GOP
|
||||||
|
floor by design.
|
||||||
|
- **Hardware-accelerated compositing / multiple heads** — future.
|
||||||
@@ -0,0 +1,126 @@
|
|||||||
|
# The display service v2: a pluggable scanout backend
|
||||||
|
|
||||||
|
v1 ([display.md](display.md)) is a compositor that owns the **GOP framebuffer** — it
|
||||||
|
composites a layer stack into a cacheable back buffer and streams damage to the linear
|
||||||
|
framebuffer the firmware handed over. That path is portable and good: it drives any GPU,
|
||||||
|
including a real NVIDIA card at an ultrawide's native resolution, with zero GPU-specific
|
||||||
|
code. v2 keeps it as the **floor** and makes *scanout* — how a finished frame reaches the
|
||||||
|
panel — a **pluggable backend**, so the compositor can **upgrade to a real GPU driver when
|
||||||
|
one is present** and fall back to the framebuffer when it isn't.
|
||||||
|
|
||||||
|
The compositor itself (layers, back buffer, damage) does not change. Only the last step —
|
||||||
|
"put this frame on screen" — becomes swappable.
|
||||||
|
|
||||||
|
## The shape
|
||||||
|
|
||||||
|
```
|
||||||
|
compositor (display service) ── layer stack + back buffer + damage (unchanged)
|
||||||
|
│ composites a frame, then: backend.present(damage)
|
||||||
|
▼
|
||||||
|
scanout backend (selected at runtime — GOP by default, native when it appears)
|
||||||
|
│
|
||||||
|
├─ GopBackend the v1 path: WC copy back→front to the firmware LFB.
|
||||||
|
│ Always available. No mode-set, no vsync. THE FLOOR.
|
||||||
|
│
|
||||||
|
└─ VirtioGpuBackend talks to a virtio-gpu driver process over a `scanout`
|
||||||
|
service: present via a shared resource + flush (real vsync),
|
||||||
|
EDID mode list, runtime mode-set.
|
||||||
|
```
|
||||||
|
|
||||||
|
A **backend** is a small interface the compositor calls:
|
||||||
|
|
||||||
|
- `surface()` → the pixels to compose into and their geometry `{ptr, pitch, format, w, h}`
|
||||||
|
(the LFB for GOP; a shared scanout resource for virtio-gpu),
|
||||||
|
- `present(damage: Rect)` → make the damaged region visible (a no-op-ish WC copy for GOP;
|
||||||
|
a virtio flush, optionally vsync-fenced, for the native path),
|
||||||
|
- capability queries — `canModeSet`, `hasVsync` — and, when supported, `modes()` /
|
||||||
|
`setMode(m)`.
|
||||||
|
|
||||||
|
The compositor composes into `surface()` and calls `present(damage)` exactly as it does
|
||||||
|
today; everything device-specific lives behind the interface.
|
||||||
|
|
||||||
|
## Selection and hot-attach
|
||||||
|
|
||||||
|
The choice is **dynamic**, because a GPU driver is spawned asynchronously (the device
|
||||||
|
manager brings it up after boot), and because danos is meant to be resilient:
|
||||||
|
|
||||||
|
1. **Boot on GOP.** The compositor starts on `GopBackend` immediately, so there is never a
|
||||||
|
blank screen while drivers load — the exact v1 behaviour.
|
||||||
|
2. **Upgrade on announce.** When the virtio-gpu driver has claimed its device and set up a
|
||||||
|
scanout, it **announces itself to the display service** (a `push`: the driver looks up
|
||||||
|
`.display` and sends an *attach-scanout* message carrying its `scanout` endpoint as a
|
||||||
|
capability). The compositor switches to `VirtioGpuBackend` and re-presents the current
|
||||||
|
frame full-screen. Push beats polling — the compositor doesn't know a priori which
|
||||||
|
driver, if any, exists, and danos has no service-registration pub/sub.
|
||||||
|
3. **Native is restartable, not fallback-on-crash.** Once a native driver has reprogrammed
|
||||||
|
the device, the firmware's GOP framebuffer is **stale** — "native → GOP" is not a clean
|
||||||
|
fall-back. So a native driver that **crashes** is *restarted* by its supervisor (the
|
||||||
|
resilience work already merged), re-announces, and the compositor **re-attaches**
|
||||||
|
(native → native). The screen freezes on the last frame during the gap — acceptable.
|
||||||
|
4. **GOP is the floor for "no driver was ever there."** On a real GPU (NVIDIA/AMD/Intel)
|
||||||
|
the class-0x03 device matches nothing in the driver table, no `scanout` is ever
|
||||||
|
announced, and the compositor stays on GOP forever — no special-casing. Only if a
|
||||||
|
native driver *permanently* gives up (crash-loop cap) does the compositor attempt GOP
|
||||||
|
again, and even then only if the LFB is still mappable.
|
||||||
|
|
||||||
|
## The shared-memory primitive this needs
|
||||||
|
|
||||||
|
virtio-gpu's scanout resource is **guest RAM** — the driver allocates it and attaches it
|
||||||
|
to a virtio resource, and the compositor composes into it. That means the compositor
|
||||||
|
writing into the driver's buffer is **cross-process memory sharing**, the primitive v1
|
||||||
|
deferred (docs/display.md, "What v1 does not do"). v2 builds it: the natural generalization
|
||||||
|
of M13 capability-passing from *endpoints* to *memory objects* —
|
||||||
|
|
||||||
|
```
|
||||||
|
shm_create(len) -> {handle, vaddr} // a shareable, page-aligned RAM region
|
||||||
|
… pass `handle` as the send_cap on an ipc_call …
|
||||||
|
shm_map(cap) -> vaddr // the receiver maps the same physical pages
|
||||||
|
```
|
||||||
|
|
||||||
|
The payoff is leverage: the **same** primitive unlocks **both** native GPU drivers *and*
|
||||||
|
client-rendered surfaces (an app composing its own bitmap and handing the compositor a
|
||||||
|
reference instead of drawing by command). One piece of kernel work, two features.
|
||||||
|
|
||||||
|
## The virtio-gpu driver
|
||||||
|
|
||||||
|
A new ring-3 driver process (the topology v1 anticipated — "split the driver from the
|
||||||
|
compositor when a second backend arrives"). It claims the virtio-gpu PCI function, and:
|
||||||
|
|
||||||
|
- sets up the **virtqueues** (control + cursor) and the device's config space,
|
||||||
|
- creates a **2D scanout resource** backed by an `shm` region, `attach_backing`s it,
|
||||||
|
`set_scanout`s it to a CRTC, and `resource_flush`es damaged rectangles,
|
||||||
|
- reads **EDID** (the `GET_EDID` control command) for the mode list, and `set_scanout`
|
||||||
|
at a chosen mode for **runtime mode-setting**,
|
||||||
|
- registers a `scanout` service and announces to the display service.
|
||||||
|
|
||||||
|
Its `resource_flush` is the real **present** — and gives a genuine **vsync/tear-free**
|
||||||
|
path a dumb GOP framebuffer can't.
|
||||||
|
|
||||||
|
## What v2 unlocks — and its honest scope
|
||||||
|
|
||||||
|
Behind the abstraction, a native backend gives runtime **mode-setting** (resolution /
|
||||||
|
refresh / bpp), **EDID** enumeration, and **vsync**. But only on devices we have a driver
|
||||||
|
for — realistically **VMs** (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs
|
||||||
|
need per-vendor KMS-class drivers that aren't getting written, so they **stay on GOP** —
|
||||||
|
which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold:
|
||||||
|
the **pluggable architecture** (a driver slots in when one exists) and a **rich, vsync'd
|
||||||
|
path in VMs**, where danos development happens. The framebuffer floor never goes away.
|
||||||
|
|
||||||
|
## Locked decisions
|
||||||
|
|
||||||
|
- **First native backend: virtio-gpu** — the VM standard; gives mode-set + a real
|
||||||
|
present/flush (and vsync), and exercises the whole pluggable design. Tested with QEMU
|
||||||
|
`-device virtio-gpu`.
|
||||||
|
- **Dynamic hot-attach** — boot on GOP, upgrade to native on the driver's announce,
|
||||||
|
re-attach across driver restarts; GOP is the floor for "no driver ever," not a live
|
||||||
|
fall-back after a reprogram.
|
||||||
|
- **Detection = push** (the driver announces to `.display`), not compositor polling.
|
||||||
|
- **v2 builds the `shm` capability** (endpoints → memory objects), shared with the future
|
||||||
|
client-surface path.
|
||||||
|
|
||||||
|
## See also
|
||||||
|
|
||||||
|
- [display.md](display.md) — v1: the compositor, the GOP-vs-device split, the WC discipline.
|
||||||
|
- [display-v2-plan.md](display-v2-plan.md) — the ordered build-out.
|
||||||
|
- [driver-model.md](driver-model.md) — claim / `mmio_map` / MSI / capability passing (M13).
|
||||||
|
- [resilience.md](resilience.md) — the restart machinery the hot-attach leans on.
|
||||||
@@ -2,12 +2,15 @@
|
|||||||
//! into the linear framebuffer the bootloader handed us. No firmware, no driver
|
//! into the linear framebuffer the bootloader handed us. No firmware, no driver
|
||||||
//! — just pixels.
|
//! — just pixels.
|
||||||
//!
|
//!
|
||||||
//! This is a **bootstrap** console — a stop-gap so early boot has something on
|
//! This is a **bootstrap / fatal-fallback** console. The driver machinery now exists — the
|
||||||
//! screen. The framebuffer is a general graphics surface, *not* inherently a text
|
//! user-space **display service** ([../services/display](../services/display/display.zig),
|
||||||
//! terminal; once the driver machinery exists it becomes a proper graphics device
|
//! docs/display.md) owns the framebuffer in normal operation — so this no longer paints
|
||||||
//! driver and this text-grid crutch goes away. It is therefore kept **separate
|
//! routine status. It exists for the two cases the display service can't cover: **early
|
||||||
//! from the diagnostic [log](log.zig)** — the log fans out to serial/debugcon/file,
|
//! boot**, before the service has claimed the framebuffer, and **fatal errors** (a kernel
|
||||||
//! while this only paints the handful of user-facing status lines and panics.
|
//! panic or a kernel-mode fault), which force it back on (`setSuppressed`) so a dying
|
||||||
|
//! machine's last words reach the screen even over a live display. It is kept **separate
|
||||||
|
//! from the diagnostic [log](log.zig)** — the log fans out to serial/debugcon/file and
|
||||||
|
//! carries all routine kernel output; this only paints those fatal cases.
|
||||||
//!
|
//!
|
||||||
//! The module owns a single console and a `present` flag; `write` is a no-op when
|
//! The module owns a single console and a `present` flag; `write` is a no-op when
|
||||||
//! the firmware handed over no framebuffer (a headless machine), so the kernel
|
//! the firmware handed over no framebuffer (a headless machine), so the kernel
|
||||||
|
|||||||
+46
-21
@@ -9,6 +9,7 @@ const wall_clock = @import("wall-clock.zig");
|
|||||||
const pmm = @import("pmm.zig");
|
const pmm = @import("pmm.zig");
|
||||||
const heap = @import("heap.zig");
|
const heap = @import("heap.zig");
|
||||||
const scheduler = @import("scheduler.zig");
|
const scheduler = @import("scheduler.zig");
|
||||||
|
const sync = @import("sync.zig");
|
||||||
const process = @import("process.zig");
|
const process = @import("process.zig");
|
||||||
const devices_broker = @import("devices-broker.zig");
|
const devices_broker = @import("devices-broker.zig");
|
||||||
const irq = @import("irq.zig");
|
const irq = @import("irq.zig");
|
||||||
@@ -156,12 +157,13 @@ fn kmain(boot_information: *const BootInformation) noreturn {
|
|||||||
log.print(" page tables: root = 0x{x:0>16}\n", .{architecture.activePageTable()});
|
log.print(" page tables: root = 0x{x:0>16}\n", .{architecture.activePageTable()});
|
||||||
log.print(" kernel segs: {d} (mapped with W^X permissions)\n", .{boot_information.kernel_segment_count});
|
log.print(" kernel segs: {d} (mapped with W^X permissions)\n", .{boot_information.kernel_segment_count});
|
||||||
|
|
||||||
// Now on our own tables, the framebuffer window is write-combining: bring up
|
// Now on our own tables, the framebuffer window is write-combining: bring up the
|
||||||
// the on-screen console and clear it (a fast burst here, not the loader's
|
// on-screen console and clear it to a blank canvas (a fast burst here, not the loader's
|
||||||
// uncached crawl). From here `status` reaches the screen as well as the log.
|
// uncached crawl). Routine boot output goes only to the log; this console now exists for
|
||||||
|
// early-boot and fatal (`fatal`/panic) output, until the display service takes over.
|
||||||
console.init(fb);
|
console.init(fb);
|
||||||
log.write(if (console.present())
|
log.write(if (console.present())
|
||||||
"/system/kernel: framebuffer console online (bootstrap; graphics driver later)\n"
|
"/system/kernel: framebuffer ready (early-boot + fatal fallback; the display service drives it in normal operation)\n"
|
||||||
else
|
else
|
||||||
"/system/kernel: no framebuffer (headless) -> logging to serial/debugcon only\n");
|
"/system/kernel: no framebuffer (headless) -> logging to serial/debugcon only\n");
|
||||||
|
|
||||||
@@ -414,11 +416,22 @@ fn bringUpSecondaries() void {
|
|||||||
log.print("/system/kernel: {d}/{d} cores online\n", .{ scheduler.onlineCount(), cores.len });
|
log.print("/system/kernel: {d}/{d} cores online\n", .{ scheduler.onlineCount(), cores.len });
|
||||||
}
|
}
|
||||||
|
|
||||||
/// A user-facing status line: to the diagnostic `log` *and* the on-screen console
|
/// A user-facing status line. Now that the user-space **display service** owns the
|
||||||
/// (if a framebuffer is present). The verbose log uses `log.*` directly and never
|
/// framebuffer in normal operation (docs/display.md), routine kernel output goes to the
|
||||||
/// touches the framebuffer.
|
/// diagnostic `log` (serial/debugcon/RAM) *only* — never to the on-screen console, which
|
||||||
|
/// the compositor is about to paint over. For a message that must reach the screen even so
|
||||||
|
/// — a panic or a fatal fault, when the machine is going down — use `fatal`.
|
||||||
fn status(message: []const u8) void {
|
fn status(message: []const u8) void {
|
||||||
log.write(message);
|
log.write(message);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A fatal, user-facing message: to the diagnostic log *and* the on-screen console, forcing
|
||||||
|
/// the console back on (`setSuppressed(false)`) first — a dying machine's last words outrank
|
||||||
|
/// any display service holding the framebuffer. The console is otherwise silent in normal
|
||||||
|
/// operation (see `status`); it exists now only for early-boot and fatal output.
|
||||||
|
fn fatal(message: []const u8) void {
|
||||||
|
log.write(message);
|
||||||
|
console.setSuppressed(false);
|
||||||
console.write(message);
|
console.write(message);
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -427,6 +440,11 @@ fn statusPrint(comptime fmt: []const u8, args: anytype) void {
|
|||||||
status(std.fmt.bufPrint(&buffer, fmt, args) catch return);
|
status(std.fmt.bufPrint(&buffer, fmt, args) catch return);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn fatalPrint(comptime fmt: []const u8, args: anytype) void {
|
||||||
|
var buffer: [256]u8 = undefined;
|
||||||
|
fatal(std.fmt.bufPrint(&buffer, fmt, args) catch return);
|
||||||
|
}
|
||||||
|
|
||||||
/// Frames (4 KiB pages) to whole MiB.
|
/// Frames (4 KiB pages) to whole MiB.
|
||||||
fn mib(pages: u64) u64 {
|
fn mib(pages: u64) u64 {
|
||||||
return pages * abi.page_size / (1024 * 1024);
|
return pages * abi.page_size / (1024 * 1024);
|
||||||
@@ -486,20 +504,26 @@ fn onException(state: *const architecture.CpuState) noreturn {
|
|||||||
}
|
}
|
||||||
|
|
||||||
log.checkpoint(cp_exception);
|
log.checkpoint(cp_exception);
|
||||||
// The machine is going down: force the console back on even if a display service
|
|
||||||
// was holding the framebuffer, so the exception actually reaches the screen.
|
|
||||||
console.setSuppressed(false);
|
|
||||||
const core = scheduler.currentCpuIndex();
|
const core = scheduler.currentCpuIndex();
|
||||||
// A fault is user-facing enough to paint on screen too (via statusPrint), on
|
// The machine is going down: paint the exception on screen too — `fatalPrint` forces the
|
||||||
// top of the diagnostic log.
|
// console back on even if a display service was holding the framebuffer — on top of the
|
||||||
statusPrint("\nCPU EXCEPTION on core {d}: {s} (vector {d})\n", .{ core, architecture.exceptionName(state.vector), state.vector });
|
// diagnostic log.
|
||||||
statusPrint(" error code : 0x{x}\n", .{state.error_code});
|
fatalPrint("\nCPU EXCEPTION on core {d}: {s} (vector {d})\n", .{ core, architecture.exceptionName(state.vector), state.vector });
|
||||||
statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
|
// Name the culprit: which task, and whether it faulted in ring 3 (a process the
|
||||||
statusPrint(" SP : 0x{x:0>16}\n", .{architecture.stackPointer(state)});
|
// kernel would normally kill — landing here means it had no address space) or ring 0
|
||||||
if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address});
|
// (the trusted base itself). Without this the fatal report is anonymous.
|
||||||
|
fatalPrint(" task : {d} ({s}), {s}\n", .{ scheduler.currentIdSafe(), scheduler.currentNameSafe(), if (architecture.fromUser(state)) "ring 3 (user)" else "ring 0 (kernel)" });
|
||||||
|
fatalPrint(" error code : 0x{x}\n", .{state.error_code});
|
||||||
|
fatalPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
|
||||||
|
fatalPrint(" SP : 0x{x:0>16}\n", .{architecture.stackPointer(state)});
|
||||||
|
if (architecture.faultAddress(state)) |address| fatalPrint(" fault addr : 0x{x:0>16}\n", .{address});
|
||||||
|
|
||||||
var buffer: [128]u8 = undefined;
|
var buffer: [128]u8 = undefined;
|
||||||
log.recordPanic(std.fmt.bufPrint(&buffer, "CPU exception {s} (vector {d}) on core {d} at IP 0x{x}", .{ architecture.exceptionName(state.vector), state.vector, core, architecture.instructionPointer(state) }) catch "cpu exception");
|
log.recordPanic(std.fmt.bufPrint(&buffer, "CPU exception {s} (vector {d}) on core {d} at IP 0x{x}", .{ architecture.exceptionName(state.vector), state.vector, core, architecture.instructionPointer(state) }) catch "cpu exception");
|
||||||
|
// Free the BKL if this core held it (a kernel-mode fault, or a nested fault in the
|
||||||
|
// recovery teardown), so halting this one core doesn't deadlock every other core on
|
||||||
|
// the lock. Only that core stops; the rest — and the supervisor — keep running.
|
||||||
|
sync.releaseIfHeldHere();
|
||||||
architecture.halt();
|
architecture.halt();
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -511,10 +535,11 @@ pub const panic = std.debug.FullPanic(struct {
|
|||||||
_ = first_trace_address;
|
_ = first_trace_address;
|
||||||
log.checkpoint(cp_panic);
|
log.checkpoint(cp_panic);
|
||||||
log.recordPanic(message);
|
log.recordPanic(message);
|
||||||
console.setSuppressed(false); // a panic outranks any display service holding the screen
|
fatal("\nKERNEL PANIC: "); // a panic outranks any display service holding the screen
|
||||||
status("\nKERNEL PANIC: ");
|
fatal(message);
|
||||||
status(message);
|
fatal("\n");
|
||||||
status("\n");
|
fatalPrint(" task : {d} ({s})\n", .{ scheduler.currentIdSafe(), scheduler.currentNameSafe() });
|
||||||
|
sync.releaseIfHeldHere(); // don't deadlock the other cores on the lock we may hold
|
||||||
architecture.halt();
|
architecture.halt();
|
||||||
}
|
}
|
||||||
}.panic);
|
}.panic);
|
||||||
|
|||||||
@@ -791,6 +791,21 @@ pub fn currentCpuIndex() u32 {
|
|||||||
return thisCpu().index;
|
return thisCpu().index;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The running task's id, or 0 if this core's scheduler isn't up yet (early boot, no GS
|
||||||
|
/// base). Safe for a fault reporter to call unconditionally — like `currentCpuIndex`,
|
||||||
|
/// it never dereferences an unpublished per-CPU pointer and so can't fault a second time.
|
||||||
|
pub fn currentIdSafe() u32 {
|
||||||
|
if (architecture.cpuLocal() == 0) return 0;
|
||||||
|
return thisCpu().current.id;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The running task's name (argv[0]), or "" if this core's scheduler isn't up yet.
|
||||||
|
/// The companion to `currentIdSafe` for naming the culprit in a fatal fault report.
|
||||||
|
pub fn currentNameSafe() []const u8 {
|
||||||
|
if (architecture.cpuLocal() == 0) return "";
|
||||||
|
return thisCpu().current.name();
|
||||||
|
}
|
||||||
|
|
||||||
/// Change the running task's priority (takes effect next time it's enqueued).
|
/// Change the running task's priority (takes effect next time it's enqueued).
|
||||||
pub fn setPriority(p: Priority) void {
|
pub fn setPriority(p: Priority) void {
|
||||||
current().priority = p;
|
current().priority = p;
|
||||||
|
|||||||
@@ -33,6 +33,13 @@ const architecture = @import("architecture");
|
|||||||
/// 0 = free, 1 = held. A single global lock for the whole kernel.
|
/// 0 = free, 1 = held. A single global lock for the whole kernel.
|
||||||
var held = std.atomic.Value(u32).init(0);
|
var held = std.atomic.Value(u32).init(0);
|
||||||
|
|
||||||
|
/// The per-CPU base pointer (`architecture.cpuLocal()`) of the core currently holding
|
||||||
|
/// the lock, or 0 when free. Metadata only — `held` is what enforces exclusion — read
|
||||||
|
/// solely by `releaseIfHeldHere` on the fatal-fault path. `cpuLocal()` is a unique,
|
||||||
|
/// architecture-level token per core (0 before this core's GS base is published, which
|
||||||
|
/// is fine: that window is single-core early boot, where no other core can deadlock).
|
||||||
|
var owner = std.atomic.Value(usize).init(0);
|
||||||
|
|
||||||
/// Enter the kernel: disable interrupts on this core, then spin until we own the
|
/// Enter the kernel: disable interrupts on this core, then spin until we own the
|
||||||
/// lock. Returns the caller's prior interrupt flags for `leave` to restore.
|
/// lock. Returns the caller's prior interrupt flags for `leave` to restore.
|
||||||
/// Interrupts stay off for the whole critical section so this core's timer tick
|
/// Interrupts stay off for the whole critical section so this core's timer tick
|
||||||
@@ -67,14 +74,29 @@ export fn releaseForFreshTask() callconv(.c) void {
|
|||||||
release();
|
release();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Release the big kernel lock **only if this core is the one holding it** — a no-op
|
||||||
|
/// otherwise. For the fatal-fault path (a kernel-mode fault, or a nested fault inside the
|
||||||
|
/// recovery teardown, both of which run under the lock): a core that dies holding the BKL
|
||||||
|
/// must free it, or every other core spins forever in `acquire` and the whole machine
|
||||||
|
/// deadlocks instead of just that core stopping. It must NOT free a lock another core
|
||||||
|
/// owns, hence the owner check. Caveat: if we held it mid-mutation the shared state may be
|
||||||
|
/// inconsistent — but letting the other cores (and the supervisor) run on possibly-degraded
|
||||||
|
/// state is strictly more recoverable than a guaranteed total hang.
|
||||||
|
pub fn releaseIfHeldHere() void {
|
||||||
|
const me = architecture.cpuLocal();
|
||||||
|
if (me != 0 and owner.load(.monotonic) == me) release();
|
||||||
|
}
|
||||||
|
|
||||||
fn acquire() void {
|
fn acquire() void {
|
||||||
// Test-and-test-and-set: try once, then spin read-only until the lock looks
|
// Test-and-test-and-set: try once, then spin read-only until the lock looks
|
||||||
// free before retrying the (bus-locked) swap — cheaper on the coherency fabric.
|
// free before retrying the (bus-locked) swap — cheaper on the coherency fabric.
|
||||||
while (held.swap(1, .acquire) != 0) {
|
while (held.swap(1, .acquire) != 0) {
|
||||||
while (held.load(.monotonic) != 0) architecture.cpuRelax();
|
while (held.load(.monotonic) != 0) architecture.cpuRelax();
|
||||||
}
|
}
|
||||||
|
owner.store(architecture.cpuLocal(), .monotonic);
|
||||||
}
|
}
|
||||||
|
|
||||||
fn release() void {
|
fn release() void {
|
||||||
|
owner.store(0, .monotonic);
|
||||||
held.store(0, .release);
|
held.store(0, .release);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -30,12 +30,22 @@ const log_path = "/mnt/usb/DANOS.LOG";
|
|||||||
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent
|
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent
|
||||||
/// on purpose: the device manager owns those. (A future init reads this from a
|
/// on purpose: the device manager owns those. (A future init reads this from a
|
||||||
/// manifest under /system/services instead of a hardcoded list.)
|
/// manifest under /system/services instead of a hardcoded list.)
|
||||||
const boot_services = [_][]const u8{ "vfs", "input", "device-manager", "fat", "display" };
|
const boot_services = [_][]const u8{ "vfs", "input", "device-manager", "fat", "display", "display-demo" };
|
||||||
|
|
||||||
var children: [boot_services.len]u32 = .{0} ** boot_services.len;
|
/// The live process id of each boot service (0 = not running), indexed by its position
|
||||||
var child_count: usize = 0;
|
/// in `boot_services`, plus how many times init has restarted it. init supervises these:
|
||||||
|
/// it spawns them against `supervision_endpoint` and, on a child's death, restarts it (up
|
||||||
|
/// to `maximum_restarts`) — the reincarnation half of resilience (docs/resilience.md), the
|
||||||
|
/// service-level counterpart to the device manager's driver restarts.
|
||||||
|
var child_ids: [boot_services.len]u32 = .{0} ** boot_services.len;
|
||||||
|
var restart_counts: [boot_services.len]u32 = .{0} ** boot_services.len;
|
||||||
|
var shutting_down = false;
|
||||||
var supervision_endpoint: runtime.ipc.Handle = 0;
|
var supervision_endpoint: runtime.ipc.Handle = 0;
|
||||||
|
|
||||||
|
/// Give up restarting a service after this many crashes — a crash-loop cap, so a service
|
||||||
|
/// that faults immediately on every spawn doesn't respawn forever.
|
||||||
|
const maximum_restarts = 3;
|
||||||
|
|
||||||
pub fn main() void {
|
pub fn main() void {
|
||||||
// Prove the heap end to end: allocate through the runtime allocator (which
|
// Prove the heap end to end: allocate through the runtime allocator (which
|
||||||
// mmaps pages from the kernel and carves them with the free list), write into
|
// mmaps pages from the kernel and carves them with the free list), write into
|
||||||
@@ -63,11 +73,8 @@ pub fn main() void {
|
|||||||
// Bring up the boot services, supervised so init can stop them cleanly.
|
// Bring up the boot services, supervised so init can stop them cleanly.
|
||||||
// Best-effort and silent: each service announces its own readiness, and in
|
// Best-effort and silent: each service announces its own readiness, and in
|
||||||
// an isolation test with no initial-ramdisk the spawns simply no-op.
|
// an isolation test with no initial-ramdisk the spawns simply no-op.
|
||||||
for (boot_services) |service| {
|
for (boot_services, 0..) |service, i| {
|
||||||
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| {
|
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| child_ids[i] = id;
|
||||||
children[child_count] = id;
|
|
||||||
child_count += 1;
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// Once the storage stack is up, a one-shot copies the boot log to the USB
|
// Once the storage stack is up, a one-shot copies the boot log to the USB
|
||||||
@@ -105,11 +112,47 @@ pub fn main() void {
|
|||||||
if (receive[1] == @intFromEnum(power.Event.power_button)) shutDown();
|
if (receive[1] == @intFromEnum(power.Event.power_button)) shutDown();
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
// Child-exit notifications and anything else: keep waiting.
|
if (got.isChildExit()) {
|
||||||
|
restartChild(got.childProcessId());
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
// Anything else: keep waiting.
|
||||||
if (got.isNotification()) continue;
|
if (got.isNotification()) continue;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// A supervised boot service died. Find which one and restart it — unless it exited
|
||||||
|
/// cleanly (it chose to stop, e.g. a driver with no hardware) or has hit the crash-loop
|
||||||
|
/// cap. Reclaiming the dead process is already the kernel's job (docs/process-lifecycle.md
|
||||||
|
/// iron rule 1); init only decides whether to bring it back.
|
||||||
|
fn restartChild(id: u32) void {
|
||||||
|
if (shutting_down) return; // deaths during the stop sequence are expected, not crashes
|
||||||
|
for (boot_services, 0..) |service, i| {
|
||||||
|
if (child_ids[i] != id) continue;
|
||||||
|
child_ids[i] = 0;
|
||||||
|
// An unknown reason (the record aged out) is treated as a crash worth restarting.
|
||||||
|
const reason = runtime.process.exitReason(id) orelse .fault;
|
||||||
|
if (reason == .exited) {
|
||||||
|
logLine("/system/services/init: {s} exited cleanly; not restarting\n", .{service});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
restart_counts[i] += 1;
|
||||||
|
if (restart_counts[i] > maximum_restarts) {
|
||||||
|
logLine("/system/services/init: {s} keeps crashing; giving up after {d} restarts\n", .{ service, maximum_restarts });
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
logLine("/system/services/init: {s} died ({s}); restarting ({d}/{d})\n", .{ service, @tagName(reason), restart_counts[i], maximum_restarts });
|
||||||
|
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |new_id| child_ids[i] = new_id;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
// An untracked child (e.g. the log-flush one-shot): nothing to restart.
|
||||||
|
}
|
||||||
|
|
||||||
|
fn logLine(comptime fmt: []const u8, args: anytype) void {
|
||||||
|
var line: [128]u8 = undefined;
|
||||||
|
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, args) catch return);
|
||||||
|
}
|
||||||
|
|
||||||
/// Look up the power service and subscribe our endpoint (handed over as the
|
/// Look up the power service and subscribe our endpoint (handed over as the
|
||||||
/// call's capability) so events arrive as buffered messages here.
|
/// call's capability) so events arrive as buffered messages here.
|
||||||
fn subscribePower() void {
|
fn subscribePower() void {
|
||||||
@@ -153,15 +196,16 @@ fn flushKernelLog() void {
|
|||||||
/// it), waiting up to a deadline for each to exit before killing it, then ask the
|
/// it), waiting up to a deadline for each to exit before killing it, then ask the
|
||||||
/// power service to enter S5.
|
/// power service to enter S5.
|
||||||
fn shutDown() void {
|
fn shutDown() void {
|
||||||
|
shutting_down = true; // the stop loop below kills children — those deaths aren't crashes
|
||||||
_ = runtime.system.write("/system/services/init: shutting down\n");
|
_ = runtime.system.write("/system/services/init: shutting down\n");
|
||||||
// Persist the fullest log to the USB volume BEFORE tearing anything down: the
|
// Persist the fullest log to the USB volume BEFORE tearing anything down: the
|
||||||
// reverse-order stop loop below kills the fat server (children[3]) first, so
|
// reverse-order stop loop below kills the fat server first, so /mnt/usb must be
|
||||||
// /mnt/usb must be written while it is still mounted.
|
// written while it is still mounted.
|
||||||
flushKernelLog();
|
flushKernelLog();
|
||||||
var i = child_count;
|
var i = boot_services.len;
|
||||||
while (i > 0) {
|
while (i > 0) {
|
||||||
i -= 1;
|
i -= 1;
|
||||||
if (children[i] != 0) runtime.process.stop(children[i], 2000, supervision_endpoint);
|
if (child_ids[i] != 0) runtime.process.stop(child_ids[i], 2000, supervision_endpoint);
|
||||||
}
|
}
|
||||||
if (runtime.ipc.lookup(.power)) |h| {
|
if (runtime.ipc.lookup(.power)) |h| {
|
||||||
const request = power.Shutdown{};
|
const request = power.Shutdown{};
|
||||||
|
|||||||
Reference in New Issue
Block a user