diff --git a/docs/README.md b/docs/README.md index 77cde6c..2b16d96 100644 --- a/docs/README.md +++ b/docs/README.md @@ -42,9 +42,14 @@ rather than restate it. Roughly in the order things happen at runtime: Start with the north star: -- **[vision.md](vision.md) — the vision.** danos is aiming to be a real-time - microkernel: minimal kernel, drivers/services isolated in user space, preemptive - scheduling with timing guarantees. The *why* that shapes everything below. +- **[vision.md](vision.md) — the vision.** danos is a **learning-by-doing** + microkernel: minimal kernel, drivers/services isolated in user space, chosen for + **resilience** (restartable components). Win condition: runs on the author's PC and + both Raspberry Pis, ideally with a GUI. Real-time is an option to explore, not a + requirement. The *why* that shapes everything below. +- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on + fault isolation + live restart — the reincarnation-server + capability model that + makes "if I break it, I can restart it" real. danos's core motivation. Cutting across all of these: diff --git a/docs/resilience.md b/docs/resilience.md new file mode 100644 index 0000000..22878a8 --- /dev/null +++ b/docs/resilience.md @@ -0,0 +1,152 @@ +# Resilience: fault isolation and live restart + +A design/research note, not built yet. This is the property danos is really chasing: +**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.** +A crashed driver gets restarted; a wedged service gets killed and brought back. It's +the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal +from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) — +one that's less pervasive to build (see [vision.md](vision.md)). + +## The idea: "let it crash" + supervision + +The philosophy is older than microkernels and shows up across systems: don't try to +make every component perfect — make failures **contained and recoverable**. Isolate +each component, watch it, and when it dies, restart it from a known-good state. A +small trusted core supervises a fleet of restartable, untrusted parts. + +Prior art worth studying (see Further reading): **MINIX 3's reincarnation server** (a +driver crashes, a supervisor restarts it live — the closest thing to your goal), +**QNX** (restartable drivers on a message-passing microkernel), **Erlang/OTP +supervision trees** ("let it crash", not a kernel but the canonical design), and the +historical **Tandem NonStop** (fault-tolerant by process pairs). + +## Why a microkernel makes this possible + +The blast radius of a fault is the address space it happens in. In a monolith, a +driver bug can corrupt anything — the kernel *is* the driver. In a microkernel, +drivers and services are **isolated user-space processes**, so a fault is trapped by +the kernel and confined to that one process. The kernel — the one thing that *can't* +be restarted, because it's the trusted base — stays tiny, which is precisely why a +small kernel is a *more recoverable* kernel: less code that can take the whole system +down. **Keeping the kernel minimal is a resilience strategy, not just an aesthetic.** + +## The building blocks + +1. **Address-space isolation.** A fault in one component can't corrupt another or the + kernel. This is the [user-mode milestone](vision.md) (ring 3, per-process page + tables) — the shared prerequisite for *any* of this, and it's needed regardless. +2. **Fault detection** — how the system notices a component is dead or sick: + - **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps + to the kernel, which kills the process and notifies the supervisor. The clean, + easy case — and danos already reports CPU faults (see + [interrupts.md](interrupts.md)); user mode turns "halt on fault" into "kill the + process and tell the supervisor." + - **Hang**: a livelocked or infinite-looping component needs a **watchdog / + heartbeat** and the ability to **preempt and kill** it. The preemptive scheduler + already built ([scheduling.md](scheduling.md)) is what makes a runaway component + killable — a nice case of a scheduling mechanism serving resilience without any + real-time *guarantee*. + - **Misbehaviour**: IPC timeouts, failed health checks. +3. **A supervisor / reincarnation server.** A user-space server holding the *policy*: + what components exist, their dependencies, and each one's restart strategy. When a + component dies, it decides whether/how to restart it. (MINIX 3 calls this the + reincarnation server; Erlang calls it a supervisor.) +4. **A resource model that supports clean teardown.** When a component dies, its + resources — memory, MMIO grants, IPC channels, IRQ routes — must be **reclaimed**, + and a restarted replacement must be able to **re-acquire** them. This is where a + **capability** model shines (seL4's reference design): a component holds + capabilities to its resources; killing it **revokes** them, which frees everything + in one clean sweep, and the supervisor hands the replacement fresh caps. A simpler + grant/ownership table can work too — capabilities are the principled version. +5. **Re-initialisable drivers.** A driver must start from a known state and + re-establish its hardware. Some hardware is easy to reset; some holds state that's + hard to recover — a real limit on what "just restart it" can fix. + +## The hard part: restarting *correctly* + +Detecting and killing is the easy half. The genuinely tricky questions are about the +*rest of the system* when a component dies: + +- **In-flight IPC**: messages sent to the dead component, or replies its clients are + blocked waiting for. The channel has to break cleanly and unblock the waiters with + an error rather than hang them forever (a design constraint that reaches back into + [ipc.md](ipc.md) — channels need a "peer died" outcome). +- **Clients**: how does a client discover the service it was talking to is gone and + has been replaced? Options: capability revocation makes stale handles fail; or a + **name server** re-binds clients to the new instance; or clients retry through a + stable endpoint. +- **State**: the cheapest model is **stateless restart** — the replacement starts + fresh and clients re-establish whatever they need. Richer options (checkpointed + state, state handed to a standby) are more work and more failure modes. Start + stateless. + +These are the constraints most worth *bumping into and researching* — they're where +resilience gets genuinely interesting. + +## Kernel mechanism vs user-space policy + +The microkernel split applies to fault management itself: + +- **Kernel (mechanism):** isolation, trapping faults, enforcing capabilities/grants, + IPC, creating/destroying address spaces, granting/revoking resources, preempting a + runaway task. +- **User space (policy):** the supervisor decides *what* to restart, *when*, and + *how* — dependency order, retry limits, escalation. None of that belongs in the + kernel. + +So the kernel gains a few primitives (kill an address space, reclaim its resources, +deliver a "child died" notification); everything smart lives in a user-space server. + +## What's *not* recoverable this way + +Honest boundaries: + +- **The kernel itself.** It's the trusted base; if it faults, this mechanism can't + save it. The mitigation is to keep it tiny — the microkernel bet. +- **Corrupted hardware state.** Isolation limits the blast radius to one process, but + if a driver wedged the device itself, a restart may not un-wedge it. +- **Shared-resource corruption** that happened *before* the fault was detected. Clean + capability revocation limits this, but it's why fault *detection latency* matters. + +## Suggested ordering + +1. **User mode + address-space isolation** — the shared prerequisite (also on the + path for everything else). +2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into + "confine to the process and report it." +3. **A minimal supervisor server** that can (re)start a process. +4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant + table. +5. **First restartable driver** — the keyboard — as the end-to-end proof: crash it on + purpose, watch it come back. + +## Relationship to real-time + +Resilience needs **structural** features (isolation + supervision + a resource +model); real-time needs a **pervasive** timing invariant. They're separable, and +resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](vision.md)). +Note the overlap, though: **preemptive scheduling** and **priorities** — already +built — serve resilience too (you can preempt and kill a misbehaving component, and +run the supervisor at high priority). So danos keeps the useful *mechanisms* of the +real-time work without owing anyone a timing *guarantee*. + +## Further reading + +- Herder, Bos, Gras, Homburg, Tanenbaum — the **MINIX 3** papers, esp. *"Construction + of a Highly Dependable Operating System"* and *"Fault Isolation for Device + Drivers"* — the reincarnation server, the closest match to danos's goal. +- **QNX** architecture — a shipping microkernel with restartable drivers. +- **Erlang/OTP** supervision trees and the *"let it crash"* philosophy — the design + pattern, distilled. +- **seL4** capability model — the principled basis for clean resource teardown. +- **Tandem NonStop** (historical) — fault tolerance via process pairs. + +## Related + +- [vision.md](vision.md) — the goals this serves (learning by doing; resilience over + hard real-time). +- [scheduling.md](scheduling.md) — preemption, which makes runaway components killable. +- [ipc.md](ipc.md) — channels that need a "peer died" outcome for clean restart. +- [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and + restart" instead of "halt". +- [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context. diff --git a/docs/vision.md b/docs/vision.md index 05b9798..cc184b4 100644 --- a/docs/vision.md +++ b/docs/vision.md @@ -1,84 +1,108 @@ -# Vision: a real-time microkernel +# Vision: a microkernel, built to learn -danos is aiming to be a **real-time operating system built on a microkernel** — -where drivers and services run isolated in user space for maximum stability, and -scheduling gives real guarantees about timing. This page is the north star: the -*why* that shapes every design decision below it. Read it before adding anything -structural. +danos exists first and foremost as a **learning-by-doing project**: the point is to +build a real operating system, bump into the hard constraints for real, and research +them from a position of having actually hit them. The docs in this folder are part of +that — they're where a constraint gets understood once it's been met. -## Microkernel +That framing sets the priorities. danos is not chasing a spec or a product; it's +chasing understanding, with a concrete, motivating **win condition** to aim at. -The kernel stays **minimal** — only what genuinely must run in privileged mode: +## The win condition + +danos is a "win" when it: + +- **boots and runs on real hardware** — the author's **PC** (x86-64) and **both + Raspberry Pis**: the **Zero 2 W** and the **Pi 5** (both `aarch64`, one backend — + see [arm.md](arm.md)), +- **has a graphical user interface**, ideally — building on the framebuffer it + already draws to. + +Everything below serves that, or serves the curiosity that the project runs on. + +## Why a microkernel: resilience + +The kernel stays **minimal** — only what genuinely must run privileged: - scheduling, - inter-process communication (IPC), - memory management (address spaces, page tables), - low-level interrupt dispatch. -Everything else — device drivers, filesystems, the network stack — runs as an -**isolated user-space server**, each in its own address space with only the +Everything else — device drivers, filesystems, the GUI, the network stack — runs as +an **isolated user-space server**, each in its own address space with only the privileges it needs. -The payoff is **stability through isolation**. A driver bug can't corrupt the -kernel or another driver; a crashing service is contained and can be restarted, -while the rest of the system keeps running. That's the opposite of a monolithic -kernel, where a single driver fault can take everything down. +The reason for this shape is **resilience**: the ability to **re-initialise parts of +the OS while it runs**. A driver bug can't corrupt the kernel or another driver; a +crashed or wedged component is contained, killed, and **restarted** — "if I break +something, I can just fix it," without rebooting. Keeping the kernel tiny is part of +that strategy: the one thing that *can't* be restarted is the trusted base, so the +less code in it, the less that can take the whole system down. This is the project's +real motivation, and it has its own design note: [resilience.md](resilience.md). -The cost is that **IPC becomes the backbone**: whatever used to be a function call -across a monolithic kernel is now a message between address spaces. In a -microkernel, IPC performance essentially *is* system performance (the lesson of -L4). So IPC must be fast, and it's a first-class concern, not an afterthought. -Hardware interrupts, too, become IPC: the kernel turns an IRQ into a message to the -driver task that owns that device. +The cost is that **IPC becomes the backbone**: what used to be a function call inside +a monolithic kernel is now a message between address spaces. In a microkernel, IPC +performance essentially *is* system performance (the lesson of L4), so it's a +first-class concern. Hardware interrupts become IPC too: the kernel turns an IRQ into +a message to the driver that owns the device. -## Real-time +## On real-time: an option, not a commitment -danos schedules **preemptively, with guarantees about quanta** — the system must -be able to promise that a task runs when it's supposed to, within bounded time. -That imposes concrete requirements: +danos was originally framed as a hard **real-time** OS. That's now held as **one +interesting constraint to explore, not a requirement** — because real-time is a +*pervasive* invariant (every operation must be provably time-bounded, everywhere) +that would slow every milestone, whereas resilience is a set of *structural* features +that's lighter to build and is what the project actually wants. The trade-off is +written up in [smp.md](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience). -- **Fixed-priority preemptive scheduling.** The highest-priority ready task always - runs; a higher-priority task that becomes ready preempts a lower one immediately. - Not round-robin (which is fair but not predictable). -- **A calibrated, deterministic clock.** Guarantees measured in "quanta" are - meaningless on an arbitrary tick rate — real time requires a timer calibrated to - a known frequency. -- **Bounded interrupt latency.** Interrupt-disabled sections must be short and - bounded, so a ready high-priority task is never delayed by an unbounded kernel - operation. -- **Deterministic kernel operations.** Scheduling decisions should be O(1) (e.g. a - priority bitmap), not "walk a list of unknown length." -- **Priority inheritance** (once there are locks/IPC), so a high-priority task - blocked on a resource held by a low-priority one can't be delayed indefinitely by - a middle-priority task — bounding priority inversion. +What danos keeps from the real-time direction, because it's cheap and useful anyway: -A consequence worth stating early: the current [kernel heap](heap.md) is a -first-fit free list, which has **unbounded allocation time** and can fragment — it -is *not* real-time safe. It's fine for one-time kernel setup, but real-time paths -must pre-allocate or use a bounded (fixed-size pool) allocator. Don't allocate on a -hot real-time path. +- **Fixed-priority preemptive scheduling** — the highest-priority ready task runs, and + preemption lets a runaway component be interrupted and killed (which *serves + resilience*). Already built ([scheduling.md](scheduling.md)). +- **A calibrated, deterministic clock** — already built ([device-interrupts.md](device-interrupts.md)). -## What this means for the roadmap +What danos does *not* owe anyone unless it deliberately chooses real-time later: +timing *guarantees*, priority inheritance, bounded allocators, tickless timers, MCS +scheduling contexts. Concretely, the current [heap](heap.md) is a first-fit free list +with unbounded allocation time — fine here, and only a problem *if* a hard-real-time +path is ever added. Note that **QNX is both** a real-time and a restartable +microkernel, so choosing resilience now doesn't close the real-time door — it just +doesn't pay the tax yet. -The vision reorders the obvious hobby-kernel path. Notably, **drivers are not -built into the kernel** — so an in-kernel keyboard driver would be throwaway work. -Input devices arrive later, as the *first user-space drivers*, once the machinery -to isolate them exists. The trajectory: +## The roadmap — tracks, not a strict line -1. **Calibrated timer / clock** — a known-frequency, deterministic tick. The - foundation real-time quanta rest on. *(next)* -2. **Real-time scheduler** — fixed-priority preemptive, kernel threads first: - context switch, task struct, priority run-queue, timer-driven preemption. -3. **User mode + address-space isolation** — higher-half kernel, ring 3, per-process - page tables. The substrate for isolated servers. -4. **IPC** — fast message passing between address spaces. The microkernel's heart. -5. **User-space drivers** — interrupts delivered as IPC, plus MMIO/port-access - grants. The keyboard becomes the first one, validating the whole model. +Because the driver is curiosity plus the win condition, the roadmap is a set of +**tracks** with dependencies, not a rigid sequence. Pick by interest; mind the +prerequisites. -## Where we are +**Done:** UEFI boot, framebuffer + [serial](testing.md), [physical frames](frame-allocator.md) +(with boot-services memory reclaimed), [paging](paging.md) with W^X, [exceptions and +interrupts](interrupts.md), a [calibrated timer + ns clock](device-interrupts.md), a +[heap](heap.md), a [fixed-priority preemptive scheduler](scheduling.md) with blocking, +and in-kernel [IPC channels](ipc.md) — plus a [test harness](testing.md). -The foundation is in place: UEFI boot, framebuffer + [serial](testing.md), -[physical frames](frame-allocator.md), [paging](paging.md) with W^X, [exceptions -and interrupts](interrupts.md), a [timer](device-interrupts.md), and a -[heap](heap.md) — plus a [test harness](testing.md). The kernel boots and has its -core services; the next milestones make it *schedule*, then *isolate*. +- **Isolation track** — **user mode + address-space isolation** (higher-half kernel, + ring 3, per-process page tables). The substrate everything else needs. *Next, and a + prerequisite for the resilience and driver tracks.* +- **Resilience track** — fault → kill → notify, a supervisor/reincarnation server, + resource cleanup on death, then a restartable driver as proof. Needs isolation. + See [resilience.md](resilience.md). +- **ARM track** — the `aarch64` port so danos runs on the Zero 2 W and Pi 5. Largely + independent of the others (it's the [arch layer](arch.md)); directly serves the win + condition. Likely via aarch64-UEFI first (QEMU `virt` + AAVMF), then real boards. + See [arm.md](arm.md), and [discovery.md](discovery.md) for the device tree it needs. +- **GUI track** — a framebuffer-based windowing/compositor, and the input + display + drivers under it. Builds on the neutral framebuffer (so it's arch-independent), and + on the driver model from the isolation/resilience tracks. The visible payoff. + +The natural spine is **isolation → (resilience + drivers) → GUI**, with the **ARM +track** pursued alongside whenever the itch to see it boot on a Pi wins out. + +## How to use this page + +Read it before adding anything structural. When a design decision comes up, the +question is: does it serve the **win condition** (runs on the three machines, with a +GUI), or the **learning** (a constraint worth meeting)? If it serves neither — e.g. +paying the full real-time tax with no payoff in sight — it can wait.