documenting things to research

This commit is contained in:
2026-07-03 21:40:59 +01:00
parent a6de418e42
commit d055de7141
3 changed files with 247 additions and 66 deletions
+8 -3
View File
@@ -42,9 +42,14 @@ rather than restate it. Roughly in the order things happen at runtime:
Start with the north star:
- **[vision.md](vision.md) — the vision.** danos is aiming to be a real-time
microkernel: minimal kernel, drivers/services isolated in user space, preemptive
scheduling with timing guarantees. The *why* that shapes everything below.
- **[vision.md](vision.md) — the vision.** danos is a **learning-by-doing**
microkernel: minimal kernel, drivers/services isolated in user space, chosen for
**resilience** (restartable components). Win condition: runs on the author's PC and
both Raspberry Pis, ideally with a GUI. Real-time is an option to explore, not a
requirement. The *why* that shapes everything below.
- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on
fault isolation + live restart — the reincarnation-server + capability model that
makes "if I break it, I can restart it" real. danos's core motivation.
Cutting across all of these:
+152
View File
@@ -0,0 +1,152 @@
# Resilience: fault isolation and live restart
A design/research note, not built yet. This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) —
one that's less pervasive to build (see [vision.md](vision.md)).
## The idea: "let it crash" + supervision
The philosophy is older than microkernels and shows up across systems: don't try to
make every component perfect — make failures **contained and recoverable**. Isolate
each component, watch it, and when it dies, restart it from a known-good state. A
small trusted core supervises a fleet of restartable, untrusted parts.
Prior art worth studying (see Further reading): **MINIX 3's reincarnation server** (a
driver crashes, a supervisor restarts it live — the closest thing to your goal),
**QNX** (restartable drivers on a message-passing microkernel), **Erlang/OTP
supervision trees** ("let it crash", not a kernel but the canonical design), and the
historical **Tandem NonStop** (fault-tolerant by process pairs).
## Why a microkernel makes this possible
The blast radius of a fault is the address space it happens in. In a monolith, a
driver bug can corrupt anything — the kernel *is* the driver. In a microkernel,
drivers and services are **isolated user-space processes**, so a fault is trapped by
the kernel and confined to that one process. The kernel — the one thing that *can't*
be restarted, because it's the trusted base — stays tiny, which is precisely why a
small kernel is a *more recoverable* kernel: less code that can take the whole system
down. **Keeping the kernel minimal is a resilience strategy, not just an aesthetic.**
## The building blocks
1. **Address-space isolation.** A fault in one component can't corrupt another or the
kernel. This is the [user-mode milestone](vision.md) (ring 3, per-process page
tables) — the shared prerequisite for *any* of this, and it's needed regardless.
2. **Fault detection** — how the system notices a component is dead or sick:
- **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps
to the kernel, which kills the process and notifies the supervisor. The clean,
easy case — and danos already reports CPU faults (see
[interrupts.md](interrupts.md)); user mode turns "halt on fault" into "kill the
process and tell the supervisor."
- **Hang**: a livelocked or infinite-looping component needs a **watchdog /
heartbeat** and the ability to **preempt and kill** it. The preemptive scheduler
already built ([scheduling.md](scheduling.md)) is what makes a runaway component
killable — a nice case of a scheduling mechanism serving resilience without any
real-time *guarantee*.
- **Misbehaviour**: IPC timeouts, failed health checks.
3. **A supervisor / reincarnation server.** A user-space server holding the *policy*:
what components exist, their dependencies, and each one's restart strategy. When a
component dies, it decides whether/how to restart it. (MINIX 3 calls this the
reincarnation server; Erlang calls it a supervisor.)
4. **A resource model that supports clean teardown.** When a component dies, its
resources — memory, MMIO grants, IPC channels, IRQ routes — must be **reclaimed**,
and a restarted replacement must be able to **re-acquire** them. This is where a
**capability** model shines (seL4's reference design): a component holds
capabilities to its resources; killing it **revokes** them, which frees everything
in one clean sweep, and the supervisor hands the replacement fresh caps. A simpler
grant/ownership table can work too — capabilities are the principled version.
5. **Re-initialisable drivers.** A driver must start from a known state and
re-establish its hardware. Some hardware is easy to reset; some holds state that's
hard to recover — a real limit on what "just restart it" can fix.
## The hard part: restarting *correctly*
Detecting and killing is the easy half. The genuinely tricky questions are about the
*rest of the system* when a component dies:
- **In-flight IPC**: messages sent to the dead component, or replies its clients are
blocked waiting for. The channel has to break cleanly and unblock the waiters with
an error rather than hang them forever (a design constraint that reaches back into
[ipc.md](ipc.md) — channels need a "peer died" outcome).
- **Clients**: how does a client discover the service it was talking to is gone and
has been replaced? Options: capability revocation makes stale handles fail; or a
**name server** re-binds clients to the new instance; or clients retry through a
stable endpoint.
- **State**: the cheapest model is **stateless restart** — the replacement starts
fresh and clients re-establish whatever they need. Richer options (checkpointed
state, state handed to a standby) are more work and more failure modes. Start
stateless.
These are the constraints most worth *bumping into and researching* — they're where
resilience gets genuinely interesting.
## Kernel mechanism vs user-space policy
The microkernel split applies to fault management itself:
- **Kernel (mechanism):** isolation, trapping faults, enforcing capabilities/grants,
IPC, creating/destroying address spaces, granting/revoking resources, preempting a
runaway task.
- **User space (policy):** the supervisor decides *what* to restart, *when*, and
*how* — dependency order, retry limits, escalation. None of that belongs in the
kernel.
So the kernel gains a few primitives (kill an address space, reclaim its resources,
deliver a "child died" notification); everything smart lives in a user-space server.
## What's *not* recoverable this way
Honest boundaries:
- **The kernel itself.** It's the trusted base; if it faults, this mechanism can't
save it. The mitigation is to keep it tiny — the microkernel bet.
- **Corrupted hardware state.** Isolation limits the blast radius to one process, but
if a driver wedged the device itself, a restart may not un-wedge it.
- **Shared-resource corruption** that happened *before* the fault was detected. Clean
capability revocation limits this, but it's why fault *detection latency* matters.
## Suggested ordering
1. **User mode + address-space isolation** — the shared prerequisite (also on the
path for everything else).
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
"confine to the process and report it."
3. **A minimal supervisor server** that can (re)start a process.
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
table.
5. **First restartable driver** — the keyboard — as the end-to-end proof: crash it on
purpose, watch it come back.
## Relationship to real-time
Resilience needs **structural** features (isolation + supervision + a resource
model); real-time needs a **pervasive** timing invariant. They're separable, and
resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](vision.md)).
Note the overlap, though: **preemptive scheduling** and **priorities** — already
built — serve resilience too (you can preempt and kill a misbehaving component, and
run the supervisor at high priority). So danos keeps the useful *mechanisms* of the
real-time work without owing anyone a timing *guarantee*.
## Further reading
- Herder, Bos, Gras, Homburg, Tanenbaum — the **MINIX 3** papers, esp. *"Construction
of a Highly Dependable Operating System"* and *"Fault Isolation for Device
Drivers"* — the reincarnation server, the closest match to danos's goal.
- **QNX** architecture — a shipping microkernel with restartable drivers.
- **Erlang/OTP** supervision trees and the *"let it crash"* philosophy — the design
pattern, distilled.
- **seL4** capability model — the principled basis for clean resource teardown.
- **Tandem NonStop** (historical) — fault tolerance via process pairs.
## Related
- [vision.md](vision.md) — the goals this serves (learning by doing; resilience over
hard real-time).
- [scheduling.md](scheduling.md) — preemption, which makes runaway components killable.
- [ipc.md](ipc.md) — channels that need a "peer died" outcome for clean restart.
- [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and
restart" instead of "halt".
- [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context.
+87 -63
View File
@@ -1,84 +1,108 @@
# Vision: a real-time microkernel
# Vision: a microkernel, built to learn
danos is aiming to be a **real-time operating system built on a microkernel** —
where drivers and services run isolated in user space for maximum stability, and
scheduling gives real guarantees about timing. This page is the north star: the
*why* that shapes every design decision below it. Read it before adding anything
structural.
danos exists first and foremost as a **learning-by-doing project**: the point is to
build a real operating system, bump into the hard constraints for real, and research
them from a position of having actually hit them. The docs in this folder are part of
that — they're where a constraint gets understood once it's been met.
## Microkernel
That framing sets the priorities. danos is not chasing a spec or a product; it's
chasing understanding, with a concrete, motivating **win condition** to aim at.
The kernel stays **minimal** — only what genuinely must run in privileged mode:
## The win condition
danos is a "win" when it:
- **boots and runs on real hardware** — the author's **PC** (x86-64) and **both
Raspberry Pis**: the **Zero 2 W** and the **Pi 5** (both `aarch64`, one backend —
see [arm.md](arm.md)),
- **has a graphical user interface**, ideally — building on the framebuffer it
already draws to.
Everything below serves that, or serves the curiosity that the project runs on.
## Why a microkernel: resilience
The kernel stays **minimal** — only what genuinely must run privileged:
- scheduling,
- inter-process communication (IPC),
- memory management (address spaces, page tables),
- low-level interrupt dispatch.
Everything else — device drivers, filesystems, the network stack — runs as an
**isolated user-space server**, each in its own address space with only the
Everything else — device drivers, filesystems, the GUI, the network stack — runs as
an **isolated user-space server**, each in its own address space with only the
privileges it needs.
The payoff is **stability through isolation**. A driver bug can't corrupt the
kernel or another driver; a crashing service is contained and can be restarted,
while the rest of the system keeps running. That's the opposite of a monolithic
kernel, where a single driver fault can take everything down.
The reason for this shape is **resilience**: the ability to **re-initialise parts of
the OS while it runs**. A driver bug can't corrupt the kernel or another driver; a
crashed or wedged component is contained, killed, and **restarted** — "if I break
something, I can just fix it," without rebooting. Keeping the kernel tiny is part of
that strategy: the one thing that *can't* be restarted is the trusted base, so the
less code in it, the less that can take the whole system down. This is the project's
real motivation, and it has its own design note: [resilience.md](resilience.md).
The cost is that **IPC becomes the backbone**: whatever used to be a function call
across a monolithic kernel is now a message between address spaces. In a
microkernel, IPC performance essentially *is* system performance (the lesson of
L4). So IPC must be fast, and it's a first-class concern, not an afterthought.
Hardware interrupts, too, become IPC: the kernel turns an IRQ into a message to the
driver task that owns that device.
The cost is that **IPC becomes the backbone**: what used to be a function call inside
a monolithic kernel is now a message between address spaces. In a microkernel, IPC
performance essentially *is* system performance (the lesson of L4), so it's a
first-class concern. Hardware interrupts become IPC too: the kernel turns an IRQ into
a message to the driver that owns the device.
## Real-time
## On real-time: an option, not a commitment
danos schedules **preemptively, with guarantees about quanta** — the system must
be able to promise that a task runs when it's supposed to, within bounded time.
That imposes concrete requirements:
danos was originally framed as a hard **real-time** OS. That's now held as **one
interesting constraint to explore, not a requirement** — because real-time is a
*pervasive* invariant (every operation must be provably time-bounded, everywhere)
that would slow every milestone, whereas resilience is a set of *structural* features
that's lighter to build and is what the project actually wants. The trade-off is
written up in [smp.md](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience).
- **Fixed-priority preemptive scheduling.** The highest-priority ready task always
runs; a higher-priority task that becomes ready preempts a lower one immediately.
Not round-robin (which is fair but not predictable).
- **A calibrated, deterministic clock.** Guarantees measured in "quanta" are
meaningless on an arbitrary tick rate — real time requires a timer calibrated to
a known frequency.
- **Bounded interrupt latency.** Interrupt-disabled sections must be short and
bounded, so a ready high-priority task is never delayed by an unbounded kernel
operation.
- **Deterministic kernel operations.** Scheduling decisions should be O(1) (e.g. a
priority bitmap), not "walk a list of unknown length."
- **Priority inheritance** (once there are locks/IPC), so a high-priority task
blocked on a resource held by a low-priority one can't be delayed indefinitely by
a middle-priority task — bounding priority inversion.
What danos keeps from the real-time direction, because it's cheap and useful anyway:
A consequence worth stating early: the current [kernel heap](heap.md) is a
first-fit free list, which has **unbounded allocation time** and can fragment — it
is *not* real-time safe. It's fine for one-time kernel setup, but real-time paths
must pre-allocate or use a bounded (fixed-size pool) allocator. Don't allocate on a
hot real-time path.
- **Fixed-priority preemptive scheduling** — the highest-priority ready task runs, and
preemption lets a runaway component be interrupted and killed (which *serves
resilience*). Already built ([scheduling.md](scheduling.md)).
- **A calibrated, deterministic clock** — already built ([device-interrupts.md](device-interrupts.md)).
## What this means for the roadmap
What danos does *not* owe anyone unless it deliberately chooses real-time later:
timing *guarantees*, priority inheritance, bounded allocators, tickless timers, MCS
scheduling contexts. Concretely, the current [heap](heap.md) is a first-fit free list
with unbounded allocation time — fine here, and only a problem *if* a hard-real-time
path is ever added. Note that **QNX is both** a real-time and a restartable
microkernel, so choosing resilience now doesn't close the real-time door — it just
doesn't pay the tax yet.
The vision reorders the obvious hobby-kernel path. Notably, **drivers are not
built into the kernel** — so an in-kernel keyboard driver would be throwaway work.
Input devices arrive later, as the *first user-space drivers*, once the machinery
to isolate them exists. The trajectory:
## The roadmap — tracks, not a strict line
1. **Calibrated timer / clock** — a known-frequency, deterministic tick. The
foundation real-time quanta rest on. *(next)*
2. **Real-time scheduler** — fixed-priority preemptive, kernel threads first:
context switch, task struct, priority run-queue, timer-driven preemption.
3. **User mode + address-space isolation** — higher-half kernel, ring 3, per-process
page tables. The substrate for isolated servers.
4. **IPC** — fast message passing between address spaces. The microkernel's heart.
5. **User-space drivers** — interrupts delivered as IPC, plus MMIO/port-access
grants. The keyboard becomes the first one, validating the whole model.
Because the driver is curiosity plus the win condition, the roadmap is a set of
**tracks** with dependencies, not a rigid sequence. Pick by interest; mind the
prerequisites.
## Where we are
**Done:** UEFI boot, framebuffer + [serial](testing.md), [physical frames](frame-allocator.md)
(with boot-services memory reclaimed), [paging](paging.md) with W^X, [exceptions and
interrupts](interrupts.md), a [calibrated timer + ns clock](device-interrupts.md), a
[heap](heap.md), a [fixed-priority preemptive scheduler](scheduling.md) with blocking,
and in-kernel [IPC channels](ipc.md) — plus a [test harness](testing.md).
The foundation is in place: UEFI boot, framebuffer + [serial](testing.md),
[physical frames](frame-allocator.md), [paging](paging.md) with W^X, [exceptions
and interrupts](interrupts.md), a [timer](device-interrupts.md), and a
[heap](heap.md) — plus a [test harness](testing.md). The kernel boots and has its
core services; the next milestones make it *schedule*, then *isolate*.
- **Isolation track** — **user mode + address-space isolation** (higher-half kernel,
ring 3, per-process page tables). The substrate everything else needs. *Next, and a
prerequisite for the resilience and driver tracks.*
- **Resilience track** — fault → kill → notify, a supervisor/reincarnation server,
resource cleanup on death, then a restartable driver as proof. Needs isolation.
See [resilience.md](resilience.md).
- **ARM track** — the `aarch64` port so danos runs on the Zero 2 W and Pi 5. Largely
independent of the others (it's the [arch layer](arch.md)); directly serves the win
condition. Likely via aarch64-UEFI first (QEMU `virt` + AAVMF), then real boards.
See [arm.md](arm.md), and [discovery.md](discovery.md) for the device tree it needs.
- **GUI track** — a framebuffer-based windowing/compositor, and the input + display
drivers under it. Builds on the neutral framebuffer (so it's arch-independent), and
on the driver model from the isolation/resilience tracks. The visible payoff.
The natural spine is **isolation → (resilience + drivers) → GUI**, with the **ARM
track** pursued alongside whenever the itch to see it boot on a Pi wins out.
## How to use this page
Read it before adding anything structural. When a design decision comes up, the
question is: does it serve the **win condition** (runs on the three machines, with a
GUI), or the **learning** (a constraint worth meeting)? If it serves neither — e.g.
paying the full real-time tax with no payoff in sight — it can wait.