documenting things to research
This commit is contained in:
+8
-3
@@ -42,9 +42,14 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
|
||||
Start with the north star:
|
||||
|
||||
- **[vision.md](vision.md) — the vision.** danos is aiming to be a real-time
|
||||
microkernel: minimal kernel, drivers/services isolated in user space, preemptive
|
||||
scheduling with timing guarantees. The *why* that shapes everything below.
|
||||
- **[vision.md](vision.md) — the vision.** danos is a **learning-by-doing**
|
||||
microkernel: minimal kernel, drivers/services isolated in user space, chosen for
|
||||
**resilience** (restartable components). Win condition: runs on the author's PC and
|
||||
both Raspberry Pis, ideally with a GUI. Real-time is an option to explore, not a
|
||||
requirement. The *why* that shapes everything below.
|
||||
- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on
|
||||
fault isolation + live restart — the reincarnation-server + capability model that
|
||||
makes "if I break it, I can restart it" real. danos's core motivation.
|
||||
|
||||
Cutting across all of these:
|
||||
|
||||
|
||||
@@ -0,0 +1,152 @@
|
||||
# Resilience: fault isolation and live restart
|
||||
|
||||
A design/research note, not built yet. This is the property danos is really chasing:
|
||||
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
|
||||
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
|
||||
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
|
||||
from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) —
|
||||
one that's less pervasive to build (see [vision.md](vision.md)).
|
||||
|
||||
## The idea: "let it crash" + supervision
|
||||
|
||||
The philosophy is older than microkernels and shows up across systems: don't try to
|
||||
make every component perfect — make failures **contained and recoverable**. Isolate
|
||||
each component, watch it, and when it dies, restart it from a known-good state. A
|
||||
small trusted core supervises a fleet of restartable, untrusted parts.
|
||||
|
||||
Prior art worth studying (see Further reading): **MINIX 3's reincarnation server** (a
|
||||
driver crashes, a supervisor restarts it live — the closest thing to your goal),
|
||||
**QNX** (restartable drivers on a message-passing microkernel), **Erlang/OTP
|
||||
supervision trees** ("let it crash", not a kernel but the canonical design), and the
|
||||
historical **Tandem NonStop** (fault-tolerant by process pairs).
|
||||
|
||||
## Why a microkernel makes this possible
|
||||
|
||||
The blast radius of a fault is the address space it happens in. In a monolith, a
|
||||
driver bug can corrupt anything — the kernel *is* the driver. In a microkernel,
|
||||
drivers and services are **isolated user-space processes**, so a fault is trapped by
|
||||
the kernel and confined to that one process. The kernel — the one thing that *can't*
|
||||
be restarted, because it's the trusted base — stays tiny, which is precisely why a
|
||||
small kernel is a *more recoverable* kernel: less code that can take the whole system
|
||||
down. **Keeping the kernel minimal is a resilience strategy, not just an aesthetic.**
|
||||
|
||||
## The building blocks
|
||||
|
||||
1. **Address-space isolation.** A fault in one component can't corrupt another or the
|
||||
kernel. This is the [user-mode milestone](vision.md) (ring 3, per-process page
|
||||
tables) — the shared prerequisite for *any* of this, and it's needed regardless.
|
||||
2. **Fault detection** — how the system notices a component is dead or sick:
|
||||
- **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps
|
||||
to the kernel, which kills the process and notifies the supervisor. The clean,
|
||||
easy case — and danos already reports CPU faults (see
|
||||
[interrupts.md](interrupts.md)); user mode turns "halt on fault" into "kill the
|
||||
process and tell the supervisor."
|
||||
- **Hang**: a livelocked or infinite-looping component needs a **watchdog /
|
||||
heartbeat** and the ability to **preempt and kill** it. The preemptive scheduler
|
||||
already built ([scheduling.md](scheduling.md)) is what makes a runaway component
|
||||
killable — a nice case of a scheduling mechanism serving resilience without any
|
||||
real-time *guarantee*.
|
||||
- **Misbehaviour**: IPC timeouts, failed health checks.
|
||||
3. **A supervisor / reincarnation server.** A user-space server holding the *policy*:
|
||||
what components exist, their dependencies, and each one's restart strategy. When a
|
||||
component dies, it decides whether/how to restart it. (MINIX 3 calls this the
|
||||
reincarnation server; Erlang calls it a supervisor.)
|
||||
4. **A resource model that supports clean teardown.** When a component dies, its
|
||||
resources — memory, MMIO grants, IPC channels, IRQ routes — must be **reclaimed**,
|
||||
and a restarted replacement must be able to **re-acquire** them. This is where a
|
||||
**capability** model shines (seL4's reference design): a component holds
|
||||
capabilities to its resources; killing it **revokes** them, which frees everything
|
||||
in one clean sweep, and the supervisor hands the replacement fresh caps. A simpler
|
||||
grant/ownership table can work too — capabilities are the principled version.
|
||||
5. **Re-initialisable drivers.** A driver must start from a known state and
|
||||
re-establish its hardware. Some hardware is easy to reset; some holds state that's
|
||||
hard to recover — a real limit on what "just restart it" can fix.
|
||||
|
||||
## The hard part: restarting *correctly*
|
||||
|
||||
Detecting and killing is the easy half. The genuinely tricky questions are about the
|
||||
*rest of the system* when a component dies:
|
||||
|
||||
- **In-flight IPC**: messages sent to the dead component, or replies its clients are
|
||||
blocked waiting for. The channel has to break cleanly and unblock the waiters with
|
||||
an error rather than hang them forever (a design constraint that reaches back into
|
||||
[ipc.md](ipc.md) — channels need a "peer died" outcome).
|
||||
- **Clients**: how does a client discover the service it was talking to is gone and
|
||||
has been replaced? Options: capability revocation makes stale handles fail; or a
|
||||
**name server** re-binds clients to the new instance; or clients retry through a
|
||||
stable endpoint.
|
||||
- **State**: the cheapest model is **stateless restart** — the replacement starts
|
||||
fresh and clients re-establish whatever they need. Richer options (checkpointed
|
||||
state, state handed to a standby) are more work and more failure modes. Start
|
||||
stateless.
|
||||
|
||||
These are the constraints most worth *bumping into and researching* — they're where
|
||||
resilience gets genuinely interesting.
|
||||
|
||||
## Kernel mechanism vs user-space policy
|
||||
|
||||
The microkernel split applies to fault management itself:
|
||||
|
||||
- **Kernel (mechanism):** isolation, trapping faults, enforcing capabilities/grants,
|
||||
IPC, creating/destroying address spaces, granting/revoking resources, preempting a
|
||||
runaway task.
|
||||
- **User space (policy):** the supervisor decides *what* to restart, *when*, and
|
||||
*how* — dependency order, retry limits, escalation. None of that belongs in the
|
||||
kernel.
|
||||
|
||||
So the kernel gains a few primitives (kill an address space, reclaim its resources,
|
||||
deliver a "child died" notification); everything smart lives in a user-space server.
|
||||
|
||||
## What's *not* recoverable this way
|
||||
|
||||
Honest boundaries:
|
||||
|
||||
- **The kernel itself.** It's the trusted base; if it faults, this mechanism can't
|
||||
save it. The mitigation is to keep it tiny — the microkernel bet.
|
||||
- **Corrupted hardware state.** Isolation limits the blast radius to one process, but
|
||||
if a driver wedged the device itself, a restart may not un-wedge it.
|
||||
- **Shared-resource corruption** that happened *before* the fault was detected. Clean
|
||||
capability revocation limits this, but it's why fault *detection latency* matters.
|
||||
|
||||
## Suggested ordering
|
||||
|
||||
1. **User mode + address-space isolation** — the shared prerequisite (also on the
|
||||
path for everything else).
|
||||
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
|
||||
"confine to the process and report it."
|
||||
3. **A minimal supervisor server** that can (re)start a process.
|
||||
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
|
||||
table.
|
||||
5. **First restartable driver** — the keyboard — as the end-to-end proof: crash it on
|
||||
purpose, watch it come back.
|
||||
|
||||
## Relationship to real-time
|
||||
|
||||
Resilience needs **structural** features (isolation + supervision + a resource
|
||||
model); real-time needs a **pervasive** timing invariant. They're separable, and
|
||||
resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](vision.md)).
|
||||
Note the overlap, though: **preemptive scheduling** and **priorities** — already
|
||||
built — serve resilience too (you can preempt and kill a misbehaving component, and
|
||||
run the supervisor at high priority). So danos keeps the useful *mechanisms* of the
|
||||
real-time work without owing anyone a timing *guarantee*.
|
||||
|
||||
## Further reading
|
||||
|
||||
- Herder, Bos, Gras, Homburg, Tanenbaum — the **MINIX 3** papers, esp. *"Construction
|
||||
of a Highly Dependable Operating System"* and *"Fault Isolation for Device
|
||||
Drivers"* — the reincarnation server, the closest match to danos's goal.
|
||||
- **QNX** architecture — a shipping microkernel with restartable drivers.
|
||||
- **Erlang/OTP** supervision trees and the *"let it crash"* philosophy — the design
|
||||
pattern, distilled.
|
||||
- **seL4** capability model — the principled basis for clean resource teardown.
|
||||
- **Tandem NonStop** (historical) — fault tolerance via process pairs.
|
||||
|
||||
## Related
|
||||
|
||||
- [vision.md](vision.md) — the goals this serves (learning by doing; resilience over
|
||||
hard real-time).
|
||||
- [scheduling.md](scheduling.md) — preemption, which makes runaway components killable.
|
||||
- [ipc.md](ipc.md) — channels that need a "peer died" outcome for clean restart.
|
||||
- [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and
|
||||
restart" instead of "halt".
|
||||
- [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context.
|
||||
+87
-63
@@ -1,84 +1,108 @@
|
||||
# Vision: a real-time microkernel
|
||||
# Vision: a microkernel, built to learn
|
||||
|
||||
danos is aiming to be a **real-time operating system built on a microkernel** —
|
||||
where drivers and services run isolated in user space for maximum stability, and
|
||||
scheduling gives real guarantees about timing. This page is the north star: the
|
||||
*why* that shapes every design decision below it. Read it before adding anything
|
||||
structural.
|
||||
danos exists first and foremost as a **learning-by-doing project**: the point is to
|
||||
build a real operating system, bump into the hard constraints for real, and research
|
||||
them from a position of having actually hit them. The docs in this folder are part of
|
||||
that — they're where a constraint gets understood once it's been met.
|
||||
|
||||
## Microkernel
|
||||
That framing sets the priorities. danos is not chasing a spec or a product; it's
|
||||
chasing understanding, with a concrete, motivating **win condition** to aim at.
|
||||
|
||||
The kernel stays **minimal** — only what genuinely must run in privileged mode:
|
||||
## The win condition
|
||||
|
||||
danos is a "win" when it:
|
||||
|
||||
- **boots and runs on real hardware** — the author's **PC** (x86-64) and **both
|
||||
Raspberry Pis**: the **Zero 2 W** and the **Pi 5** (both `aarch64`, one backend —
|
||||
see [arm.md](arm.md)),
|
||||
- **has a graphical user interface**, ideally — building on the framebuffer it
|
||||
already draws to.
|
||||
|
||||
Everything below serves that, or serves the curiosity that the project runs on.
|
||||
|
||||
## Why a microkernel: resilience
|
||||
|
||||
The kernel stays **minimal** — only what genuinely must run privileged:
|
||||
|
||||
- scheduling,
|
||||
- inter-process communication (IPC),
|
||||
- memory management (address spaces, page tables),
|
||||
- low-level interrupt dispatch.
|
||||
|
||||
Everything else — device drivers, filesystems, the network stack — runs as an
|
||||
**isolated user-space server**, each in its own address space with only the
|
||||
Everything else — device drivers, filesystems, the GUI, the network stack — runs as
|
||||
an **isolated user-space server**, each in its own address space with only the
|
||||
privileges it needs.
|
||||
|
||||
The payoff is **stability through isolation**. A driver bug can't corrupt the
|
||||
kernel or another driver; a crashing service is contained and can be restarted,
|
||||
while the rest of the system keeps running. That's the opposite of a monolithic
|
||||
kernel, where a single driver fault can take everything down.
|
||||
The reason for this shape is **resilience**: the ability to **re-initialise parts of
|
||||
the OS while it runs**. A driver bug can't corrupt the kernel or another driver; a
|
||||
crashed or wedged component is contained, killed, and **restarted** — "if I break
|
||||
something, I can just fix it," without rebooting. Keeping the kernel tiny is part of
|
||||
that strategy: the one thing that *can't* be restarted is the trusted base, so the
|
||||
less code in it, the less that can take the whole system down. This is the project's
|
||||
real motivation, and it has its own design note: [resilience.md](resilience.md).
|
||||
|
||||
The cost is that **IPC becomes the backbone**: whatever used to be a function call
|
||||
across a monolithic kernel is now a message between address spaces. In a
|
||||
microkernel, IPC performance essentially *is* system performance (the lesson of
|
||||
L4). So IPC must be fast, and it's a first-class concern, not an afterthought.
|
||||
Hardware interrupts, too, become IPC: the kernel turns an IRQ into a message to the
|
||||
driver task that owns that device.
|
||||
The cost is that **IPC becomes the backbone**: what used to be a function call inside
|
||||
a monolithic kernel is now a message between address spaces. In a microkernel, IPC
|
||||
performance essentially *is* system performance (the lesson of L4), so it's a
|
||||
first-class concern. Hardware interrupts become IPC too: the kernel turns an IRQ into
|
||||
a message to the driver that owns the device.
|
||||
|
||||
## Real-time
|
||||
## On real-time: an option, not a commitment
|
||||
|
||||
danos schedules **preemptively, with guarantees about quanta** — the system must
|
||||
be able to promise that a task runs when it's supposed to, within bounded time.
|
||||
That imposes concrete requirements:
|
||||
danos was originally framed as a hard **real-time** OS. That's now held as **one
|
||||
interesting constraint to explore, not a requirement** — because real-time is a
|
||||
*pervasive* invariant (every operation must be provably time-bounded, everywhere)
|
||||
that would slow every milestone, whereas resilience is a set of *structural* features
|
||||
that's lighter to build and is what the project actually wants. The trade-off is
|
||||
written up in [smp.md](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience).
|
||||
|
||||
- **Fixed-priority preemptive scheduling.** The highest-priority ready task always
|
||||
runs; a higher-priority task that becomes ready preempts a lower one immediately.
|
||||
Not round-robin (which is fair but not predictable).
|
||||
- **A calibrated, deterministic clock.** Guarantees measured in "quanta" are
|
||||
meaningless on an arbitrary tick rate — real time requires a timer calibrated to
|
||||
a known frequency.
|
||||
- **Bounded interrupt latency.** Interrupt-disabled sections must be short and
|
||||
bounded, so a ready high-priority task is never delayed by an unbounded kernel
|
||||
operation.
|
||||
- **Deterministic kernel operations.** Scheduling decisions should be O(1) (e.g. a
|
||||
priority bitmap), not "walk a list of unknown length."
|
||||
- **Priority inheritance** (once there are locks/IPC), so a high-priority task
|
||||
blocked on a resource held by a low-priority one can't be delayed indefinitely by
|
||||
a middle-priority task — bounding priority inversion.
|
||||
What danos keeps from the real-time direction, because it's cheap and useful anyway:
|
||||
|
||||
A consequence worth stating early: the current [kernel heap](heap.md) is a
|
||||
first-fit free list, which has **unbounded allocation time** and can fragment — it
|
||||
is *not* real-time safe. It's fine for one-time kernel setup, but real-time paths
|
||||
must pre-allocate or use a bounded (fixed-size pool) allocator. Don't allocate on a
|
||||
hot real-time path.
|
||||
- **Fixed-priority preemptive scheduling** — the highest-priority ready task runs, and
|
||||
preemption lets a runaway component be interrupted and killed (which *serves
|
||||
resilience*). Already built ([scheduling.md](scheduling.md)).
|
||||
- **A calibrated, deterministic clock** — already built ([device-interrupts.md](device-interrupts.md)).
|
||||
|
||||
## What this means for the roadmap
|
||||
What danos does *not* owe anyone unless it deliberately chooses real-time later:
|
||||
timing *guarantees*, priority inheritance, bounded allocators, tickless timers, MCS
|
||||
scheduling contexts. Concretely, the current [heap](heap.md) is a first-fit free list
|
||||
with unbounded allocation time — fine here, and only a problem *if* a hard-real-time
|
||||
path is ever added. Note that **QNX is both** a real-time and a restartable
|
||||
microkernel, so choosing resilience now doesn't close the real-time door — it just
|
||||
doesn't pay the tax yet.
|
||||
|
||||
The vision reorders the obvious hobby-kernel path. Notably, **drivers are not
|
||||
built into the kernel** — so an in-kernel keyboard driver would be throwaway work.
|
||||
Input devices arrive later, as the *first user-space drivers*, once the machinery
|
||||
to isolate them exists. The trajectory:
|
||||
## The roadmap — tracks, not a strict line
|
||||
|
||||
1. **Calibrated timer / clock** — a known-frequency, deterministic tick. The
|
||||
foundation real-time quanta rest on. *(next)*
|
||||
2. **Real-time scheduler** — fixed-priority preemptive, kernel threads first:
|
||||
context switch, task struct, priority run-queue, timer-driven preemption.
|
||||
3. **User mode + address-space isolation** — higher-half kernel, ring 3, per-process
|
||||
page tables. The substrate for isolated servers.
|
||||
4. **IPC** — fast message passing between address spaces. The microkernel's heart.
|
||||
5. **User-space drivers** — interrupts delivered as IPC, plus MMIO/port-access
|
||||
grants. The keyboard becomes the first one, validating the whole model.
|
||||
Because the driver is curiosity plus the win condition, the roadmap is a set of
|
||||
**tracks** with dependencies, not a rigid sequence. Pick by interest; mind the
|
||||
prerequisites.
|
||||
|
||||
## Where we are
|
||||
**Done:** UEFI boot, framebuffer + [serial](testing.md), [physical frames](frame-allocator.md)
|
||||
(with boot-services memory reclaimed), [paging](paging.md) with W^X, [exceptions and
|
||||
interrupts](interrupts.md), a [calibrated timer + ns clock](device-interrupts.md), a
|
||||
[heap](heap.md), a [fixed-priority preemptive scheduler](scheduling.md) with blocking,
|
||||
and in-kernel [IPC channels](ipc.md) — plus a [test harness](testing.md).
|
||||
|
||||
The foundation is in place: UEFI boot, framebuffer + [serial](testing.md),
|
||||
[physical frames](frame-allocator.md), [paging](paging.md) with W^X, [exceptions
|
||||
and interrupts](interrupts.md), a [timer](device-interrupts.md), and a
|
||||
[heap](heap.md) — plus a [test harness](testing.md). The kernel boots and has its
|
||||
core services; the next milestones make it *schedule*, then *isolate*.
|
||||
- **Isolation track** — **user mode + address-space isolation** (higher-half kernel,
|
||||
ring 3, per-process page tables). The substrate everything else needs. *Next, and a
|
||||
prerequisite for the resilience and driver tracks.*
|
||||
- **Resilience track** — fault → kill → notify, a supervisor/reincarnation server,
|
||||
resource cleanup on death, then a restartable driver as proof. Needs isolation.
|
||||
See [resilience.md](resilience.md).
|
||||
- **ARM track** — the `aarch64` port so danos runs on the Zero 2 W and Pi 5. Largely
|
||||
independent of the others (it's the [arch layer](arch.md)); directly serves the win
|
||||
condition. Likely via aarch64-UEFI first (QEMU `virt` + AAVMF), then real boards.
|
||||
See [arm.md](arm.md), and [discovery.md](discovery.md) for the device tree it needs.
|
||||
- **GUI track** — a framebuffer-based windowing/compositor, and the input + display
|
||||
drivers under it. Builds on the neutral framebuffer (so it's arch-independent), and
|
||||
on the driver model from the isolation/resilience tracks. The visible payoff.
|
||||
|
||||
The natural spine is **isolation → (resilience + drivers) → GUI**, with the **ARM
|
||||
track** pursued alongside whenever the itch to see it boot on a Pi wins out.
|
||||
|
||||
## How to use this page
|
||||
|
||||
Read it before adding anything structural. When a design decision comes up, the
|
||||
question is: does it serve the **win condition** (runs on the three machines, with a
|
||||
GUI), or the **learning** (a constraint worth meeting)? If it serves neither — e.g.
|
||||
paying the full real-time tax with no payoff in sight — it can wait.
|
||||
|
||||
Reference in New Issue
Block a user