documenting things to research
This commit is contained in:
+8
-3
@@ -42,9 +42,14 @@ rather than restate it. Roughly in the order things happen at runtime:
|
|||||||
|
|
||||||
Start with the north star:
|
Start with the north star:
|
||||||
|
|
||||||
- **[vision.md](vision.md) — the vision.** danos is aiming to be a real-time
|
- **[vision.md](vision.md) — the vision.** danos is a **learning-by-doing**
|
||||||
microkernel: minimal kernel, drivers/services isolated in user space, preemptive
|
microkernel: minimal kernel, drivers/services isolated in user space, chosen for
|
||||||
scheduling with timing guarantees. The *why* that shapes everything below.
|
**resilience** (restartable components). Win condition: runs on the author's PC and
|
||||||
|
both Raspberry Pis, ideally with a GUI. Real-time is an option to explore, not a
|
||||||
|
requirement. The *why* that shapes everything below.
|
||||||
|
- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on
|
||||||
|
fault isolation + live restart — the reincarnation-server + capability model that
|
||||||
|
makes "if I break it, I can restart it" real. danos's core motivation.
|
||||||
|
|
||||||
Cutting across all of these:
|
Cutting across all of these:
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,152 @@
|
|||||||
|
# Resilience: fault isolation and live restart
|
||||||
|
|
||||||
|
A design/research note, not built yet. This is the property danos is really chasing:
|
||||||
|
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
|
||||||
|
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
|
||||||
|
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
|
||||||
|
from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) —
|
||||||
|
one that's less pervasive to build (see [vision.md](vision.md)).
|
||||||
|
|
||||||
|
## The idea: "let it crash" + supervision
|
||||||
|
|
||||||
|
The philosophy is older than microkernels and shows up across systems: don't try to
|
||||||
|
make every component perfect — make failures **contained and recoverable**. Isolate
|
||||||
|
each component, watch it, and when it dies, restart it from a known-good state. A
|
||||||
|
small trusted core supervises a fleet of restartable, untrusted parts.
|
||||||
|
|
||||||
|
Prior art worth studying (see Further reading): **MINIX 3's reincarnation server** (a
|
||||||
|
driver crashes, a supervisor restarts it live — the closest thing to your goal),
|
||||||
|
**QNX** (restartable drivers on a message-passing microkernel), **Erlang/OTP
|
||||||
|
supervision trees** ("let it crash", not a kernel but the canonical design), and the
|
||||||
|
historical **Tandem NonStop** (fault-tolerant by process pairs).
|
||||||
|
|
||||||
|
## Why a microkernel makes this possible
|
||||||
|
|
||||||
|
The blast radius of a fault is the address space it happens in. In a monolith, a
|
||||||
|
driver bug can corrupt anything — the kernel *is* the driver. In a microkernel,
|
||||||
|
drivers and services are **isolated user-space processes**, so a fault is trapped by
|
||||||
|
the kernel and confined to that one process. The kernel — the one thing that *can't*
|
||||||
|
be restarted, because it's the trusted base — stays tiny, which is precisely why a
|
||||||
|
small kernel is a *more recoverable* kernel: less code that can take the whole system
|
||||||
|
down. **Keeping the kernel minimal is a resilience strategy, not just an aesthetic.**
|
||||||
|
|
||||||
|
## The building blocks
|
||||||
|
|
||||||
|
1. **Address-space isolation.** A fault in one component can't corrupt another or the
|
||||||
|
kernel. This is the [user-mode milestone](vision.md) (ring 3, per-process page
|
||||||
|
tables) — the shared prerequisite for *any* of this, and it's needed regardless.
|
||||||
|
2. **Fault detection** — how the system notices a component is dead or sick:
|
||||||
|
- **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps
|
||||||
|
to the kernel, which kills the process and notifies the supervisor. The clean,
|
||||||
|
easy case — and danos already reports CPU faults (see
|
||||||
|
[interrupts.md](interrupts.md)); user mode turns "halt on fault" into "kill the
|
||||||
|
process and tell the supervisor."
|
||||||
|
- **Hang**: a livelocked or infinite-looping component needs a **watchdog /
|
||||||
|
heartbeat** and the ability to **preempt and kill** it. The preemptive scheduler
|
||||||
|
already built ([scheduling.md](scheduling.md)) is what makes a runaway component
|
||||||
|
killable — a nice case of a scheduling mechanism serving resilience without any
|
||||||
|
real-time *guarantee*.
|
||||||
|
- **Misbehaviour**: IPC timeouts, failed health checks.
|
||||||
|
3. **A supervisor / reincarnation server.** A user-space server holding the *policy*:
|
||||||
|
what components exist, their dependencies, and each one's restart strategy. When a
|
||||||
|
component dies, it decides whether/how to restart it. (MINIX 3 calls this the
|
||||||
|
reincarnation server; Erlang calls it a supervisor.)
|
||||||
|
4. **A resource model that supports clean teardown.** When a component dies, its
|
||||||
|
resources — memory, MMIO grants, IPC channels, IRQ routes — must be **reclaimed**,
|
||||||
|
and a restarted replacement must be able to **re-acquire** them. This is where a
|
||||||
|
**capability** model shines (seL4's reference design): a component holds
|
||||||
|
capabilities to its resources; killing it **revokes** them, which frees everything
|
||||||
|
in one clean sweep, and the supervisor hands the replacement fresh caps. A simpler
|
||||||
|
grant/ownership table can work too — capabilities are the principled version.
|
||||||
|
5. **Re-initialisable drivers.** A driver must start from a known state and
|
||||||
|
re-establish its hardware. Some hardware is easy to reset; some holds state that's
|
||||||
|
hard to recover — a real limit on what "just restart it" can fix.
|
||||||
|
|
||||||
|
## The hard part: restarting *correctly*
|
||||||
|
|
||||||
|
Detecting and killing is the easy half. The genuinely tricky questions are about the
|
||||||
|
*rest of the system* when a component dies:
|
||||||
|
|
||||||
|
- **In-flight IPC**: messages sent to the dead component, or replies its clients are
|
||||||
|
blocked waiting for. The channel has to break cleanly and unblock the waiters with
|
||||||
|
an error rather than hang them forever (a design constraint that reaches back into
|
||||||
|
[ipc.md](ipc.md) — channels need a "peer died" outcome).
|
||||||
|
- **Clients**: how does a client discover the service it was talking to is gone and
|
||||||
|
has been replaced? Options: capability revocation makes stale handles fail; or a
|
||||||
|
**name server** re-binds clients to the new instance; or clients retry through a
|
||||||
|
stable endpoint.
|
||||||
|
- **State**: the cheapest model is **stateless restart** — the replacement starts
|
||||||
|
fresh and clients re-establish whatever they need. Richer options (checkpointed
|
||||||
|
state, state handed to a standby) are more work and more failure modes. Start
|
||||||
|
stateless.
|
||||||
|
|
||||||
|
These are the constraints most worth *bumping into and researching* — they're where
|
||||||
|
resilience gets genuinely interesting.
|
||||||
|
|
||||||
|
## Kernel mechanism vs user-space policy
|
||||||
|
|
||||||
|
The microkernel split applies to fault management itself:
|
||||||
|
|
||||||
|
- **Kernel (mechanism):** isolation, trapping faults, enforcing capabilities/grants,
|
||||||
|
IPC, creating/destroying address spaces, granting/revoking resources, preempting a
|
||||||
|
runaway task.
|
||||||
|
- **User space (policy):** the supervisor decides *what* to restart, *when*, and
|
||||||
|
*how* — dependency order, retry limits, escalation. None of that belongs in the
|
||||||
|
kernel.
|
||||||
|
|
||||||
|
So the kernel gains a few primitives (kill an address space, reclaim its resources,
|
||||||
|
deliver a "child died" notification); everything smart lives in a user-space server.
|
||||||
|
|
||||||
|
## What's *not* recoverable this way
|
||||||
|
|
||||||
|
Honest boundaries:
|
||||||
|
|
||||||
|
- **The kernel itself.** It's the trusted base; if it faults, this mechanism can't
|
||||||
|
save it. The mitigation is to keep it tiny — the microkernel bet.
|
||||||
|
- **Corrupted hardware state.** Isolation limits the blast radius to one process, but
|
||||||
|
if a driver wedged the device itself, a restart may not un-wedge it.
|
||||||
|
- **Shared-resource corruption** that happened *before* the fault was detected. Clean
|
||||||
|
capability revocation limits this, but it's why fault *detection latency* matters.
|
||||||
|
|
||||||
|
## Suggested ordering
|
||||||
|
|
||||||
|
1. **User mode + address-space isolation** — the shared prerequisite (also on the
|
||||||
|
path for everything else).
|
||||||
|
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
|
||||||
|
"confine to the process and report it."
|
||||||
|
3. **A minimal supervisor server** that can (re)start a process.
|
||||||
|
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
|
||||||
|
table.
|
||||||
|
5. **First restartable driver** — the keyboard — as the end-to-end proof: crash it on
|
||||||
|
purpose, watch it come back.
|
||||||
|
|
||||||
|
## Relationship to real-time
|
||||||
|
|
||||||
|
Resilience needs **structural** features (isolation + supervision + a resource
|
||||||
|
model); real-time needs a **pervasive** timing invariant. They're separable, and
|
||||||
|
resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](vision.md)).
|
||||||
|
Note the overlap, though: **preemptive scheduling** and **priorities** — already
|
||||||
|
built — serve resilience too (you can preempt and kill a misbehaving component, and
|
||||||
|
run the supervisor at high priority). So danos keeps the useful *mechanisms* of the
|
||||||
|
real-time work without owing anyone a timing *guarantee*.
|
||||||
|
|
||||||
|
## Further reading
|
||||||
|
|
||||||
|
- Herder, Bos, Gras, Homburg, Tanenbaum — the **MINIX 3** papers, esp. *"Construction
|
||||||
|
of a Highly Dependable Operating System"* and *"Fault Isolation for Device
|
||||||
|
Drivers"* — the reincarnation server, the closest match to danos's goal.
|
||||||
|
- **QNX** architecture — a shipping microkernel with restartable drivers.
|
||||||
|
- **Erlang/OTP** supervision trees and the *"let it crash"* philosophy — the design
|
||||||
|
pattern, distilled.
|
||||||
|
- **seL4** capability model — the principled basis for clean resource teardown.
|
||||||
|
- **Tandem NonStop** (historical) — fault tolerance via process pairs.
|
||||||
|
|
||||||
|
## Related
|
||||||
|
|
||||||
|
- [vision.md](vision.md) — the goals this serves (learning by doing; resilience over
|
||||||
|
hard real-time).
|
||||||
|
- [scheduling.md](scheduling.md) — preemption, which makes runaway components killable.
|
||||||
|
- [ipc.md](ipc.md) — channels that need a "peer died" outcome for clean restart.
|
||||||
|
- [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and
|
||||||
|
restart" instead of "halt".
|
||||||
|
- [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context.
|
||||||
+87
-63
@@ -1,84 +1,108 @@
|
|||||||
# Vision: a real-time microkernel
|
# Vision: a microkernel, built to learn
|
||||||
|
|
||||||
danos is aiming to be a **real-time operating system built on a microkernel** —
|
danos exists first and foremost as a **learning-by-doing project**: the point is to
|
||||||
where drivers and services run isolated in user space for maximum stability, and
|
build a real operating system, bump into the hard constraints for real, and research
|
||||||
scheduling gives real guarantees about timing. This page is the north star: the
|
them from a position of having actually hit them. The docs in this folder are part of
|
||||||
*why* that shapes every design decision below it. Read it before adding anything
|
that — they're where a constraint gets understood once it's been met.
|
||||||
structural.
|
|
||||||
|
|
||||||
## Microkernel
|
That framing sets the priorities. danos is not chasing a spec or a product; it's
|
||||||
|
chasing understanding, with a concrete, motivating **win condition** to aim at.
|
||||||
|
|
||||||
The kernel stays **minimal** — only what genuinely must run in privileged mode:
|
## The win condition
|
||||||
|
|
||||||
|
danos is a "win" when it:
|
||||||
|
|
||||||
|
- **boots and runs on real hardware** — the author's **PC** (x86-64) and **both
|
||||||
|
Raspberry Pis**: the **Zero 2 W** and the **Pi 5** (both `aarch64`, one backend —
|
||||||
|
see [arm.md](arm.md)),
|
||||||
|
- **has a graphical user interface**, ideally — building on the framebuffer it
|
||||||
|
already draws to.
|
||||||
|
|
||||||
|
Everything below serves that, or serves the curiosity that the project runs on.
|
||||||
|
|
||||||
|
## Why a microkernel: resilience
|
||||||
|
|
||||||
|
The kernel stays **minimal** — only what genuinely must run privileged:
|
||||||
|
|
||||||
- scheduling,
|
- scheduling,
|
||||||
- inter-process communication (IPC),
|
- inter-process communication (IPC),
|
||||||
- memory management (address spaces, page tables),
|
- memory management (address spaces, page tables),
|
||||||
- low-level interrupt dispatch.
|
- low-level interrupt dispatch.
|
||||||
|
|
||||||
Everything else — device drivers, filesystems, the network stack — runs as an
|
Everything else — device drivers, filesystems, the GUI, the network stack — runs as
|
||||||
**isolated user-space server**, each in its own address space with only the
|
an **isolated user-space server**, each in its own address space with only the
|
||||||
privileges it needs.
|
privileges it needs.
|
||||||
|
|
||||||
The payoff is **stability through isolation**. A driver bug can't corrupt the
|
The reason for this shape is **resilience**: the ability to **re-initialise parts of
|
||||||
kernel or another driver; a crashing service is contained and can be restarted,
|
the OS while it runs**. A driver bug can't corrupt the kernel or another driver; a
|
||||||
while the rest of the system keeps running. That's the opposite of a monolithic
|
crashed or wedged component is contained, killed, and **restarted** — "if I break
|
||||||
kernel, where a single driver fault can take everything down.
|
something, I can just fix it," without rebooting. Keeping the kernel tiny is part of
|
||||||
|
that strategy: the one thing that *can't* be restarted is the trusted base, so the
|
||||||
|
less code in it, the less that can take the whole system down. This is the project's
|
||||||
|
real motivation, and it has its own design note: [resilience.md](resilience.md).
|
||||||
|
|
||||||
The cost is that **IPC becomes the backbone**: whatever used to be a function call
|
The cost is that **IPC becomes the backbone**: what used to be a function call inside
|
||||||
across a monolithic kernel is now a message between address spaces. In a
|
a monolithic kernel is now a message between address spaces. In a microkernel, IPC
|
||||||
microkernel, IPC performance essentially *is* system performance (the lesson of
|
performance essentially *is* system performance (the lesson of L4), so it's a
|
||||||
L4). So IPC must be fast, and it's a first-class concern, not an afterthought.
|
first-class concern. Hardware interrupts become IPC too: the kernel turns an IRQ into
|
||||||
Hardware interrupts, too, become IPC: the kernel turns an IRQ into a message to the
|
a message to the driver that owns the device.
|
||||||
driver task that owns that device.
|
|
||||||
|
|
||||||
## Real-time
|
## On real-time: an option, not a commitment
|
||||||
|
|
||||||
danos schedules **preemptively, with guarantees about quanta** — the system must
|
danos was originally framed as a hard **real-time** OS. That's now held as **one
|
||||||
be able to promise that a task runs when it's supposed to, within bounded time.
|
interesting constraint to explore, not a requirement** — because real-time is a
|
||||||
That imposes concrete requirements:
|
*pervasive* invariant (every operation must be provably time-bounded, everywhere)
|
||||||
|
that would slow every milestone, whereas resilience is a set of *structural* features
|
||||||
|
that's lighter to build and is what the project actually wants. The trade-off is
|
||||||
|
written up in [smp.md](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience).
|
||||||
|
|
||||||
- **Fixed-priority preemptive scheduling.** The highest-priority ready task always
|
What danos keeps from the real-time direction, because it's cheap and useful anyway:
|
||||||
runs; a higher-priority task that becomes ready preempts a lower one immediately.
|
|
||||||
Not round-robin (which is fair but not predictable).
|
|
||||||
- **A calibrated, deterministic clock.** Guarantees measured in "quanta" are
|
|
||||||
meaningless on an arbitrary tick rate — real time requires a timer calibrated to
|
|
||||||
a known frequency.
|
|
||||||
- **Bounded interrupt latency.** Interrupt-disabled sections must be short and
|
|
||||||
bounded, so a ready high-priority task is never delayed by an unbounded kernel
|
|
||||||
operation.
|
|
||||||
- **Deterministic kernel operations.** Scheduling decisions should be O(1) (e.g. a
|
|
||||||
priority bitmap), not "walk a list of unknown length."
|
|
||||||
- **Priority inheritance** (once there are locks/IPC), so a high-priority task
|
|
||||||
blocked on a resource held by a low-priority one can't be delayed indefinitely by
|
|
||||||
a middle-priority task — bounding priority inversion.
|
|
||||||
|
|
||||||
A consequence worth stating early: the current [kernel heap](heap.md) is a
|
- **Fixed-priority preemptive scheduling** — the highest-priority ready task runs, and
|
||||||
first-fit free list, which has **unbounded allocation time** and can fragment — it
|
preemption lets a runaway component be interrupted and killed (which *serves
|
||||||
is *not* real-time safe. It's fine for one-time kernel setup, but real-time paths
|
resilience*). Already built ([scheduling.md](scheduling.md)).
|
||||||
must pre-allocate or use a bounded (fixed-size pool) allocator. Don't allocate on a
|
- **A calibrated, deterministic clock** — already built ([device-interrupts.md](device-interrupts.md)).
|
||||||
hot real-time path.
|
|
||||||
|
|
||||||
## What this means for the roadmap
|
What danos does *not* owe anyone unless it deliberately chooses real-time later:
|
||||||
|
timing *guarantees*, priority inheritance, bounded allocators, tickless timers, MCS
|
||||||
|
scheduling contexts. Concretely, the current [heap](heap.md) is a first-fit free list
|
||||||
|
with unbounded allocation time — fine here, and only a problem *if* a hard-real-time
|
||||||
|
path is ever added. Note that **QNX is both** a real-time and a restartable
|
||||||
|
microkernel, so choosing resilience now doesn't close the real-time door — it just
|
||||||
|
doesn't pay the tax yet.
|
||||||
|
|
||||||
The vision reorders the obvious hobby-kernel path. Notably, **drivers are not
|
## The roadmap — tracks, not a strict line
|
||||||
built into the kernel** — so an in-kernel keyboard driver would be throwaway work.
|
|
||||||
Input devices arrive later, as the *first user-space drivers*, once the machinery
|
|
||||||
to isolate them exists. The trajectory:
|
|
||||||
|
|
||||||
1. **Calibrated timer / clock** — a known-frequency, deterministic tick. The
|
Because the driver is curiosity plus the win condition, the roadmap is a set of
|
||||||
foundation real-time quanta rest on. *(next)*
|
**tracks** with dependencies, not a rigid sequence. Pick by interest; mind the
|
||||||
2. **Real-time scheduler** — fixed-priority preemptive, kernel threads first:
|
prerequisites.
|
||||||
context switch, task struct, priority run-queue, timer-driven preemption.
|
|
||||||
3. **User mode + address-space isolation** — higher-half kernel, ring 3, per-process
|
|
||||||
page tables. The substrate for isolated servers.
|
|
||||||
4. **IPC** — fast message passing between address spaces. The microkernel's heart.
|
|
||||||
5. **User-space drivers** — interrupts delivered as IPC, plus MMIO/port-access
|
|
||||||
grants. The keyboard becomes the first one, validating the whole model.
|
|
||||||
|
|
||||||
## Where we are
|
**Done:** UEFI boot, framebuffer + [serial](testing.md), [physical frames](frame-allocator.md)
|
||||||
|
(with boot-services memory reclaimed), [paging](paging.md) with W^X, [exceptions and
|
||||||
|
interrupts](interrupts.md), a [calibrated timer + ns clock](device-interrupts.md), a
|
||||||
|
[heap](heap.md), a [fixed-priority preemptive scheduler](scheduling.md) with blocking,
|
||||||
|
and in-kernel [IPC channels](ipc.md) — plus a [test harness](testing.md).
|
||||||
|
|
||||||
The foundation is in place: UEFI boot, framebuffer + [serial](testing.md),
|
- **Isolation track** — **user mode + address-space isolation** (higher-half kernel,
|
||||||
[physical frames](frame-allocator.md), [paging](paging.md) with W^X, [exceptions
|
ring 3, per-process page tables). The substrate everything else needs. *Next, and a
|
||||||
and interrupts](interrupts.md), a [timer](device-interrupts.md), and a
|
prerequisite for the resilience and driver tracks.*
|
||||||
[heap](heap.md) — plus a [test harness](testing.md). The kernel boots and has its
|
- **Resilience track** — fault → kill → notify, a supervisor/reincarnation server,
|
||||||
core services; the next milestones make it *schedule*, then *isolate*.
|
resource cleanup on death, then a restartable driver as proof. Needs isolation.
|
||||||
|
See [resilience.md](resilience.md).
|
||||||
|
- **ARM track** — the `aarch64` port so danos runs on the Zero 2 W and Pi 5. Largely
|
||||||
|
independent of the others (it's the [arch layer](arch.md)); directly serves the win
|
||||||
|
condition. Likely via aarch64-UEFI first (QEMU `virt` + AAVMF), then real boards.
|
||||||
|
See [arm.md](arm.md), and [discovery.md](discovery.md) for the device tree it needs.
|
||||||
|
- **GUI track** — a framebuffer-based windowing/compositor, and the input + display
|
||||||
|
drivers under it. Builds on the neutral framebuffer (so it's arch-independent), and
|
||||||
|
on the driver model from the isolation/resilience tracks. The visible payoff.
|
||||||
|
|
||||||
|
The natural spine is **isolation → (resilience + drivers) → GUI**, with the **ARM
|
||||||
|
track** pursued alongside whenever the itch to see it boot on a Pi wins out.
|
||||||
|
|
||||||
|
## How to use this page
|
||||||
|
|
||||||
|
Read it before adding anything structural. When a design decision comes up, the
|
||||||
|
question is: does it serve the **win condition** (runs on the three machines, with a
|
||||||
|
GUI), or the **learning** (a constraint worth meeting)? If it serves neither — e.g.
|
||||||
|
paying the full real-time tax with no payoff in sight — it can wait.
|
||||||
|
|||||||
Reference in New Issue
Block a user