re-org docs

This commit is contained in:
Daniel Samson
2026-07-23 00:24:01 +01:00
parent f023f1cfd6
commit 52d6e372fd
53 changed files with 271 additions and 389 deletions
+9
View File
@@ -0,0 +1,9 @@
# OS Developer Guide
This document is for those who need to understand the architectural decisions behind the OS.
## Written in Zig?
The os was initially written in zig because it has excellent support for EFI. With zig, we could forgo using a third party bootloader, reducing the time to boot up the kernel. Following the "Zen of Zig", helped to produce the most readable codebase for an operating system ever created. So those, new to OS development could quickly get up to speed.
+178
View File
@@ -0,0 +1,178 @@
# ACPI: finding the tables (RSDP → RSDT/XSDT → SDTs)
ACPI describes the hardware the kernel can't assume — the interrupt controllers, the
PCIe config window, the timer, the power registers — in a set of **system
description tables** (SDTs). But before it can read any of them, danos has to *find*
them, and they aren't at a fixed address. Getting there is a short chain of pointers,
and this note explains it — in particular the question it's easy to trip on: **how
does the [platform / device module](architecture.md) know where the RSDT is?**
Short answer: it doesn't receive the RSDT. The firmware hands over the **RSDP**, and
the RSDT's address is a field *inside* the RSDP. The platform follows that pointer.
## The locator chain
```
UEFI configuration table
│ the loader reads the RSDP's physical address
▼
BootInformation.acpi_rsdp (u64, in the loader↔kernel handoff) system/boot-handoff.zig
│ the kernel forwards the whole BootInformation
▼
platform.discover(boot_information, …) system/kernel/platform.zig
│ reads boot_information.acpi_rsdp, hands it to the ACPI backend
▼
acpi.discover(rsdp_phys, …) system/kernel/acpi.zig
│ dereferences the RSDP, reads the pointer it contains
▼
RSDP ──(a field in the struct)──► RSDT / XSDT ──► SDTs (MADT, MCFG, FADT, HPET, DSDT…)
```
The **RSDP** (Root System Description Pointer) is the root of the whole ACPI tree.
Its only job is to point at the root *table* — the **RSDT** (ACPI 1.0) or its 64-bit
successor the **XSDT** (ACPI 2.0+) — which in turn lists every other SDT.
## Step 1 — the loader finds the RSDP
Only the firmware knows where ACPI lives, so the RSDP must be grabbed while UEFI is
still up. `acpiRootSystemDescriptorPointer()` in `boot/efi.zig` walks the UEFI
**configuration table** for the ACPI GUID and returns the vendor pointer — the same
"grab it before `ExitBootServices`" pattern as the [framebuffer](framebuffer.md) and
the [memory map](memory-map.md).
## Step 2 — the handoff: a physical address in `BootInformation`
The loader can't just call the device module: the bootloader binary and the kernel
binary are compiled separately, and **the loader isn't linked against the `platform`
module at all** (it imports only the `boot-handoff` contract and the
`initial-ramdisk` module). So instead of a call, it
deposits a value in the handoff struct:
```zig
// boot/efi.zig — while boot services are still up
.acpi_rsdp = if (acpiRootSystemDescriptorPointer()) |p| @intFromPtr(p) else 0,
```
Two things about what crosses the boundary:
- **It's a *physical* address, not a Zig pointer.** The loader and kernel don't share
an address space at the moment of the jump, so a raw `u64` physical address is the
only thing that survives the handoff. `BootInformation.acpi_rsdp` is `0` when the firmware
exposed no ACPI (e.g. a future device-tree machine, which would fill a different
field instead — the kernel never learns which firmware booted it).
- **The kernel can dereference it through the physmap.** The RSDP lives in
ACPI-reclaim memory, which [paging.zig](paging.md) maps — along with the rest of
RAM — into the higher-half **physmap** (there is no identity mapping; the low half
belongs to user space). So by the time discovery runs,
`@ptrFromInt(physicalToVirtual(rsdp_phys))` is a valid pointer.
This is the concrete form of the "capture the description pointer" step sketched in
[discovery.md](discovery.md) — a plain `acpi_rsdp: u64` rather than a tagged handle,
since x86 is the only backend wired up so far.
## Step 3 — the platform derives the RSDT from the RSDP
`acpi.discover` reinterprets the physical address as the RSDP struct, validates it,
and then reads the root-table pointer *out of it*. Which pointer depends on the ACPI
version, because the RSDP carries **both**:
```zig
const rsdp: *const RootSystemDescriptionPointer = @ptrFromInt(physicalToVirtual(rsdp_phys));
if (!std.mem.eql(u8, &rsdp.signature, "RSD PTR ")) return error.BadRsdpSignature;
if (!checksumOk(@ptrFromInt(physicalToVirtual(rsdp_phys)), 20)) return error.BadRsdpChecksum;
if (rsdp.revision >= 2) {
// ACPI 2.0+: use the 64-bit XSDT pointer (the 32-bit RSDT is deprecated)
const xsdp: *const ExtendedSystemDescriptorPointer = @ptrFromInt(physicalToVirtual(rsdp_phys));
try walkRoot(u64, xsdp.extended_system_descriptor_table_address, …);
} else {
// ACPI 1.0: use the 32-bit RSDT pointer
try walkRoot(u32, rsdp.root_system_description_table_address, …);
}
```
- The **`revision`** byte selects the root table. `root_system_description_table_address`
(32-bit, → RSDT) and `extended_system_descriptor_table_address` (64-bit, → XSDT) are
ordinary fields of the RSDP/XSDP structs — the platform never *receives* the RSDT
address, it *reads* it here. On QEMU q35 the RSDP is revision 2, so the XSDT path is
taken.
- The signature (`"RSD PTR "`) and one-byte checksum guard against a bad pointer before
anything downstream trusts it.
## After the root table
`walkRoot` treats the RSDT/XSDT as an array of physical pointers — 32-bit entries for
the RSDT, 64-bit for the XSDT — and hands each SDT to `handleTable`, which dispatches
on its 4-byte signature: **MADT** (CPUs + IOAPIC), **MCFG** (PCIe ECAM), **FADT**
(power registers, and the pointer to the DSDT), **HPET** (timer). That's where the
firmware-agnostic [device model](discovery.md) gets populated; this note stops at the
part that answers "where are the tables?" — everything past the RSDP is just following
more pointers the tables themselves provide.
## ACPI events: the SCI, the power button, and GPEs (M21)
The tables above are static description; ACPI is also a *live* channel. Hardware
raises the **SCI** (System Control Interrupt) — one shared, level-triggered line
whose vector the FADT names — and the OS reads status registers to learn what
happened: a fixed event like the power button, or a **General-Purpose Event**
(GPE) whose handler is an AML method. Since [discovery](discovery.md) moved AML
to ring 3, the event side lives there too, in the same **acpi service** — the
device discoverer and the event source are one process, because both need the
namespace and the port grant.
**The kernel hands the service what it needs and no more.** Reading PM1 event
blocks and GPE blocks requires the FADT, which the kernel already parses for its
own power register map (feeding reboot), the PM timer, and the SCI line — the
kernel itself has no S5/poweroff path. Rather than re-parse, the kernel appends the **FADT as one
more memory resource** on the `acpi-tables` node; the service tells it apart
from the AML blob resources by signature — the FADT keeps its intact `"FACP"`
header, while the blob resources are header-stripped bytecode that starts with
no signature. The kernel's own FADT parse is untouched; the service reads the
PM1 *event* blocks (which the kernel never parsed — it extracts only the PM1
*control* register, and it is the service, not the kernel, that writes it for
`\_S5`) and the GPE0/GPE1 blocks straight from its copy. The **SCI itself**
arrives as the node's one `len == 1` irq resource (distinct from the broad
`[0, 256)` window that covers children's legacy lines), which is how the service
finds the line to `irq_bind`.
With those in hand the service enables ACPI mode (only if `SCI_EN` is clear —
some firmwares boot with it already set), sets `PWRBTN_EN`, and on each SCI:
- **The power button** is a *fixed* event: a set `PWRBTN_STS` bit in PM1 status.
The handler clears it (write-1-to-clear), logs the press, and publishes a
[`power`](power.md) `power_button` event to subscribers.
- **GPEs** are the general path: for each set-and-enabled GPE bit `n`, the
service evaluates its `\_GPE._L%02X` (level) or `_E%02X` (edge) handler
method, drains the **Notify** queue that method produced, maps each notified
device to an event (battery, AC, lid, or a generic `notify` with its code),
and clears the status bit. A missing handler method is not an error: the
status bit is cleared and the event silently dropped. Making GPEs work
required teaching the interpreter one opcode it never
handled — `Notify` (`0x86`) — which it now folds into a bounded queue drained
per evaluation; everything else a handler needs (field access, control flow,
method calls) was already proven by the ring-3 `_STA`/`_CRS` work.
**How this is tested.** QEMU cannot raise GPEs deterministically on this config,
so GPE/Notify correctness is proven by **host unit tests** — hand-encoded AML
with a `Notify` inside a method body, run under `zig build test`. The QEMU
`power-button` scenario proves the fixed-event path end to end: a QMP
`system_powerdown` injects a real ACPI power-button press, and the service's SCI
handler must log it. Battery/AC/lid mapping is interface-complete but validated
on real hardware later; the embedded controller's `_Qxx` queries are out of
scope.
The service surface these events are *published on* — subscription, the event
vocabulary, and orderly shutdown — is the power service, [power.md](power.md).
## Related
- [efi.md](efi.md) — the loader that captures the RSDP before `ExitBootServices`.
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam, and
the ACPI-reclaim memory the RSDP lives in.
- [discovery.md](discovery.md) — the broader (still-evolving) plan for turning these
tables into one neutral device model shared with the ARM device-tree path, and how
ACPI enumeration and events moved to the ring-3 acpi service.
- [power.md](power.md) — the domain-named power service the ACPI event side publishes
to (button, lid, battery) and its orderly-shutdown path into S5.
- [architecture.md](architecture.md) — why the kernel reaches the device code through a `platform`
module and never names ACPI directly.
+109
View File
@@ -0,0 +1,109 @@
# Architecture split
danos targets x86_64 today, but is meant to grow onto other systems later — a
Raspberry Pi, say, which is AArch64 and has no UEFI. To keep that possible without
a rewrite, CPU-specific kernel code lives behind a boundary: the generic kernel
never names an architecture, and each architecture plugs in behind it.
## The seam is a build-time module named `architecture`
The mechanism is deliberately boring — no vtables, no function-pointer tables, no
runtime dispatch. `build.zig` exposes one architecture's code as a module called
`architecture`:
```zig
const architecture_module = b.addModule("architecture", .{
.root_source_file = b.path("system/kernel/architecture/x86_64/cpu.zig"),
});
```
and the generic kernel imports it by that name:
```zig
const architecture = @import("architecture");
// ...
architecture.halt(); // never says "x86_64"
```
Adding a second architecture is then a build-time choice: create
`system/kernel/architecture/aarch64/`, and point the `architecture` module at it when the target CPU is
AArch64. `kernel.zig` and `console.zig` don't change. **That compiler-checked module
boundary _is_ the architecture interface** — when a new architecture is missing a function
the generic kernel calls, the build fails and names exactly what's missing.
## What's arch-specific vs generic
The split follows a simple test: does it name a CPU instruction, a hardware
register, or a memory-management structure? If so, it's arch-specific.
| Arch-specific — `system/kernel/architecture/x86_64/` | Generic — kernel core |
|---|---|
| `cpu.zig`: CPU state, trap-frame accessors, paging, SMP | `console.zig` — pure pixel math, framebuffer drawing |
| `gdt.zig`, `idt.zig`, `tss.zig` — descriptor tables | `kernel.zig` — kernel orchestration, scheduler, IPC |
| `paging.zig` — page-table setup and management | `process.zig` — process lifecycle, address spaces |
| `apic.zig`, `ioapic.zig` — interrupt controllers | `scheduler.zig` — task scheduling and context switch |
| `serial.zig`, `io.zig` — UART, I/O primitives | `vfs.zig` — filesystem abstraction |
| `isr.s`, `smp.zig` — exceptions, AP bring-up, context switch | `irq.zig`, `ipc*.zig` — interrupt dispatch, messaging |
| `linker.ld` — kernel link layout, load address | |
Notice the framebuffer console is *generic*: it just writes pixels into whatever
framebuffer it's handed, so it needs no per-arch version. Most of the kernel
should end up on the generic side; the architecture module stays small.
## Two axes, kept separate
There are really two independent questions, and it's worth not conflating them:
- **CPU architecture** (x86_64 vs AArch64): instructions, MMU, interrupts →
`system/kernel/architecture/<cpu>/`.
- **Boot protocol** (UEFI vs Raspberry Pi firmware + device tree): handled
*separately*, because loaders are their own binaries. `boot/efi.zig` builds
`BOOTX64.efi`, a distinct executable from the kernel ELF. On a Pi there is no
separate loader at all — the firmware jumps straight into the kernel with a
device-tree pointer, so that entry work would live in the AArch64 architecture code.
Either path converges on the same neutral [`BootInformation`](memory-map.md).
## Current x86_64 contents
- **`system/kernel/architecture/x86_64/cpu.zig`** — the `architecture` module root. Exposes the trap-frame
`CpuState` and accessors, `init()` (bring up the descriptor tables), `enablePaging()`,
`enterUser()`/`userExit()` for ring-0 ↔ ring-3 transitions, address-space management, and SMP
entry points (see [halting.md](halting.md), [interrupts.md](interrupts.md), [paging.md](paging.md),
[scheduling.md](scheduling.md)).
- **`system/kernel/architecture/x86_64/gdt.zig`** / **`idt.zig`** / **`tss.zig`** — the GDT, IDT and
TSS plus CPU-exception handling (see [interrupts.md](interrupts.md)).
- **`system/kernel/architecture/x86_64/paging.zig`** — the kernel's page tables and address-space
management (see [paging.md](paging.md)).
- **`system/kernel/architecture/x86_64/apic.zig`** / **`ioapic.zig`** — the Local APIC, its timer,
and the I/O APIC for device interrupts (see [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)).
- **`system/kernel/architecture/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's
machine-readable log channel, see [testing.md](../testing.md)) and the shared port-I/O + MSR primitives.
- **`system/kernel/architecture/x86_64/smp.zig`** / **`per-cpu.zig`** — application-processor bring-up
and per-CPU state (GS base, system-call entry point, see [scheduling.md](scheduling.md)).
- **`system/kernel/architecture/x86_64/isr.s`** — the exception stubs, the `lgdt`/`lidt`/`ltr` load
helpers, ring-0 ↔ ring-3 transitions, and the context switch — real assembly, since Zig inline asm can't
express them (see [scheduling.md](scheduling.md)).
- **`system/kernel/architecture/x86_64/linker.ld`** — the kernel link layout (fixed low load
address, one PT_LOAD per permission set).
The kernel entry point `_start` lives in the architecture-specific `isr.s` (x86_64 here).
On x86_64 it sets up the kernel stack in BSS and jumps to `kmain()` in `kernel.zig`.
This is already per-architecture — an AArch64 port would have its own `isr.s` entry
that parses the device-tree pointer from a register and jumps to the same `kmain()`.
The entry interface is minimal and emerges naturally from the [boot-handoff](memory-map.md)
contract both share.
## The discipline
The thing that makes this help rather than hurt: **only extract what's provably
architecture-specific, and let the interface emerge with the second
implementation.** With a single architecture you're guessing at the seam, and a
wrong guess encoded as elaborate abstraction is expensive to undo. So:
- Move code into `architecture/` only when it genuinely names CPU-specific machinery.
- Grow the `architecture` surface one function at a time, as steps need it.
- Don't pre-design the interrupt or paging interfaces before writing them.
Directory hygiene is cheap and reversible; premature abstraction is neither. When
architecture #2 lands and something doesn't fit, reshaping a few hundred lines is nothing —
unwinding an abstraction empire is not.
+125
View File
@@ -0,0 +1,125 @@
# ARM targets (`arm` and `aarch64`)
danos aims to run on Raspberry Pi hardware eventually. "ARM" isn't one target,
though — the Pis span **two different CPU architectures** (32-bit `arm` and 64-bit
`aarch64`) and (stock) a different boot protocol from x86-64's UEFI. **danos targets
`aarch64` only** (see the decision below); the `arm`/`aarch64` distinction still
matters for understanding why. This page maps the landscape so the
[architecture split](architecture.md) and build system can be planned for it.
## `arm` vs `aarch64` — 32-bit vs 64-bit
- **`arm`** = **32-bit** ARM (the *AArch32* state, A32/T32 instruction sets).
ARMv7 and earlier, plus the 32-bit compatibility mode of newer cores. 16 × 32-bit
registers.
- **`aarch64`** = **64-bit** ARM (the *AArch64* state, A64 instruction set), from
**ARMv8-A** on. Also called **arm64**. 31 × 64-bit registers, a fixed 32-bit
instruction width, a redesigned exception model — *not* a widening of A32, a clean
new ISA.
They are as different from each other as either is from x86-64: separate registers,
page-table formats, and calling conventions. Each needs its own `system/kernel/architecture/<name>/`.
## The Raspberry Pi models
| Model | SoC | Core | Architecture | danos target |
|-------|-----|------|--------------|--------------|
| Pi Zero / Zero W | BCM2835 | ARM1176JZF-S | ARMv6, 32-bit only | `arm` (not planned) |
| **Pi Zero 2 W** ← target | BCM2710 | Cortex-A53 | ARMv8-A, 64-bit | **`aarch64`** |
| **Pi 3 / 3B+** | BCM2837 | Cortex-A53 | ARMv8-A, 64-bit | **`aarch64`** |
| **Pi 4** | BCM2711 | Cortex-A72 | ARMv8-A, 64-bit | **`aarch64`** |
| **Pi 5** | BCM2712 | Cortex-A76 | ARMv8.2-A, 64-bit | **`aarch64`** |
> **Decision: `aarch64` only.** The target small board is a **Pi Zero 2 W** (BCM2710,
> Cortex-A53) — which is **`aarch64`**, *not* the original Zero W's 32-bit ARMv6. So
> every ARM board danos targets (Zero 2 W and Pi 3-5) is `aarch64`, and the 32-bit
> `arm`/ARMv6 backend is **not planned** — one ARM CPU port, not two. The original
> Zero W (ARMv6) would only re-enter scope if that specific older board were ever
> needed; the row above is kept only to explain the distinction.
## Booting: UEFI is not x86-only
The boot protocol is a **separate axis** from the CPU (see [architecture.md](architecture.md)):
- **UEFI** exists for ARM too — ARM servers require it (SBSA/SBBR), QEMU boots it
with **AAVMF** (the AArch64 build of the same EDK2 firmware as x86's OVMF), and
the Pi can even run it with community UEFI firmware. Under UEFI the handoff is the
*same* as x86-64: system table, boot services, memory map, GOP framebuffer — so
the loader logic largely carries over.
- **Device tree / firmware boot** — the **stock** Raspberry Pi firmware (VideoCore
bootloader) is *not* UEFI: it loads the kernel and jumps to it with a **device-tree
blob (DTB)** pointer. Both the Zero W and stock Pi 3-5 boot this way.
Note that even under UEFI on ARM, the OS still gets its hardware description from
**ACPI or a device tree** (often the DTB passed via a UEFI configuration table). So
"UEFI on ARM" doesn't remove the device tree — UEFI gives you memory + framebuffer;
the DTB/ACPI tells you what devices exist.
## What danos needs, layer by layer
- **One CPU arch module: `system/kernel/architecture/aarch64/`** — covering the Zero 2 W and Pi 3-5,
providing the same `arch` interface as x86_64: `halt`, context switch,
interrupt/exception vectors, page tables, a UART, a timer. No `system/kernel/architecture/arm/` is
planned (see the decision above), so there's a single ARM backend to write.
- **A device-tree boot path.** Since stock Pis boot via DTB, danos needs an entry
that parses the DTB's `/memory` and `/reserved-memory` into the neutral
[`MemoryMap`](memory-map.md) — the same neutral handoff `efi.zig` produces, just
from a different source. This is where keeping boot-protocol knowledge on the
loader side (as we did for the UEFI memory-map classification) pays off.
- **The UEFI loader mostly carries over.** `boot/efi.zig` is largely
boot-*protocol* code (`std.os.uefi` protocol calls), not x86 code. Its truly
x86-specific bits are the ELF machine check (`.X86_64`), the SysV calling
convention for the kernel jump, the `hlt` park on failure — and, the substantial
one, the bootstrap page tables: `buildBootstrapTables` builds x86-64 4-level
tables (PML4/PDPT/PD index shifts, x86 PTE bits, 2 MiB leaves) and hands the
kernel a CR3. Page-table formats are per-architecture (see above), so an
`aarch64` loader keeps the protocol code but rewrites that builder in the
aarch64 translation-table format. Even so, an `aarch64`-UEFI target (QEMU
`virt` + AAVMF) reuses most of it — which makes **aarch64-UEFI the easiest
second target**, easier than the device-tree Pi.
## Pi hardware quirks (for when we port)
The Pi is not a "standard" ARM platform — expect Broadcom-specific peripherals:
- **Peripheral base moves per SoC**: `0x2000_0000` (BCM2835, Zero W),
`0x3F00_0000` (BCM2837, Pi 3), `0xFE00_0000` (BCM2711, Pi 4), different again on
Pi 5. Everything below is an offset from it.
- **UART**: a **PL011** (at base + `0x20_1000`) plus a mini-UART; on some boards the
PL011 is wired to Bluetooth, so which one is the console varies. This is the
`aarch64`/`arm` equivalent of our x86 [COM1 serial](../testing.md).
- **Interrupt controller**: *not* a standard ARM GIC on the older parts — the Zero W
and Pi 3 use Broadcom's own ARMCTRL controller (Pi 3 adds a per-core "local"
controller for timers/mailboxes). The **Pi 4 and 5 do have a GIC-400**. So the
interrupt backend differs even within the `aarch64` Pis.
- **Timer**: the ARM generic timer (`CNTPCT`/`CNTFRQ`) on ARMv8, or the BCM system
timer — the counterpart to our calibrated LAPIC/TSC clock.
## Building each (intended)
`build.zig` currently pins the kernel to `x86_64`; supporting these means selecting
the target and `arch` module together (e.g. a `-Darch=` option). The Zig target
queries would be roughly:
- **Pi Zero W**: `.cpu_arch = .arm`, `.cpu_model = arm1176jzf_s`, `.os_tag = .freestanding`
- **Pi 3**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a53`, `.os_tag = .freestanding`
- **Pi 4**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a72`
- **Pi 5**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a76`
## Testing in QEMU
Two routes, mirroring how we test x86-64 with OVMF:
- **Board emulation**: `qemu-system-aarch64 -machine raspi3b` (and `raspi4b` on
recent QEMU) for the Pi 3/4; `qemu-system-arm -machine raspi0`/`raspi1ap` for the
ARMv6 Zero-class board — closest to real hardware, device-tree boot.
- **Generic aarch64-UEFI**: `qemu-system-aarch64 -machine virt` + AAVMF — the
cleanest way to bring up the `aarch64` kernel via the reused UEFI loader before
tackling Pi-specific boards. A future `run-aarch64` build step would use this.
## Related
- [architecture.md](architecture.md) — the arch-module boundary these targets plug into, and the
CPU-arch vs boot-protocol "two axes".
- [efi.md](efi.md) — the UEFI loader that carries over to aarch64-UEFI.
- [vision.md](../vision.md) — why isolated, portable-across-architectures is the goal.
+259
View File
@@ -0,0 +1,259 @@
# Device discovery (ACPI / device tree), the agnostic way
"Device discovery" is how the kernel learns **what hardware exists and where** — the
MMIO addresses, IRQ numbers, CPU count, and interrupt controller it can't just
assume. On x86 that description comes from **ACPI** tables; on ARM from a **device
tree** (DTB). This note is a design plan, not built yet: *when* danos should tackle
it, and *how* to keep it architecture-agnostic — the same discipline the
[memory map](memory-map.md) and [architecture split](architecture.md) already follow.
## What the kernel assumed when this plan was written
At the time danos discovered almost nothing — it coasted on legacy PC fixtures that
are guaranteed to exist under QEMU + UEFI:
- `system/kernel/architecture/x86_64/apic.zig` assumed the **Local APIC** at the
default `0xFEE0_0000` and calibrated its timer against the **PIT** (the legacy
8254). (Today the PIT is the *last-resort* reference: calibration prefers the
CPUID-reported TSC frequency, then the HPET, then the ACPI PM timer.)
- `system/kernel/architecture/x86_64/serial.zig` hardcoded **COM1** at I/O port
`0x3F8`. (Today `0x3F8` is only the default: the kernel loopback-probes the UART
and parses ACPI's SPCR table to target the firmware's actual debug port.)
- The framebuffer and memory map come from **UEFI** — that *is* discovery, just done
by the firmware and handed over, not read from ACPI.
This worked only because PC-compatible hardware promises those legacy pieces exist at
those addresses. It was a crutch, and it did not travel.
## The forcing functions: when to build it
Two things drive the need, and they set the timing:
1. **The second architecture makes it mandatory.** ARM has *no* legacy fixtures —
no PIT, no fixed serial port, no standard interrupt controller address. You can't
find the UART to print a character without reading the device tree. So on x86 we
can defer discovery a long time (until we want the IOAPIC, PCIe, or SMP), but on
**aarch64 it's required to boot at all**. The [aarch64 port](arm.md) is what
forces the issue.
2. **Isolated user-space drivers need it.** In the [microkernel vision](../vision.md),
drivers live in user space — but something has to enumerate the hardware and hand
each driver its MMIO regions and IRQs. That enumeration *is* device discovery. So
discovery is a prerequisite for real drivers, **not** for user mode itself.
The conclusion on timing: **don't build full discovery before user mode.** User mode
+ address-space isolation needs none of it; the current assumptions are fine there.
Build the agnostic discovery layer **when the aarch64 port starts** — because that's
when a second, real implementation makes "agnostic" honest.
## Why build it *with* the second arch, not before
The same lesson as the arch split: an abstraction with only one implementation
quietly bends to that implementation. Build "agnostic discovery" x86-only first and
you'll get an ACPI-shaped interface with device tree bolted on afterward. Build it
when aarch64 lands and the small, clean DTB parser pulls the abstraction toward the
right neutral shape, which the x86 side then fills. Design it against two backends or
it isn't really agnostic.
## The one cheap step to take sooner
Have the **loader capture the description pointer** into `BootInformation` — a
neutral handle, no parsing:
```zig
pub const HardwareInfo = extern struct {
kind: enum(u32) { none, acpi, device_tree },
addr: u64, // ACPI RSDP, or the DTB blob
};
```
On x86-UEFI that's the ACPI **RSDP**, read from the UEFI configuration table *before*
`ExitBootServices` — the same "grab it before exit" pattern as the framebuffer and
memory map (already flagged in [memory-map.md](memory-map.md)). On ARM it's the DTB
pointer the firmware passes. This keeps the door open for near-zero cost without
committing to the parser.
## What "agnostic" looks like
The happy accident: **the device-tree data model is already a good neutral
representation.** A DTB is a tree of nodes, each with:
- a **`compatible`** string — what the device is,
- **`reg`** — its MMIO base(s) and size(s),
- **`interrupts`** — its IRQ number(s),
- other properties.
Even OSes running on ACPI hardware normalize into a unified device model shaped like
this. So the neutral layer is a **device model** — "here are the devices, each with a
type, MMIO regions, and IRQs" — fed by two backends behind it:
- a **DTB parser** (ARM) — a few hundred lines against a well-specified binary format;
- **static ACPI + PCI enumeration** (x86).
### The caveat that sets the effort: "ACPI" ≠ "AML interpreter"
A full ACPI namespace is **AML** (ACPI Machine Language) bytecode, and writing an AML
interpreter is an enormous undertaking. **We skip it.** Interrupt routing and PCIe
come entirely from the *static* tables:
- **MADT** — the APICs and the **IOAPIC** (interrupt routing),
- **MCFG** — PCIe configuration space (then walk the PCI bus to enumerate devices),
- **FADT**, **HPET** — power/reset and a precise timer.
Walking PCI config space finds most devices without any AML. So the x86 static-ACPI
backend is roughly comparable in scope to the DTB parser; it's full AML that's the
monster, and it isn't on the path.
## Where it sits in the seam
Discovery layers onto the existing loader↔kernel split cleanly:
| Layer | Responsibility | x86 | ARM |
|-------|----------------|-----|-----|
| **Capture** (loader) | grab the description pointer | RSDP from UEFI config table | DTB pointer from firmware |
| **Parse** (kernel, per-mechanism) | pointer → neutral device model | static ACPI + PCI | DTB parser |
| **Consume** (generic) | use the device model | — same code — | — same code — |
Boot-protocol-specific capture, mechanism-specific parse, generic consumption —
exactly like the memory map, where the loader classifies and the kernel just sees
neutral regions.
## Where it lives in a microkernel
For isolated drivers, discovery is not one lump — it splits by *who needs it and
when*:
- **Minimal, in-kernel: the interrupt controller and the timer.** The IOAPIC/GIC and
the timer are needed *before* user space exists (scheduling and preemption depend
on them), so the kernel must parse at least these from ACPI/DTB itself. This is also
the **first real consumer** of discovery — the first genuine reason to read MADT or
the DTB is "where is the interrupt controller and how do I route an IRQ?" It's the
point where x86 finally graduates from the PIT/legacy-LAPIC assumptions.
- **User-space enumeration: a device-manager server.** Everything else — PCI devices,
peripherals — is parsed (or queried from the kernel's parse) by a privileged
user-space server that hands each driver process its MMIO regions and IRQ rights
over [IPC](../device-driver-development-guide/ipc.md). Combined with **interrupts-as-messages** (an IRQ delivered to a
driver as a message on a channel — a natural extension of the wait queues and
channels already built), that's what makes drivers genuinely isolated.
So the device-manager server depends on user mode + IPC; only the interrupt-controller
slice is unavoidably in-kernel.
## A nice ARM contrast
On ARMv8 the generic timer exposes its frequency directly via the `CNTFRQ` register —
no calibration needed. That's cleaner than the x86 side, where we measure the LAPIC
and TSC against the PIT because nothing tells us their frequency (see
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). Discovery on ARM hands you more for
free; discovery on x86 is partly about *finding* what ARM just tells you.
## Suggested ordering
1. **Now (cheap):** plumb the neutral `HardwareInfo` pointer through the loader into
`BootInformation`. No parser yet.
2. **Next milestone unchanged:** user mode + address-space isolation — needs no
discovery.
3. **With the aarch64 port:** build the agnostic discovery layer, **DTB first** (the
forcing function), then x86 static-ACPI to fill the same model — starting with the
**interrupt controller + timer**.
4. **Then:** a user-space device-manager server + interrupts-as-messages → real
isolated drivers (keyboard first).
## Related
- [acpi.md](acpi.md) — the built x86 side of the "capture" step: how the loader grabs
the RSDP and the platform follows it to the RSDT/XSDT and the SDTs.
- [arm.md](arm.md) — the aarch64 target that forces genuine discovery (DTB, GIC).
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam,
and the note about grabbing the RSDP before `ExitBootServices`.
- [device-interrupts.md](../device-driver-development-guide/device-interrupts.md) — the LAPIC/timer bring-up that
discovery will eventually feed (IOAPIC, real IRQ routing).
- [ipc.md](../device-driver-development-guide/ipc.md) — the channels that interrupts-as-messages and the device manager
will ride on.
- [vision.md](../vision.md) — why drivers belong in isolated user space at all.
## Update (M19.3, 2026-07-13): PCI enumeration left the kernel
The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO
apertures derived from the memory map's holes, bus range, and the 16-bit I/O
window). The per-function walk moved to the ring-3 `pci-bus` driver
([device-manager.md](../device-driver-development-guide/device-manager.md)): it claims the bridge, repeats the
ECAM scan through its mmio grant, and `device_register`s what it finds, which
the device manager mirrors and matches. The ACPI namespace walk follows in M20;
the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side.
## Update (M20.3, 2026-07-13): ACPI enumeration left the kernel too
The kernel no longer folds the AML namespace's Device objects into the device
tree. It still parses the *static* tables (MADT for SMP, HPET for the tick, MCFG
for the host bridge, FADT); at this point it also still built the AML namespace —
but only to read the `\\_S5` sleep type for poweroff. (That remnant is gone too:
the kernel now runs no AML at all — soft-off belongs to the acpi service, and the
kernel keeps only the AML-free reboot path.) Device discovery is the ring-3 **acpi
service** ([device-manager.md](../device-driver-development-guide/device-manager.md)): it claims the `acpi-tables`
node the kernel publishes (the AML blobs, a broad io_port grant, the SCI),
re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`,
and registers + reports each `_HID` device — the device manager matches drivers
(ps2-bus) from those reports. With M19's pci-bus driver, discovery now runs
entirely in user space, anchored on two kernel-seeded nodes: the host bridge and
the acpi-tables node. (The kernel's static-table parse also seeds the processor,
interrupt-controller, and HPET timer nodes, and it publishes the boot
framebuffer as a claimable display node — but no *enumeration* happens in
ring 0.)
## Discovery is a swappable process per firmware (M19–M20)
Moving PCI and ACPI enumeration out of ring 0 was not just a relocation — it
made discovery **firmware-neutral by construction**, which is the whole reason
to do it before the second architecture rather than after. Everything at and
above the [device-manager](../device-driver-development-guide/device-manager.md) protocol — descriptors,
containment, reports, matching, supervision — is generic and may never become
x86-specific. Discovery is the single firmware-specific piece, and it is
isolated as **one swappable process per firmware**:
- **x86** boots describe hardware with ACPI, so the discoverer is the **acpi
service** ([acpi.md](acpi.md)): it claims the `acpi-tables` node and runs AML.
- **The Raspberry Pis** hand over a flattened device tree, so the discoverer is
an **fdt service**: it claims a `devicetree-blob` node and walks the tree —
pure data, no bytecode, so it needs neither a port grant nor an interpreter,
strictly simpler than ACPI. (A placeholder until the [aarch64](arm.md)
bring-up fills it in.)
The device manager spawns the discoverer under the **neutral ramdisk name
`discovery`** and never learns which firmware it is on; the build's
`-Ddiscovery=acpi|fdt` option fills that slot (x86 defaults to `acpi`, the
aarch64 target flips the default when it lands). The manager owns the device
tree as *data* and touches no hardware, ever — firmware bytecode runs only
inside the crashable, supervised discoverer, so an AML fault can never take
down the supervisor.
Two consequences of neutrality bind on later work:
- **Cross-firmware surfaces are named by domain, not firmware.** System power is
a [`power`](power.md) protocol, not an "ACPI events" protocol: on x86 the acpi
service registers it, on ARM a PSCI/mailbox service registers the same
`ServiceId.power`, and subscribers never learn the difference.
- **Identity must widen before the fdt service exists.** `DeviceDescriptor`'s
8-byte `hid` holds an EISA id but cannot hold an FDT `compatible` string
(`"brcm,bcm2835-aux-uart"`); the identity field grows before the ARM path can
report a real node.
Two supporting decisions keep the kernel's remaining slice honest:
- **The AML interpreter is a single build module**
(`library/device/acpi/aml/aml.zig`) — one source, no fork. During the ring-3 move
it was compiled into both the kernel (which linked it just for the `\_S5`
poweroff evaluation) and the acpi service, with the `acpi-parse` test
asserting the two produce the same device count. Since soft-off followed
discovery out of the kernel, only the acpi service links the module — the
kernel runs no AML — and the test now asserts a device-count *floor* for the
ring-3 parse instead, there being no kernel count left to equal.
- **Bridge apertures come from the firmware memory map, not AML.** Registered
PCI functions carry BAR resources, and `device_register` containment demands
the bridge own windows that cover them. Those apertures are derived
kernel-side from the boot memory map's MMIO holes (regions that are neither
RAM nor tables) — mechanical, AML-free, and available at boot regardless of
what later moved to user space. The acpi service's authority is likewise
exactly one node: the `acpi-tables` node, whose broad io_port grant is the
documented trust boundary for the one process allowed to run firmware
bytecode.
+211
View File
@@ -0,0 +1,211 @@
# EFI / The Boot Process
## What EFI is
**UEFI** (Unified Extensible Firmware Interface) is the software baked into your
machine's flash chip that runs the instant it powers on — the modern successor
to the legacy BIOS. Its job is to bring the hardware up to a sane state and then
find and launch an operating system. From our point of view it's a small runtime
that hands us a working CPU, a memory map, and a screen, and then gets out of the
way.
The key thing to understand: **UEFI is not our OS, it's a stepping stone.** It
exists to load *us*. Our `boot/efi.zig` is a UEFI *application* — a normal program
that the firmware runs — and its entire purpose is to gather what the kernel needs
and then jump into the kernel.
## How the firmware finds us
UEFI boots by looking for a FAT-formatted partition called the **EFI System
Partition (ESP)** and running a file at a well-known fallback path:
```
EFI/BOOT/BOOTX64.efi <- the "removable media" default for x86-64
```
The boot volume is **FHS-shaped** (see the repository-layout note in
[README.md](../README.md)): `build.zig` installs `boot/efi.zig` (built for the `uefi`
target) at `EFI/BOOT/BOOTX64.efi` — the one path UEFI firmware fixes — and lays
the rest out by FHS path: the kernel at `system/kernel`, init at
`system/services/init`, the pre-packed boot capsule at `boot/system.img`
([system-image.md](system-image.md)).
`zig-out` mirrors that tree, but what a machine actually boots is the
self-contained FAT32 image `tools/make-fat-image.py` builds from the same files
(`danos-usb.img`). The `run-x86-64` step points QEMU at OVMF (UEFI firmware for
virtual machines) and attaches that image (its serial-logging twin, built the
same way) as a USB mass-storage device on the xHCI bus — the guest never sees
`zig-out`. The firmware finds `BOOTX64.efi` on the image and runs it — that's
our `main()`, which then loads the kernel and the system binaries from their
FHS paths.
## Boot services: the firmware's API
While a UEFI app runs, it has access to **boot services** — a table of function
pointers the firmware provides for allocating memory, reading files, locating
hardware protocols, and so on. In `boot()` this is the very first thing we grab:
```zig
const bs = uefi.system_table.boot_services orelse return error.NoBootServices;
```
Everything the firmware offers hangs off tables reachable from
`uefi.system_table`: `boot_services`, `con_out` (the text console we `log()` to),
and the various *protocols* (GOP for graphics, SimpleFileSystem for disk access).
**The critical rule:** boot services are *temporary*. They stop existing the
moment we call `ExitBootServices`. So the loader's structure is dictated by one
constraint — **gather everything the kernel could ever need first, then exit.**
The comment in `boot()` says exactly this:
> Everything the kernel needs must be gathered *before* we exit boot services,
> since afterwards none of these calls are usable.
## What our loader actually does
The four milestones below are the spine of `boot()`. Along the way it also
captures the **ACPI RSDP** from the UEFI configuration table (while boot
services are still up), loads the system binaries into an in-RAM
**initial ramdisk** (`loadSystemTree` — normally a single read of the pre-packed
`boot\system.img` capsule, which already *is* the ramdisk wire format; it falls
back to opening each manifest-listed path, and walks the `/system` and `/test`
trees only as a last resort for hand-assembled sticks — the capsule's format, builder, and
fallback chain are documented in [system-image.md](system-image.md). Best-effort either way — a kernel-only
volume still boots), and builds the **bootstrap page tables** the kernel starts
life on (`buildBootstrapTables`), all before the jump:
### 1. Query the framebuffer (`queryFramebuffer`)
We ask the firmware for the **Graphics Output Protocol (GOP)**, which describes
the linear framebuffer — its address, resolution, pitch, and pixel format —
and, when we can, switch the display to its native resolution first:
- Locate GOP via its *handle* (not `locateProtocol`), because the same handle
also carries the display's **EDID**.
- Read the EDID (trying the `EDID_ACTIVE` then `EDID_DISCOVERED` protocol on each
GOP handle) and parse the first Detailed Timing Descriptor — by convention the
panel's preferred (native) resolution. This is best-effort: firmware installs
these protocols inconsistently, and OVMF with QEMU's stdvga doesn't expose them
at all.
- If we got a native resolution, enumerate the GOP modes with `queryMode` and
`setMode` to the one that matches it (must be a linear 32bpp layout we can
paint into). If EDID gave us nothing, we **keep the firmware's current default
mode** rather than guess — with a valid EDID the firmware normally defaults to
the native mode itself, so its choice beats second-guessing it with, say, the
largest advertised mode.
- Copy the resulting address/resolution/pitch/format into our own `Framebuffer`.
All of this *must* happen now, because after exit there's no GOP to ask. (See
[framebuffer.md](framebuffer.md) for what those fields mean.)
### 2. Load the kernel (`loadKernel` + `loadElf`)
- Use the **LoadedImage** protocol to discover which device we booted from, then
**SimpleFileSystem** to open that volume.
- Open the kernel ELF at its FHS path (`system\kernel`), seek to the end to learn its
size, rewind, and read the whole ELF into a firmware-allocated pool buffer. (`read`
may return short, so we loop.)
- Parse the ELF: validate the `\x7fELF` magic and the `x86_64` machine type, then
walk the program headers. For every `PT_LOAD` segment we:
- reserve the exact physical pages it asks to be loaded at (`p_paddr`) via
`allocatePages`,
- `@memcpy` the file-backed bytes to that address,
- `@memset` the `.bss` tail (the part where `p_memsz > p_filesz`) to zero.
The kernel is linked to *run* in the higher half (virtual base
`0xFFFFFFFF80000000`; `exe.image_base` in `build.zig` is the *virtual* address
`0xFFFFFFFF80100000`) but is *loaded* low: the linker script's `AT()` clauses
give every segment a low physical load address (`p_paddr`, with `.text` at
`0x100000`, 1 MiB), which is what the loader allocates and copies into. The
bootstrap page tables built before the jump map the high link addresses onto
those low physical pages. If a segment's `p_paddr` collided with
firmware-reserved memory, `allocatePages` would fail and we'd need to move the
load addresses.
`loadElf` returns `e_entry`, the kernel's entry-point address.
### 3. Exit boot services (`exitBootServices`)
This is the handoff's trickiest step. To exit, the firmware demands the current
**memory map** and its *key* — proof that we've seen the latest state of memory.
But allocating the buffer to hold the memory map can itself *change* the map,
invalidating the key. So it's a retry loop:
```
get map info -> allocate buffer (+ spare descriptors) -> get map
-> try exit with map.key
-> if it failed, the map moved: free, retry
```
Once `exitBootServices` succeeds, **the firmware's services are gone for good** —
we must never touch `bs`, `con_out`, or any protocol again. The machine is now
entirely ours.
### 4. Jump to the kernel
```zig
handoff(cr3, entry, &boot_information);
```
`handoff` is a single inline-asm block — `cli`, load the bootstrap page tables'
`cr3`, place the `boot_information` pointer in RDI, then `callq *entry` — so
nothing runs between the CR3 load and the jump. This never returns.
## The ABI subtlety: RCX vs RDI
There's a deliberate detail worth calling out. A UEFI binary is compiled with the
**Microsoft x64** calling convention (first argument in register **RCX**). Our
kernel is freestanding and uses the **SysV AMD64** convention (first argument in
**RDI**). If we let each side use its target's default, the loader would place
`boot_information` in RCX while the kernel looked for it in RDI — and the kernel
would read garbage.
So the convention is pinned explicitly to SysV via the shared
`boot_handoff.kernel_abi` (defined in `system/boot-handoff.zig`). The kernel's
`kmainEntry` declares `callconv(kernel_abi)`; the loader honours the same contract
by loading RDI by hand in `handoff`'s inline asm rather than trusting its own
Microsoft-x64 default. `kernel_abi` lives in the shared `boot-handoff` module
because it's a contract both binaries must agree on. See
[sysv.md](sysv.md) for what "SysV" means and where else it shows up.
## The handoff contract
The loader and kernel are two *separate* binaries built for two different targets,
so everything they exchange must have an identically-defined memory layout. That's
what `system/boot-handoff.zig` provides — imported by both as the `boot-handoff` module.
It is *only* the handoff: the kernel↔user ABI (`system/abi.zig`) and the device types
(`library/device/model/device-abi.zig`) are separate contracts the bootloader never sees.
- `BootInformation` — the top-level struct passed to the kernel: the
framebuffer, the memory map, the kernel's own `PT_LOAD` segments
(`kernel_segments` + `kernel_segment_count`, so the kernel can re-map itself
with correct permissions), the ACPI RSDP address, and the initial-ramdisk
base/length.
- `Framebuffer`, `PixelFormat`, `MemoryMap`/`MemoryRegion`, `KernelSegment`,
`kernel_abi` — the shared field layouts and the calling convention.
The structs are `extern struct`, giving them a stable, C-compatible layout so the
bytes the loader writes are the bytes the kernel reads.
## The whole flow at a glance
```
power on
-> UEFI firmware initialises hardware
-> finds EFI/BOOT/BOOTX64.efi on the FHS volume, runs it (our efi.zig main)
-> grab boot services (+ the ACPI RSDP from the configuration table)
-> queryFramebuffer (via GOP: EDID native res, setMode, describe fb)
-> loadKernel (read system/kernel ELF, load PT_LOAD segments low, .text at 0x100000)
-> loadSystemTree (read the boot\system.img capsule as the in-RAM initial ramdisk;
fallbacks: manifest-listed paths, then a /system + /test tree walk)
-> buildBootstrapTables (identity + physmap + higher-half kernel mappings)
-> exitBootServices (retry until the memory-map key holds)
-> handoff: load bootstrap CR3, jump to e_entry, boot_information pointer in RDI
-> kernel _start (architecture/x86_64/isr.s: switch to a kernel-owned stack,
call kmainEntry -> paging, heap, device discovery,
scheduler, SMP, user space)
```
Bottom line: **UEFI's job is to give us a CPU, memory, and a framebuffer, then
disappear.** `boot/efi.zig` is the thin bridge that collects those gifts into a
`BootInformation`, tears down the firmware, and jumps into the kernel — after
which we're on our own.
@@ -0,0 +1,134 @@
# The physical frame allocator
Once the kernel knows what RAM exists ([memory-map.md](memory-map.md)), it needs a
way to *hand out* that RAM: give me a free page of physical memory, and later,
here's one back. That's the **physical frame allocator** (a "physical memory
manager", hence `system/kernel/pmm.zig`). It deals only in fixed 4 KiB **frames** — the
natural unit because that's the granularity the CPU's paging hardware maps — and
it is the primitive everything above it stands on: page tables, the kernel heap,
per-process memory all ultimately ask the frame allocator for pages.
It's **generic kernel code**: it operates on the neutral `boot_handoff.MemoryRegion`
array, so there's no UEFI in it and nothing architecture-specific beyond the 4 KiB
page. (Contrast [architecture.md](architecture.md), which is where CPU-specific code lives.)
## Why a bitmap
There are a few classic designs; danos starts with the simplest that still
supports freeing:
- **Bitmap** (chosen): one bit per frame, `1 = used`, `0 = free`. Freeing is
trivial (clear a bit), it's very compact, and you can later extend it to
allocate *contiguous* runs by scanning for consecutive zero bits. Allocation is
a linear scan, but that's cheap and easy to reason about.
- **Intrusive free-list / stack**: store the "next free frame" pointer inside each
free frame; O(1) alloc and free. Elegant, but it can't satisfy contiguous
multi-frame requests and can't answer "is *this* frame free?".
- **Buddy allocator**: great for contiguous power-of-two blocks, but more
machinery than a first allocator needs.
Compactness matters less than clarity here, but it's a nice property: 128 MiB of
RAM is 32768 frames — a **4 KiB bitmap, a single frame**. Even 64 GiB needs only
2 MiB of bitmap.
## How it works
State lives in `system/kernel/pmm.zig`: the `bitmap` slice, `total_frames`, `used_frames`,
and a `next_hint` marking where the next allocation scan should start.
### init(map) — building it from the memory map
1. **Size it.** Find the highest address across all RAM regions — every kind
*except* `mmio` — so reserved and ACPI spans sit *inside* the bitmap, marked
used but trackable (e.g. so the boot buffers can be freed later);
`total_frames = highest / page_size`. Only MMIO (device address space —
remember the ~12 GiB of it from [memory-map.md](memory-map.md)) sits outside
the bitmap and is simply never allocatable.
2. **Place it (the bootstrap).** The bitmap needs storage before an allocator
exists — a chicken-and-egg. Solution: pick the first `usable` region big enough
to hold the bitmap and put it there, addressing it through the **physmap**
(`boot_handoff.physicalToVirtual`). The loader's bootstrap page tables already
provide the physmap and the kernel's own tables keep it, so the pointer stays
valid across the paging switch.
3. **Mark, then free.** Set the whole bitmap to `used` (`0xff`), then walk the
`usable` regions clearing their bits. Doing it in that direction means every
gap, reserved span, and hole is unallocatable *by default* — we only ever hand
back memory the firmware explicitly called usable.
4. **Take back the essentials.** Re-reserve the frames the bitmap itself occupies
(they're inside a usable region we just freed), plus **frame 0**, so an address
of `0` can keep meaning "no frame".
### alloc() → ?u64
Scan the bitmap from `next_hint` (wrapping once) for the first free bit, mark it
used, advance the hint, and return `frame * page_size`. Returns `null` when no
frame is free — genuine out-of-memory. The hint avoids rescanning the low,
long-since-allocated frames on every call.
### free(addr)
Clear the frame's bit and, if it's below `next_hint`, pull the hint back so the
reclaimed frame gets reused soon. Bogus or double frees (a frame already marked
free, or one out of range) are ignored rather than corrupting the used count.
## Correctness points worth remembering
- **Generic walk.** Because `MemoryRegion` is danos's own type, the map is a plain
slice — none of the variable descriptor-stride from the raw UEFI map.
- **Physmap addressing.** The bitmap (and the region array) is reached through
the physmap via `boot_handoff.physicalToVirtual`, which both the loader's
bootstrap tables and the kernel's own tables provide — no remapping is needed
when danos switches to its own paging. The one invariant: the bitmap must sit
under the bootstrap physmap's reach (4 GiB), which holds because the placement
scan (step 2 above) runs from the lowest usable region up and takes the first
one big enough — on the supported configurations that lands well under 4 GiB.
- **Frame 0 is reserved** so `0` stays a safe "none" sentinel — and the bitmap is
never placed there. (An early bug did exactly that: a `usable` region at physical
address 0 collided with a `0`-means-not-found sentinel and tripped a panic. The
fix was an optional plus starting the bitmap at least one page in.)
- **Everything non-usable is unallocatable by construction** — the "mark all used,
then free usable" order gives that for free, so the kernel image, the loader's
buffers, MMIO and firmware memory can never be handed out.
## Verifying it
`kmain` brings the allocator up and self-tests it. Booted in QEMU with 128 MiB:
```
/system/kernel: frame allocator online
free frames: 30520 (119 MiB) <- matches the map's 119 MiB usable
alloc x3 : 0x3000 0x4000 0x5000 <- frame 0 reserved, bitmap at 0x1000, AP trampoline at 0x2000
after free : 30520 frames free <- three freed, count restored
```
(The AP-trampoline page is claimed with `allocBelow` right after `init`, before
the demo allocations — hence they start at 0x3000.)
The `free frames` MiB agreeing with the memory map's `usable RAM`, the three
distinct consecutive addresses, and the count returning to its start after freeing
are the three signals that init, alloc and free are all correct.
## Boot-services memory comes pre-reclaimed
The UEFI boot-services memory (~44 MiB) is defunct and free once
`ExitBootServices` runs, taking usable RAM from ~76 MiB up to ~119 MiB. The frame
allocator does **nothing special** to get it: the loader already classified it as
`usable` (see [memory-map.md](memory-map.md)), so it's just part of the `usable`
regions `init` frees. Keeping that boot-protocol knowledge on the loader side is
deliberate — the kernel has no notion of "reclaimable" or of UEFI at all.
The one live piece in that memory is the boot stack the kernel starts on; the loader
leaves the single region containing it `reserved`, so `init` won't hand it out. A
later step will move task 0 onto a kernel-owned stack, freeing that last ~1 MiB
region too (and giving user mode the clean stack it wants).
## What's next (partly done since)
- **Contiguous allocation** — done: `allocContiguous` scans for a run of clear
bits, with an optional physical ceiling for DMA (`dma_alloc` is its user), and
`allocBelow` serves the SMP trampoline.
- **A kernel stack for task 0** — still open: the boot processor's idle task runs
on the boot stack to this day, so that region can't be freed.
- **Freeing the `reserved` `loader_data`** (the boot-time map buffers) — still
open: the bitmap deliberately tracks those frames so they *can* be freed, but
nothing frees them yet.
+108
View File
@@ -0,0 +1,108 @@
# The Framebuffer
## What a framebuffer is
A **framebuffer** is just a big region of memory where each element is one
pixel's color. The display hardware continuously scans this memory and turns
each value into light on the screen. There's no drawing API involved — you
write a 32-bit value to the right address, and a pixel changes color. That's
exactly what `Console.pixel` does:
```zig
self.rowPtr(y)[x] = color; // system/kernel/console.zig
```
Our `Framebuffer` struct (`system/boot-handoff.zig`) is the four facts you need to
address it:
| Field | Meaning |
|----------|---------|
| `base` | the memory address where pixel data starts |
| `width` | visible pixels per row (e.g. 1920) |
| `height` | visible rows (e.g. 1080) |
| `pitch` | **bytes** from the start of one row to the start of the next |
The bootloader (UEFI GOP, in our case) sets all this up and hands it over. The
kernel just writes into it: no firmware, no driver — just pixels.
## The mental model: it's 1D memory pretending to be 2D
The screen is a grid, but memory is a flat line of bytes. So the pixels are
stored row after row, laid end to end:
```
row 0: [px0][px1][px2]...[width-1] <padding?>
row 1: [px0][px1][px2]...[width-1] <padding?>
row 2: ...
```
To find pixel `(x, y)` you compute:
```
address = base + y * (bytes per row) + x * (bytes per pixel)
```
## So what is pitch?
**Pitch is "bytes per row"** — sometimes called *stride*. The obvious guess
would be `pitch = width * 4` (4 bytes = 32 bits per pixel). And often it is.
**But not always** — and that's the whole reason the field exists.
Hardware frequently wants each row to start at a nicely aligned address (a
multiple of 32, 64, or a page). If `width` doesn't land on that boundary, the
firmware pads the end of every row with a few extra unused bytes. That padding
is invisible — it's never shown — but it's physically there in memory between
the last pixel of one row and the first pixel of the next.
Example: a 1366-pixel-wide display at 32bpp:
- `width * 4` = 1366 × 4 = **5464 bytes** of actual pixels
- but `pitch` might be **5504 bytes** (padded up to a multiple of 64)
- those extra 40 bytes per row are dead space
This is exactly why `rowPtr` uses `pitch`, not `width`, to step between rows:
```zig
inline fn rowPtr(self: *Console, y: u32) [*]volatile u32 {
const base: [*]volatile u8 = @ptrFromInt(self.fb.base);
return @ptrCast(@alignCast(base + y * self.fb.pitch)); // <- pitch, not width*4
}
```
Note the deliberate detail: `base` is cast to a **byte** pointer
(`[*]volatile u8`) *before* adding `y * pitch`, because pitch is measured in
bytes. Then it's cast to a `u32` pointer so that `[x]` indexes whole pixels. If
you'd done the arithmetic on a `u32` pointer, `+ pitch` would step `pitch`
*pixels* (4× too far).
### Why you must use pitch, not `width * 4`
If you assumed rows were `width * 4` apart on a display where
`pitch > width * 4`, every row would start a little too early. The error
accumulates: row 0 is fine, row 1 is off by (pitch − width×4) bytes, row 2 by
twice that, and so on. The image ends up **skewed diagonally** — a slanted,
sheared picture — because each row creeps sideways relative to where the
hardware actually reads it.
Using `pitch` is what keeps each row landing exactly where the scanout expects
it.
## Two subtleties worth noting
1. **`width` vs `pitch` in the loops.** In `fillRow` we iterate `x` up to
`self.fb.width` — the *visible* count — but jump between rows with
`pitch`. That's the correct pairing: touch only real pixels, but skip the
full stride (including padding) to reach the next row. We never write into
the padding, which is right. (A `copyRow` used to sit alongside it; it's
gone — `scroll` was since rewritten as a writes-only screen clear, because
reading VRAM back is uncached-slow on real hardware.)
2. **`volatile`.** The pointer is `volatile` because this memory is special —
it's watched by the display hardware. `volatile` tells the compiler *"don't
optimize these writes away or reorder/coalesce them"*; every store must
actually hit memory, because something outside the CPU's knowledge (the
scanout engine) is reading it.
Bottom line: **width is how wide the picture is; pitch is how wide the memory
rows are.** They're usually equal (×4) but not guaranteed to be, so always
advance rows by pitch.
+71
View File
@@ -0,0 +1,71 @@
# GOP
Graphics Output Protocol (GOP) is a UEFI driver interface that replaces legacy VGA BIOS functions to provide graphics console output in the pre-OS phase.
It allows the firmware to display boot screens and setup menus by providing direct access to the hardware frame buffer, enabling multiple GPUs to function equally without proprietary INT15 handshaking.
GOP has no standardized mode numbers. Instead, the firmware (really the GPU's GOP driver) builds a list of modes at boot, and each is just an opaque index 0 .. MaxMode-1. Mode 0 might be 1920×1080 on your laptop and 800×600 on someone else's — the index carries no fixed meaning. To learn what a mode actually is, you have to ask:
- gop.mode.max_mode — how many modes exist
- gop.query_mode(n) — returns the info (resolution, pixel format, pixels_per_scan_line) for mode n
- gop.set_mode(n) — switch to it
- gop.mode.info — the info for the currently active mode
queryFramebuffer used to read gop.mode.info directly and never call set_mode — taking whatever mode the firmware selected as its default. It now works out the monitor's native resolution from EDID and, if a matching mode exists, calls set_mode to switch to it. When EDID is unavailable it keeps the firmware's default mode rather than guessing — with a valid EDID present the firmware normally defaults to the native mode itself (this is exactly how the QEMU run boots at 1280x720: OVMF's stdvga driver reads the emulated EDID and defaults to the preferred mode, even though it never exposes the EDID protocols to us). See the "picks the native resolution" note at the bottom.
Does EFI detect native resolution?
Sometimes, but it's not guaranteed by the spec. Here's the actual chain:
1. The GOP driver reads the monitor's EDID over the DDC/I²C wire — a data blob the display publishes describing its supported resolutions, including its preferred (native) timing.
2. The firmware then picks a default GOP mode. Many modern firmwares (and OVMF in a VM, driven by the emulated display) do default to the native/preferred resolution. But plenty of firmware defaults to a safe fallback like 1024×768 or 800×600 regardless of what the panel can do.
So gop.mode.info giving you native res is a common outcome, not a promise. Two more caveats:
- The mode list itself may not even contain the true native resolution — a GOP driver can expose only a handful of modes.
- On a headless VM (your QEMU/OVMF setup), there's no real EDID; the "native" resolution is whatever the emulated GPU advertises. OVMF's default is typically 800×600 or 1024×768 unless you configure it (e.g. QEMU's -device virtio-vga with a set resolution, or the OVMF Platform config).
One note worth flagging: this picks the native resolution but doesn't force a particular pixel format — a native mode is only chosen if it's a paintable linear 32bpp layout (RGBX/BGRX); a bit_mask/blt_only native mode is skipped and we keep the firmware's default instead. That's the right trade-off for now since your console assumes linear 32bpp.
## Pixel formats
Every mode also carries a pixel format, and queryFramebuffer only accepts two of the four GOP formats. The full enum:
- red_green_blue_reserved_8_bit_per_color (RGBX) — accepted
- blue_green_red_reserved_8_bit_per_color (BGRX) — accepted
- bit_mask — rejected
- blt_only — rejected
RGBX/BGRX tell you each pixel is a 32-bit value with a fixed byte order. That's what the console needs: a known layout it can write directly with rowPtr(y)[x] = color. The other two break one of those assumptions.
### bit_mask (spec: PixelBitMask)
The framebuffer is still linear memory you can write to directly — but the bits aren't in a standard RGBX/BGRX arrangement. Instead the firmware hands you a PixelBitmask struct describing where each channel lives:
```zig
pub const PixelBitmask = extern struct {
red_mask: u32,
green_mask: u32,
blue_mask: u32,
reserved_mask: u32,
};
```
Each mask marks which bits of the pixel word belong to that channel. This is how the format expresses non-standard layouts — for example 16-bit RGB565 (red = 0xF800, green = 0x07E0, blue = 0x001F: 5/6/5 bits, only 16 bits per pixel), or an odd 32-bit order. To draw a color you'd have to read the masks, work out each channel's bit position and width, shift/scale your 8-bit R/G/B into place, and OR them together — per pixel. It's fully drawable, just not with a hardcoded 32-bit write, so we reject it rather than carry that machinery.
(RGBX/BGRX are really just two hardcoded special cases of a bitmask. The spec names them separately precisely so simple loaders can skip mask-decoding in the common case.) In practice bit_mask is rare on modern PC firmware — you'll almost always get RGBX or BGRX — so rejecting it costs basically nothing.
### blt_only (spec: PixelBltOnly)
This one is more fundamental: there is no linear framebuffer you can address at all. gop.mode.frame_buffer_base is meaningless — you have no pointer to pixel memory. The only way to put pixels on screen is through GOP's Blt ("block transfer") service, the third function pointer in the protocol: you build pixels in your own buffer and ask the firmware to copy ("blit") a rectangle onto the display. The firmware owns the actual scanout memory, wherever it lives (across a bus, behind a GPU command interface, in a layout the CPU can't map directly).
The catch that matters for us: Blt is a boot service. It stops working the instant you call ExitBootServices — which is exactly when the kernel runs. So a blt_only display gives a post-exit kernel no way to draw pixels at all, and there's genuinely nothing our framebuffer console could do with it. Rejecting it is the only correct response.
### Summary
| Format | Linear memory? | Layout | Console |
| --- | --- | --- | --- |
| rgbx / bgrx | yes | fixed 32bpp byte order | works — direct writes |
| bit_mask | yes | arbitrary, described by masks | rejected — would need per-pixel mask decoding |
| blt_only | no | no CPU-visible framebuffer; Blt service only | rejected — and Blt is gone after ExitBootServices anyway |
So the two-case accept list is the right line to draw: RGBX/BGRX are the only formats that give a post-ExitBootServices kernel a flat block of pixel memory it can write to without help from firmware that no longer exists.
+144
View File
@@ -0,0 +1,144 @@
# Halting
## Why a kernel needs to halt
An ordinary program ends by *returning* — `main` finishes, the C runtime calls
`exit`, and the OS reclaims the process. A kernel has none of that. There is no
OS underneath it, no runtime to return to, and no caller waiting. `_start` is the
end of the line. So when the kernel has nothing left to do — whether it finished
its work or hit a fatal error — it can't "quit". It has to explicitly park the
CPU forever, because if execution ever ran off the end it would just keep fetching
whatever bytes follow in memory and execute garbage.
That's what halting is: deliberately stopping the processor so it does nothing,
safely, until the machine is reset or powered off.
## The core of it: `hlt`
Everything comes down to one x86 instruction. It's CPU-specific, so it lives in
the arch module, `system/kernel/architecture/x86_64/cpu.zig` (see [architecture.md](architecture.md)), and the
generic kernel calls it as `architecture.halt()`:
```zig
/// Park the core forever. `hlt` drops it into a low-power idle until the next
/// interrupt; the loop re-halts on every wake so the stop is permanent.
pub fn halt() noreturn {
while (true) asm volatile ("hlt");
}
```
**`hlt`** ("halt") tells the CPU core to stop executing instructions and drop into
a low-power idle state. It's not a busy-wait — the core genuinely stops, drawing
almost no power and generating almost no heat, until something wakes it.
This is much better than the naïve alternative, a spin loop:
```zig
while (true) {} // "busy-wait" — DON'T do this to idle
```
A bare `while (true) {}` keeps the core running flat out, executing the jump
back to the top of the loop billions of times a second — 100% CPU, hot, and (on a
laptop) draining the battery, all to accomplish nothing. `hlt` achieves the same
"do nothing" outcome while letting the core sleep.
## Why the loop around it?
Here's the subtlety the comment points at: **`hlt` is not permanent.** It halts
the core only until the *next interrupt* arrives. An interrupt is a signal — from
a timer, a keypress, a device — that wakes the CPU so it can respond. When one
fires, the core comes out of `hlt` and executes the next instruction.
If we wrote just a single `hlt`, the very first stray interrupt would wake the
core and execution would continue past it — falling off the end of the function
into whatever comes next in memory. Wrapping it in `while (true)` closes that
door: every time an interrupt wakes the core, the loop immediately runs `hlt`
again and it goes back to sleep. The net effect is a permanent halt that still
sleeps between the interrupts it can't prevent.
(danos handles plenty of interrupts through its interrupt descriptor table —
the timer tick waking a halted idle core is exactly how scheduling works — and
non-maskable and system-management interrupts can wake a halted core regardless
of what we handle. The loop makes the halt robust to all of them.)
## The `asm volatile` part
`hlt` has no equivalent in plain Zig, so we drop to inline assembly:
- **`asm`** emits the raw instruction directly into the function.
- **`volatile`** tells the compiler *"this has side effects you can't see — do
not optimize it away or reorder it."* Without it, an optimizing compiler might
reason that the assembly produces no value anyone uses and delete it, or hoist
it somewhere wrong. `volatile` pins it exactly where we wrote it.
## `noreturn`: telling the compiler it's the end
`halt()` is typed `noreturn` — a real Zig type meaning "this function never gives
control back to its caller." That isn't decoration; it changes how the compiler
treats the call:
- Code *after* a `noreturn` call is unreachable, so the compiler needn't emit a
return sequence, and won't warn about "missing return value" in the callers.
- It lets `kmain` and the exported entry shim `kmainEntry` themselves be
`noreturn`, which is the honest signature for a kernel entry point — the
bootloader jumps in and nothing ever jumps back out.
You can see the chain in the code: `_start` (an assembly stub in
`system/kernel/architecture/x86_64/isr.s`) installs a kernel-owned stack and calls
`kmainEntry` in `system/kernel/kernel.zig`, which is `noreturn`; it calls `kmain`,
also `noreturn`, which ends by calling `architecture.halt()`, again `noreturn`.
The "never returns" property is threaded all the way down.
## Where danos halts
There are three halt sites, and they're all the same idea:
1. **The BSP's idle loop** — `kmain` no longer runs out of work: it spawns
`/system/services/init` as PID 1, drops itself to priority 0, and ends as the
bootstrap core's idle task — still by calling `architecture.halt()`:
```zig
scheduler.setPriority(0);
status("\n/system/kernel: kernel idle; user space is running.\n");
architecture.halt();
```
The timer keeps preempting the idle context into init and whatever else is
ready; between those interrupts, the halt loop is exactly the low-power park
described above.
2. **Kernel panic** — the freestanding panic handler has no OS to report to, so
it prints the message to the diagnostic log and the on-screen console —
forcing the console back on even if a display service had it suppressed — and
halts via the same `architecture.halt()`. A panic is unrecoverable here, so
stopping the machine — rather than limping on with corrupted state — is the
safe response.
3. **Bootloader failure** — in `boot/efi.zig`, if `boot()` fails *before* handing
off to the kernel, `main` logs the error and parks the machine with the same
loop so the message stays on screen:
```zig
boot() catch |err| {
log("\r\nEFI: boot failed: ");
logBytes(@errorName(err));
log("\r\n");
while (true) asm volatile ("hlt");
};
```
(Here it's an inline loop rather than `architecture.halt()` because that lives in the
kernel's arch module, and the loader is a separate binary from the kernel.)
## Summary
- A kernel can't "exit" — it must explicitly stop the CPU or it runs off into
garbage.
- **`hlt`** parks the core in a low-power idle until the next interrupt — far
better than a 100%-CPU spin loop.
- **`while (true) hlt`** makes that halt permanent, since any interrupt would
otherwise wake the core and let execution continue.
- **`asm volatile`** emits the instruction and forbids the compiler from removing
it; **`noreturn`** encodes "control never comes back" into the type system.
- danos halts in the kernel's idle loop, on a kernel panic, and on a bootloader
error — the same "park the core safely" in all three.
+78
View File
@@ -0,0 +1,78 @@
# The kernel heap
The [frame allocator](frame-allocator.md) hands out fixed 4 KiB physical frames;
the [VMM](paging.md) maps pages into virtual addresses. The **kernel heap** sits on
top of both to provide what the rest of the kernel actually wants: `alloc(n)` /
`free(p)` for arbitrary byte sizes. It's the first real consumer of `map()`, and
the thing that unlocks dynamic data structures — lists, hash maps, driver state,
eventually a process table.
It's generic kernel code (`system/kernel/heap.zig`): the allocator logic is
architecture-neutral, using `architecture.mapPage` and the frame allocator underneath.
## A growable free-list allocator
The algorithm is a classic **first-fit free list**:
- The heap owns a virtual region. Free space is tracked as an **address-ordered
singly linked list** of free blocks; each block begins with a 16-byte header
(`size`, and a `next` link used while free).
- **alloc(n)** walks the list for the first block big enough. If the block is much
larger it's **split** — the front becomes the allocation, the remainder stays
free. If nothing fits, the heap **grows** (below) and the search retries.
- **free(p)** finds the block header just before `p` and inserts it back into the
list, **coalescing** with the physically adjacent free blocks on either side so
the space can be reused as one region rather than fragmenting away.
Allocations are 16-byte aligned; larger alignments aren't supported yet (the
`std.mem.Allocator` `alloc` returns `null` for them).
## Growing on demand
The heap lives in the **higher half** of the address space (virtual base
`0xFFFF_8000_0000_0000`) — unmapped, well clear of the low half, which belongs
to user space (unmapped in the kernel's own tables; per-process user address
spaces now map into it). (That base is x86_64-canonical; another architecture
would pick its own.)
When the free list can't satisfy a request, `grow` extends the mapped region: it
pulls fresh frames from the [frame allocator](frame-allocator.md) and `map`s each
onto the end of the heap, then adds the new span as a free block (coalescing with
the current tail). So the heap starts at one page and expands page-by-page as
demand requires, up to a cap. This is exactly what the VMM's on-demand `map` was
built for.
## A std.mem.Allocator
The heap is exposed as a **`std.mem.Allocator`** (`heap.allocator()`), Zig's
standard allocator interface. That's a deliberate multiplier: it means the whole of
Zig's standard library — `ArrayList`, `AutoHashMap`, `std.fmt.allocPrint`, and the
rest — works directly on the kernel heap, no bespoke containers required.
## Verifying it
The `heap` test (see [testing.md](../testing.md)) exercises the allocator end to end:
```
[PASS] alloc 4096 bytes
[PASS] heap memory is writable and reads back
[PASS] freed block is reused <- free list + coalescing works
[PASS] many allocations (heap growth) stay valid <- grow() maps fresh frames
[PASS] std.ArrayList on the kernel heap <- std containers work on it
```
The "freed block is reused" check (free then re-alloc returns the same address) is
the proof that free and the free list actually work, not just alloc; "heap growth"
forces allocation past the initial page so `grow`/`map` runs; and the `ArrayList`
check is the std-integration payoff.
## What's next (largely still true)
- **Thread/interrupt safety** — overtaken by the big kernel lock: SMP arrived
with a single kernel lock taken at every kernel entry, which serializes all
heap access. The heap still has no lock of its own, and needs none unless the
big lock is ever split.
- **Larger alignments** than 16 — still unsupported; page-aligned and DMA
buffers come straight from the frame allocator instead.
- **`resize`/`remap` in place** — still not done; growing an `ArrayList` copies.
- **Reclaiming empty tail pages** — still not done; the heap only ever grows.
+151
View File
@@ -0,0 +1,151 @@
# Interrupts and exceptions
When something goes wrong on the CPU — a bad pointer, a divide by zero, a
malformed page table — the processor raises an **exception**. If nothing is set
up to catch it, the fault escalates: the CPU tries to invoke a handler, finds
none, faults again trying to handle *that*, and on the third strike triple-faults,
which on real hardware and in QEMU means a silent reset. Debugging by spontaneous
reboot is miserable.
This is the machinery that catches those faults and prints what happened instead.
It's all x86_64-specific, so it lives behind the [architecture](architecture.md) boundary in
`system/kernel/architecture/x86_64/`. Only the 32 CPU-defined exception vectors were wired up at
this stage; device interrupts (timer, keyboard, via the APIC) came later, on the
same IDT — today it installs 48 gates (vectors 0-47) plus the ring-3 syscall gate
at vector 128.
## First the GDT
In 64-bit long mode, segmentation is mostly switched off — but the CPU still
requires valid **segment descriptors** for code and data, and, crucially, every
IDT gate names a code-segment *selector* that must resolve in the current GDT. The
firmware left a GDT in place, but we don't control it, so we install our own with
known selectors: `0x08` kernel code, `0x10` kernel data.
`system/kernel/architecture/x86_64/gdt.zig` held three flat descriptors at this point — a required
null entry, plus code and data — where the only bits that matter in long mode are
the access byte and the code segment's long-mode (`L`) flag. (The table has since
grown to seven entries: ring-3 user data and user code descriptors arrived with
user mode, and the TSS descriptor — below — spans two slots.) Loading it (`gdt_flush` in
`isr.s`) does two things: `lgdt`, then reload the segment registers. The data
registers take a plain `mov`, but **CS can't** — so we reload it with a far
return, pushing the new selector and a return address and letting `lretq` pop them
into CS:RIP.
## Then the IDT
The **Interrupt Descriptor Table** maps each of 256 vectors to a handler. Each
entry is a 16-byte *gate* holding the handler's address (split across three
fields, a quirk of the format), the code selector (`0x08`), and flags: `0x8E`
means present, ring 0, 64-bit interrupt gate. `system/kernel/architecture/x86_64/idt.zig` builds the
table, points the first 32 vectors at their stubs (since grown to gates 0-47,
plus the ring-3 syscall gate at vector 128), and loads it with `lidt`
(`idt_flush`).
## The TSS and the double-fault stack
There's one more table, the **Task State Segment**. In long mode its main
remaining job is the **Interrupt Stack Table (IST)**: an IDT gate can name an IST
slot, and when that vector fires the CPU switches to the stack recorded there —
*regardless* of what the interrupted stack looked like.
This matters most for the **double fault** (#DF, vector 8). A #DF means the CPU
hit a fault *while trying to deliver another fault* — very often because the
current stack pointer is bad, so pushing the exception frame itself faulted. If
the #DF handler then tried to push onto that same bad stack, it would fault a
third time and **triple-fault** — an instant reset. So the #DF gate is pointed at
**IST1**, a small dedicated stack (`system/kernel/architecture/x86_64/tss.zig`) that's always valid.
Bringing it up: fill in the TSS's IST1 pointer, publish the TSS through a
descriptor in the GDT (`gdt.setTssFor`), and load it into the task register with
`ltr`. The TSS descriptor is a 16-byte system descriptor spanning two GDT slots,
which is why the GDT grew from three entries to five (and later to seven, when
user mode slotted the ring-3 user data/user code descriptors between the kernel
data entry and the TSS pair).
## The stubs and the trap frame
On an exception the CPU pushes a small frame (SS, RSP, RFLAGS, CS, RIP) and, for
*some* vectors, an **error code**. That inconsistency is a nuisance, so each stub
in `system/kernel/architecture/x86_64/isr.s` normalises it: vectors that don't get a hardware error
code push a dummy `0`, then every stub pushes its **vector number** and jumps to a
shared tail, `isr_common`. The tail pushes all the general registers and calls the
Zig handler with a pointer to the whole thing.
The result on the stack is a uniform **`CpuState`** — register block, then vector
and error code, then the CPU's frame. Its field order in `idt.zig` is exactly the
push order in `isr.s`; the two must stay in sync.
### Why a separate `.s` file
The stubs and table-loads are real assembly rather than Zig inline asm because
they need things inline asm on this toolchain can't express: cross-symbol
`jmp`/`call` (a stub jumping to `isr_common`, which calls the exported
`interruptDispatch`), and the `lgdt`/`lidt` memory operands (which LLVM rejects
inline). `build.zig` adds `isr.s` to the arch module.
## Reporting a fault
`isr_common` calls `interruptDispatch`, which forwards to a swappable `on_fault`
hook. The generic kernel installs a reporter (`onException` in `kernel.zig`) that
prints the exception name and vector, the error code, the faulting RIP
and RSP, and — for a page fault (#PF, vector 14) — the faulting address from
**CR2**. What happens next depends on where the fault came from:
- **User mode (CPL 3): kill the process, keep the machine.** The kernel is intact
(the CPU trapped onto the task's kernel stack), so the faulting process is
killed — address space, IRQ bindings, and IPC handles reclaimed; a client it
owed a reply to is failed with `-EPEER` — and the core reschedules. A crashing
driver takes itself down, never the OS. This is fault recovery step 2 of
[resilience.md](resilience.md). NMI, double fault, and machine check are
excluded: they report machine trouble regardless of what was running.
- **Kernel mode: halt this core.** The trusted base itself is broken, so there is
nothing safe to kill; the fault is still *contained* to the core (an
application-processor fault leaves the rest of the system running), and the
report makes it **visible** instead of a silent reset.
The hook is set before `arch.init()` in `kmain`, so a fault during setup is still
caught.
## Verifying it
A temporary `ud2` (unconditional invalid-opcode instruction) in `kmain` produced,
in red:
```
CPU EXCEPTION: invalid opcode (vector 6)
error code : 0x0
RIP : 0x000000000010dadd <- the ud2, in the kernel image at 0x100000+
RSP : 0x0000000007e8aed0
```
Vector 6 with no error code, a RIP inside the loaded kernel, and a sane RSP
together confirm the whole path: the GDT is active (we're still executing), the
IDT vectored to the right stub, the stub built a correct `CpuState`, and the Zig
handler read it and reported instead of triple-faulting.
Separately, pointing RSP at an unmapped address and faulting forced a **double
fault** — reported cleanly (`double fault (vector 8)`) rather than triple-faulting
into a reset, which only works because #DF ran on IST1. That's the proof the
TSS/IST is wired up: the handler survived a completely broken stack.
## What's next (since done)
Both items originally deferred here have landed:
- **The IO-APIC**: [ioapic.zig](../../system/kernel/architecture/x86_64/ioapic.zig)
routes external device lines onto vectors — discovered via ACPI's MADT, every
input masked at init, lines unmasked one at a time as user-space drivers bind
them (see [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). The keyboard followed
exactly as predicted: the PS/2 bus driver (`system/drivers/ps2-bus/`) claims
the 8042 controller and binds its IRQ 1 (and the aux mouse's IRQ 12) through
this routing. USB HID keyboards arrive over xHCI instead, which interrupts via
MSI, and the HPET's GSI routing exercises the same path.
- **SSE state**: `isr_common` (and the syscall entry) now `fxsave`/`fxrstor` the
full SSE/x87 register file around dispatch. This stopped being optional the
moment kernel code touched XMM — a 16-byte struct copy is a `movdqu` — and its
absence was the root cause of a long-lived corruption Heisenbug; the comments
in `isr.s` tell the story.
With faults now debuggable, the paging work that follows — where a wrong
page-table entry means an instant #PF — is far less painful.
+100
View File
@@ -0,0 +1,100 @@
# Logging
Output is a *diagnostic convenience, never a correctness dependency*: the kernel
and every service must run correctly with zero output channels. On top of that
rule, danos has **per-process logging** — every process's output is attributed
by the kernel and lands in its own file on the flash volume, which is what makes
a headless real machine (no serial port) debuggable. The display (the
framebuffer surface) is a separate concern with one bootstrap exception: the
kernel's framebuffer console (`system/kernel/console.zig`) joins the log sinks
at boot, so the whole transcript shows on screen until the display service
claims the framebuffer and silences it; after that only panic/fatal messages
are mirrored to it explicitly (`system/kernel/kernel.zig`).
## The pipeline
```
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /var/log/<boot-stamp>/<binary-path>.log
kernel log.print ─┘ │
└▶ serial / 0xE9 sinks (QEMU, -Dserial)
```
1. **Emit.** A program calls `std.log.info("mounted {s}", .{path})` — the
runtime's `logFn` (installed for every binary by the root shim,
`library/kernel/logging.zig`) formats one line and issues one `debug_write`
carrying the level. The payload does NOT contain the process's name.
`logging.write` remains as the raw/bring-up path (panics, test
fixtures); raw bytes ride the same ring, attributed all the same.
2. **Stamp.** The kernel wraps every payload LINE in a record stamped with the
sender's pid, task name (its binary path, e.g. `/system/services/fat`),
level, a per-boot sequence number, and a monotonic timestamp
(`system/kernel/log.zig` + `log-ring.zig`). Attribution is structural — a
payload cannot forge another sender's tag, and an embedded newline just ends
the record, so the forged "prefix" lands inside the forger's own next line.
3. **Retain.** The 512 KiB ring overwrites oldest-first; sequence gaps make any
loss countable. `klog_read` (#32) copies stream bytes from a free-running
offset; `klog_status` (#45) returns the cursors plus the wall-clock time of
boot. The framing (`abi.KlogRecordHeader`) is 32 bytes + name + payload,
8-byte aligned.
4. **Render.** Registered sinks (serial under `-Dserial`, the 0xE9 debug
console, and the framebuffer console until the display service claims the
screen) get a live transcript: kernel/raw output verbatim, leveled records
as `<binary path>: message` — one composed write per line, under the log's
own spinlock (never the big kernel lock; panic paths try-acquire with a
bound and fall back to sinks-only). Sinks are best-effort and self-guarding;
a serial-less machine just goes quiet.
5. **Persist.** The **logger service** (`system/services/logger`) drains the
ring every 250 ms and demultiplexes records into one file per source under
`/var/log/<boot-stamp>/`, e.g.
```
/var/log/2026-07-21T150434Z/kernel.log
/var/log/2026-07-21T150434Z/system/services/fat.log
/var/log/2026-07-21T150434Z/system/drivers/usb-storage.log
```
The boot stamp is the RTC anchor from `klog_status` (FAT-safe: no colons; a
dead RTC yields the 1970 directory rather than no logs). Each line carries
the record's monotonic timestamp and level. Storage is best-effort and late:
the ring buffers a whole boot many times over, and the first successful
`makePath` of the per-boot directory (also the readiness probe) triggers a
full backlog write. Files close — which is the fat server's SCSI cache
flush — after a ~2 s quiet period, bounding data-at-risk without per-record
flush thrash. At shutdown init stops the logger FIRST (it is last in the
boot order), so its final drain runs over a live storage chain.
## Why a ring in the kernel, not a logging server
The storage stack must be able to log. If the fat server wrote its own log file
through the VFS it would rendezvous-deadlock on itself; if processes sent
records to a logging server over IPC, early boot would need a buffer that is —
a ring, one hop later. The kernel ring is that buffer, placed where every
process (and the kernel itself) can reach it with one syscall, before any
service exists. The logger service is a *reader*, not a hop.
Two disciplines keep it honest:
- the logger announces itself **once** — a periodic status line would feed the
very stream it drains;
- lost records surface as an explicit `-- N records lost --` line, computed
from sequence gaps, never silently.
## Last-resort channels
Unchanged, and independent of the sink list so they survive a total output
failure: `checkpoint` (a one-byte POST code on port 0x80) and `recordPanic`
(a fixed breadcrumb record, `log.panic_record`, findable in a RAM dump; magic
written last so a reader only trusts a complete record).
## Accepted gaps
- A write-spamming process can evict other processes' unread records from the
ring (a per-process quota is future work); the loss is at least visible via
sequence gaps in every affected file.
- `/var/log` files have no privacy until the VFS grows permissions.
- Records emitted after the logger's final shutdown drain reach serial and the
ring but not the files.
+182
View File
@@ -0,0 +1,182 @@
# The memory map
Before a kernel can manage memory, it has to *know what memory exists*: which
physical address ranges are real RAM it may use, and which are firmware, hardware
registers, or already occupied. That inventory is the **memory map**, and the
firmware is the only thing that knows it. This page covers how danos gets that map
from the firmware and hands it to the kernel — deliberately without dragging UEFI
into the kernel.
## Why not just pass UEFI's map through?
UEFI hands the loader a perfectly good memory map. The tempting shortcut is to
forward it to the kernel as-is. We don't, for two reasons:
1. **It would tie the kernel to UEFI.** The kernel would compare against UEFI's
memory-type numbers and walk the array using UEFI's variable descriptor stride.
That's UEFI vocabulary bleeding across the handoff — and danos wants to boot on
systems that have no UEFI at all (a Raspberry Pi describes its memory with a
*device tree* instead). See [architecture.md](architecture.md) for the same "keep the kernel
platform-agnostic" principle applied to CPU code.
2. **We already established the better pattern.** The loader doesn't hand the
kernel a raw UEFI GOP either — [`queryFramebuffer`](gop.md) converts it to
danos's own `Framebuffer`. The memory map follows the same discipline.
So the boundary is: **each boot path translates its native memory description into
danos's own neutral format, and the kernel only ever sees that.**
## The neutral format
Defined in `system/boot-handoff.zig`, the shared loader↔kernel contract:
```zig
pub const MemoryKind = enum(u32) {
usable, // free RAM the kernel may allocate
reserved, // firmware / kernel image / boot stack — real RAM, never hand out
acpi_tables, // parse, then reclaim
acpi_nvs, // preserve across sleep
mmio, // device registers / reserved address space — not RAM at all
};
pub const MemoryRegion = extern struct {
base: u64, // physical start
pages: u64, // length in page_size (4 KiB) units
kind: MemoryKind,
_pad: u32 = 0,
};
pub const MemoryMap = extern struct {
regions: usize, // pointer to a [len]MemoryRegion
len: usize,
};
```
`MemoryKind` is danos's *own* vocabulary — not UEFI's ~15 types, just the
distinctions the kernel actually acts on. And because danos defines `MemoryRegion`
itself, `@sizeOf` is authoritative: the kernel walks a plain `[]MemoryRegion` with
no variable-stride subtlety (that stride problem is a UEFI-ism, and it stays in the
loader).
`BootInformation` carries it alongside the framebuffer (trimmed here to the
fields this page is about — the full struct has since grown the kernel's
PT_LOAD segments, the ACPI RSDP, and the initial-ramdisk span):
```zig
pub const BootInformation = extern struct {
framebuffer: Framebuffer,
memory_map: MemoryMap,
// ...kernel_segments, acpi_rsdp, initial_ramdisk_base/len
};
```
## The loader side (UEFI)
Two functions in `boot/efi.zig`: `exitBootServices` calls `convertMemoryMap`,
which runs `classify` on each descriptor:
- **`classify`** maps each UEFI descriptor to a `MemoryKind`:
`conventional_memory` **and** `boot_services_code`/`boot_services_data → usable`;
`acpi_reclaim_memory → acpi_tables`; `acpi_memory_nvs → acpi_nvs`;
`memory_mapped_io`/`memory_mapped_io_port_space → mmio`; **everything
else → reserved** (the safe default). Our own `loader_data` — the kernel image and
these buffers — falls into `reserved`.
Folding boot-services memory into `usable` is deliberate: we've already called
ExitBootServices, so it's free RAM now, and doing the classification *here* (in
the loader) means the kernel never learns about a UEFI-specific "reclaimable"
state — it just sees usable RAM. The one catch is that our stack lives in
boot-services memory and the kernel starts out running on it, so
`convertMemoryMap` keeps the single region containing the current stack pointer
`reserved`. All the boot-protocol knowledge stays on the loader side of the
boundary; the kernel's frame allocator has no idea any of this happened.
One subtlety: **a region that isn't writeback-cacheable (the descriptor's `wb`
attribute) is classified `mmio` regardless of type.** UEFI overloads
`reserved_memory_type` for both reserved RAM *and* reserved address-space windows
(PCIe config space, device BARs); the cache attribute is what actually tells them
apart, since only real RAM is writeback-cacheable. Without this, a QEMU q35 guest
reports ~12 GiB of "reserved" that is really a PCIe address hole near the 1 TB
mark — not memory at all.
- **`convertMemoryMap`** walks the UEFI descriptors (striding by
`descriptor_size`, *not* `@sizeOf`), classifies each, and writes danos
`MemoryRegion`s into an output buffer, coalescing adjacent same-kind regions.
### The ordering that makes it correct
This is the fiddly part, dictated by two UEFI rules: you can only allocate memory
*before* `ExitBootServices`, and the memory map is only final *at* the moment you
exit (its "key" proves you've seen the latest state). So `exitBootServices` does,
per attempt:
1. `getMemoryMapInfo` to size things, then `allocatePool` **two** LoaderData
buffers — one for the raw UEFI map, one for the converted regions. Allocating
now, before exit, is mandatory.
2. `getMemoryMap` then `exitBootServices(key)`. If either fails (allocating can
perturb the map and invalidate the key), free both buffers and retry.
3. **After** the exit succeeds, convert. Conversion is pure computation on memory
we already hold — no boot-services calls — so it's safe once services are gone.
Both buffers are `LoaderData`, which survives `ExitBootServices`, so the converted
array the kernel is pointed at stays valid. (The raw UEFI buffer is just scratch
for the conversion.)
## The kernel side
The kernel receives a plain array and reads it with zero UEFI knowledge:
```zig
const mm = boot_information.memory_map;
const regions = @as(
[*]const boot_handoff.MemoryRegion,
@ptrFromInt(boot_handoff.physicalToVirtual(mm.regions)),
)[0..mm.len];
for (regions) |r| {
if (r.kind == .usable) usable_pages += r.pages;
}
```
(`mm.regions` is a physical address, so it's dereferenced through the physmap —
`physicalToVirtual` — since the kernel no longer runs under the loader's
identity map.)
`kmain` summarises the map to prove the handoff works. Booted in QEMU with
128 MiB, it reports:
```
/system/kernel: physical memory
total RAM : 0.12 GiB (127 MiB) - RAM the firmware reported
usable : 121 MiB - free RAM (incl. reclaimed boot-services memory)
reserved : 6 MiB - kernel image, boot stack, ACPI, runtime services
regions : 28 - entries in the firmware memory map
```
`usable` is ~121 of ~127 MiB because the loader already folded the boot-services
memory into it — so the frame allocator gets it all with no special step. The ~6 MiB
`reserved` is the kernel image, the boot stack's region, ACPI, and runtime services.
`total` counts only writeback-cacheable RAM, so the ~12 GiB PCIe address hole is
excluded (it's `mmio`), and the RAM categories summing back to the firmware's total
is the sanity check that nothing was dropped.
## How Raspberry Pi will fit
No UEFI there, but the boundary is unchanged. The Pi's firmware jumps into the
kernel with a **device-tree blob**; the AArch64 entry code will parse its
`/memory` and `/reserved-memory` nodes and produce the *same* `MemoryRegion`
array. The kernel's memory code — the frame allocator and everything above it —
never knows the difference.
## What's next
This page is plumbing plus classification only. The map's first consumer, the
**physical frame allocator**, is built directly on the `usable` regions here — which
already include the reclaimed boot-services memory the loader folded in (see
[frame-allocator.md](frame-allocator.md)). Of the two items once listed here, one is done:
- Freeing the `reserved` `loader_data` (these boot-time buffers) once the kernel
is done reading the map — still open: the frame allocator's bitmap tracks those
frames so they can be freed, but nothing frees them yet.
- Capturing the ACPI RSDP from the UEFI configuration table before exit (the same
"grab it before ExitBootServices" pattern) — done: the loader stows it in the
boot handoff, and ACPI parsing consumes it from there ([acpi.md](acpi.md)).
See the roadmap in [efi.md](efi.md) for where this sits in the boot flow.
+145
View File
@@ -0,0 +1,145 @@
# Paging: the kernel's page tables and VMM
Every memory access the CPU makes goes through the **page tables**: hardware walks
them to translate a virtual address into a physical one, and faults if there's no
valid mapping. danos builds its own tables (rather than staying on the firmware's,
which live in memory we'd like to reclaim and don't control), switches CR3 onto
them, and — crucially — maps with **real permissions**.
It's x86_64-specific (the 4-level table format is an Intel/AMD thing), so it lives
behind the [architecture](architecture.md) boundary in `system/kernel/architecture/x86_64/paging.zig`.
## The format
x86_64 uses **4 levels**: PML4 → PDPT → PD → PT, each a 512-entry table, with 9
bits of the virtual address indexing each level and the low 12 bits the offset into
the final 4 KiB page. Each entry holds a physical address plus flag bits —
present, writable, and (bit 63) **no-execute**. danos maps nearly everything with
4 KiB pages: precise, and the extra table memory is negligible against available
RAM. (The physmap has since become the one exception: 2 MiB-aligned RAM there is
mapped with **2 MiB huge pages** — a PS-bit leaf at the PD level — with 4 KiB
pages filling the unaligned edges, so the table footprint scales sanely with big
RAM. Kernel segments, heap, user space, and on-demand MMIO stay 4 KiB.)
## Higher half: the address-space layout
danos is a **higher-half kernel**. The kernel is linked to run at
`0xFFFF_FFFF_8000_0000` but loaded low (the linker script's `AT()` gives each
segment a physical load address at 1 MiB up; the bootloader maps the high link
address to the low load address in its bootstrap tables and jumps in). The entire
**low canonical half is reserved for user space**; the kernel lives in the top half
alongside a **physmap** — a straight window onto all of physical memory at
`physmap_base + phys`. Wherever the kernel needs to touch a physical address (a
page-table frame, an ACPI table, a device register), it adds that constant:
`boot_handoff.physicalToVirtual(phys)`. The layout constants live in `system/boot-handoff.zig`:
| region | virtual base | PML4 slot |
|--------|--------------|-----------|
| user image + stack | `0x0000_7000_0000_0000` | 224 (low half) |
| kernel heap | `0xFFFF_8000_0000_0000` | 256 |
| physmap (all RAM + MMIO windows) | `0xFFFF_8800_0000_0000` + phys | 272 |
| kernel image | `0xFFFF_FFFF_8000_0000` | 511 |
The bootloader builds temporary **bootstrap tables** (identity + a 4 GiB physmap +
the high kernel) so it can switch CR3 and jump to the high entry; the kernel then
builds its own precise tables below and abandons them. Because both use the same
`physmap_base`, any physmap pointer minted before the switch stays valid after it.
## What gets mapped, and with what permissions
The address space is built in four passes (`init`):
1. **All RAM in the physmap, RW + NX.** Every non-MMIO region from the
[memory map](memory-map.md) is mapped at `physicalToVirtual(phys)`, read-write and
*non-executable*. There is **no low/identity mapping** — the low half is user
space. (Frames the kernel touches while still building these tables are reached
through the loader's bootstrap physmap, which covers the low 4 GiB; both the
frame allocator and the table builder scan low-address-up, so those frames stay
under that limit.)
2. **The framebuffer and the Local APIC**, the device memory the kernel touches
directly, as physmap windows (RW + NX). Other MMIO is mapped on demand by
`mapMmio`, also into the physmap; everything else is left unmapped, so a stray
access faults instead of silently succeeding.
3. **The kernel's own segments, overlaid with their true ELF permissions**, at
their high link addresses mapped to their low physical load addresses. This is
the interesting part.
4. **Every higher-half PML4 entry pre-created** (an empty PDPT where none exists
yet). The kernel half is then a fixed set of top-level slots, so a per-process
address space can share it by copying `PML4[256..512)` once — growth beneath
those slots (heap, on-demand MMIO) propagates to every address space because
they share the PDPTs. `init` asserts no new higher-half PML4 entry appears
afterward.
### W^X from the ELF program headers
Blanket RW+NX is fine for data but wrong for the kernel's own code, which must be
executable — and its code must *not* be writable (W^X: no page is both). We get the
right permissions per region straight from the kernel ELF: the **loader already
parses the program headers**, so `efi.zig` records each `PT_LOAD` segment's
address, size and R/W/X flags into `BootInformation`. Pass 3 re-maps those ranges with
flags derived from the ELF flags:
| segment | ELF flags | mapped as |
|---------|-----------|-----------|
| `.text` | R + X | present, **not** writable, **not** NX |
| `.rodata` | R | present, not writable, NX |
| `.data`/`.bss` | R + W | present, writable, NX |
So code can execute but not be written, and data can be written but not executed.
(Intermediate table entries are left writable and executable so the *leaf's* bits
govern — a page is writable only if every level is, and non-executable if any level
is.) NX itself has to be switched on first via `EFER.NXE`, or the NX bit would be a
reserved bit and fault.
### The null guard
The whole low half is unmapped except for explicit user mappings, so page 0 (and
every near-null address) is unmapped by construction. A null (or near-null) pointer
dereference in the kernel takes a page fault instead of quietly reading or writing
real memory — turning a whole class of silent bugs into an immediate, located crash.
## Switching on, and the on-demand API
Loading the PML4's physical address into **CR3** switches address spaces and
flushes the TLB in one step. This works because *while we build* the kernel is
still running on the loader's bootstrap tables, whose 4 GiB physmap makes freshly
allocated table frames reachable at the same `physmap_base + phys` addresses;
afterwards they're covered by pass 1.
`init` keeps the PML4 and the frame allocator around and exposes `map(virt, phys,
writable)` / `unmap(virt)` (with `invlpg` TLB invalidation) — the primitive the
kernel heap will build on to map pages on demand.
## Verifying it
Four tests (see [testing.md](../testing.md)) pin down the guarantees:
- **`vmm`** — map a fresh frame at an unused virtual address, write and read it
back. Proves `map` works end to end.
- **`fault-pf`** — an access far above all mapped RAM faults, with the address in
CR2. Proves we're on our own (deliberately sparse) map.
- **`fault-nx`** — calling into a data page (NX) faults on the instruction fetch.
Proves NX is enforced.
- **`fault-null`** — writing to address 0 faults. Proves the null guard.
> Toolchain notes, both hit while writing the tests: a volatile access to a
> compile-time-*constant* address either trips the self-hosted backend's
> `mov moffs` gap or (for address 0) Zig's null-pointer safety check — so the
> null-guard test launders the address through empty asm and uses an `allowzero`
> pointer to force a real hardware access. And `invlpg`, like `lgdt`, needs its
> operand staged through a register in inline asm.
## What's next (mostly done since)
- **A kernel heap** — done, built on `map` exactly as anticipated
([heap.md](heap.md)).
- **A higher-half kernel** — done: the kernel is linked at
`0xFFFFFFFF80000000` (`linker.ld`), loaded low and running high, and user
processes own the low half.
- **Per-address-space tables** — done: each user process gets its own root with
the kernel half shared, and refcounted shared-memory mappings exist
([ipc.md](../device-driver-development-guide/ipc.md)). Copy-on-write remains unbuilt — nothing has needed it yet.
- **Uncacheable MMIO** — half done: user-space device and DMA mappings are
strong-uncacheable and the framebuffer is write-combining via the PAT, but the
kernel's own `mapMmio` path is still writeback — the LAPIC included (see
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md)).
+130
View File
@@ -0,0 +1,130 @@
# The power service: events and shutdown
A laptop lid closes, a battery drains, someone presses the power button — and
several parts of the system might care: a session manager dims the screen, a
logger notes it, and ultimately *something* has to turn the machine off. None of
them owns the hardware that reported the event, and the reporter should not know
who is listening. So system power is a **service**: an event source **publishes**
button/lid/battery/AC events, interested processes **subscribe**, and one
privileged caller — init — can ask it to power the machine off. It is the same
publish/subscribe shape as the [input service](../device-driver-development-guide/input.md), applied to power.
## Why a service, and why it is named for the domain, not the firmware
Where the events come from is firmware-specific — on x86 they ride the ACPI SCI
([acpi.md](acpi.md)); on a Raspberry Pi they would come from PSCI or a mailbox.
What subscribers want is not: *the lid closed* means the same thing regardless of
who noticed. So the surface is **domain-named**. There is a `power-protocol`
module and a well-known `ServiceId.power = 5`; on x86 the **acpi service**
registers it, and on ARM a PSCI/mailbox service will register the *same* id.
Subscribers call `ipc.lookup(.power)` and never learn which firmware they
are on — the neutrality the whole [discovery](discovery.md) migration exists to
preserve, carried one layer up into a running-system surface.
This is why the protocol is `power`, not "ACPI events": naming a cross-firmware
surface after one firmware would leak x86 into code the ARM port must reuse
unchanged.
## The protocol
The `power-protocol` module ([library/protocol/power/power-protocol.zig](../../library/protocol/power/power-protocol.zig))
follows the vfs-protocol pattern — extern-struct messages, a version, reserved
fields. Three operations:
| Direction | Operation | Purpose |
|---|---|---|
| subscriber → service | `subscribe` | receive published events; the subscriber's endpoint rides as the call's **capability** (the input/device-manager pattern) |
| init → service | `shutdown` | orderly shutdown's last step: enter S5 (soft off) |
| service → subscriber | `event` | a published `EventMessage`, delivered as a buffered message (never sent *to* the service) |
Events are published, not polled: like the input service, the service holds
subscriber endpoints as capabilities and `ipc_send`s each event as a buffered
message, so a slow or dead subscriber can never wedge the source. The event
vocabulary is hardware-neutral:
- `power_button` — the button was pressed (a fixed ACPI event on x86).
- `lid`, `ac`, `battery` — the named GPE-driven events.
- `notify` — a device notification that maps to none of the above; its `code`
(the ACPI `Notify` argument) and the notifying device's `hid` say which device
and what happened.
An `EventMessage` carries the `event` tag plus `code` and an 8-byte `hid`, so a
generic `notify` is fully described without a second round trip.
**`shutdown` is authority, not information.** It is the only operation that
*does* something irreversible, so it is gated: the contract is that only init
(PID 1) may request it, because init is the process that has already run the stop
sequence over everything else. The acpi service implements this as a **soft
gate** — it honors `shutdown` only from a process that is a *subscriber*, and
init is the one subscriber. That stands in for "only the system supervisor may
power off" without hard-coding a pid, so it still holds under tests where PID 1
is not init.
## Orderly shutdown
Powering off cleanly is where the power service, the [process
lifecycle](process-lifecycle.md), and [ACPI events](acpi.md) compose. init
already supervises the services it starts; for shutdown it runs **one event loop
over one endpoint** that carries three things at once: its children's exit
notifications, the lifecycle **signals** it can receive (`terminate`), and the
**power events** it subscribes to — plus a re-arming heartbeat timer proving PID
1 is alive. (init subscribes with retries, because the power service registers
`.power` well after init starts; a missing power service is not fatal — a
`terminate` signal drives the same path.)
On a `power_button` event or a `terminate` signal, init:
1. logs that it is shutting down,
2. runs the standard stop sequence — `process.stop(child, deadline,
endpoint)` — over its children **in reverse spawn order**, so the VFS stops
last (other services may flush through it), each child getting the
*terminate → deadline → kill* escalation from
[process-lifecycle.md](process-lifecycle.md), and
3. requests `.power` `shutdown`.
The service then enters **S5** (soft off) by writing `SLP_TYP | SLP_EN` to the
PM1 control register(s) from ring 3, with the `SLP_TYP` values taken from its own
AML parse of the `_S5` object — the kernel has no S5 path of its own. If the
write returns instead of powering the machine off, it logs loudly so a test
fails rather than hangs.
**No new system call was needed for S5.** The broad io_port grant on the
`acpi-tables` node ([discovery.md](discovery.md)) already put the PM1 control
ports in the acpi service's hands, so writing S5 from ring 3 is something it
could physically already do; formalizing it as a protocol operation added a
contract, not authority. The kernel keeps only **reboot** (`acpi.reboot` in
`system/kernel/acpi.zig` — the FADT reset register plus the legacy fallbacks, which
need no AML); it has no poweroff path at all — S5 is not a kernel operation.
## Verifying it
Two QEMU scenarios exercise the path, both injecting a real ACPI power-button
press via QMP `system_powerdown` (there is no other deterministic power event on
this config):
- `power-button` proves the source: the acpi service's SCI handler logs the
press and publishes `power_button` (the ACPI half is in [acpi.md](acpi.md)).
- `orderly-shutdown` proves the whole composition: button → init logs shutting
down → children stopped → the service enters S5 → QEMU exits. The ordered
regex is the proof, and QEMU's self-exit through S5 is the pass.
## Scope
Interface-complete but validated on real hardware (the author's laptop) later,
because QEMU does not emulate them: battery `_BST`/`_BIF` evaluation beyond the
interface stubs, lid and AC events, and the embedded controller's `_Qxx`
queries. Deliberately out of scope for now: reboot over the power protocol, S3
sleep, per-device D-states (a future lifecycle-vocabulary extension, since
"suspend" has the shape of a signal every driver must answer and has no consumer
until laptop sleep), and thermal zones.
## See also
- [acpi.md](acpi.md) — where the events come from on x86: the SCI, the power
button fixed event, and GPE/Notify dispatch in the acpi service.
- [discovery.md](discovery.md) — why the surface is domain-named, and the
firmware neutrality that makes a PSCI backend drop-in on ARM.
- [process-lifecycle.md](process-lifecycle.md) — the stop sequence
(`terminate → deadline → kill`) and signals init composes into shutdown.
- [device-manager.md](../device-driver-development-guide/device-manager.md) — the supervision model init mirrors for
its own children.
@@ -0,0 +1,342 @@
# Process lifecycle: signals over IPC
**Status: increments 1–4 built** (2026-07-12): claim release on death, exit
reasons, published exit events, and signals + one-shot timers + the service
harness are all in — the interface below is as-built. The primitives underneath
predate this design ([process-management.md](process-management.md):
spawn, the supervision link, kill, child-exit notifications); this document designs
the layer above them — the standard vocabulary a danos process speaks about its own
life, and the stable `process` interface that carries it. Nothing here is
device- or driver-specific: a driver, the VFS, and a user application all stop,
reload, and die the same way. The device manager is simply this design's first
serious customer ([device-manager.md](../device-driver-development-guide/device-manager.md)).
**"POSIX" in this document means the concepts, never the letter of the standard.**
danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE
was a mistake) without inheriting the mechanism, the API, or the names. The naming
rule is danos's own and it is strict: plain words that communicate intent
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
a concept that already has one. Literal POSIX arrives later and lives elsewhere: the
`std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C
layer** on the same native surface (see [zig-self-hosting.md](../zig-self-hosting.md)) —
musl's syscall surface retargeted at danos system calls and IPC protocols (files onto
the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever
networking becomes). Ported programs see POSIX; the system underneath never does.
## Why a standard vocabulary
A supervisor can only manage processes it has never heard of if "please exit" means
the same thing to all of them. That is the one thing POSIX signals got deeply right:
`SIGTERM` means the same thing to nginx and to a five-line script, which is why
process supervision on Unix (init systems, container runtimes) is possible at all.
danos wants that property from day one, because supervision-and-restart is the
system's core motivation ([resilience.md](resilience.md)).
What POSIX got wrong — for a system like this — is the **delivery mechanism**:
asynchronous control-flow hijack. A Unix handler runs on a stolen stack at an
arbitrary instruction boundary, which is why the async-signal-safe function list
exists, why `errno` must be saved, and why the canonical signal bug is a SIGTERM
handler innocently calling `printf` mid-`malloc`. That entire bug class comes from
the mechanism, not the vocabulary, and none of it is worth importing.
A microkernel already has the right channel: **a signal is a message.** QNX delivers
POSIX signals over its message passing; seL4 has notification objects; Erlang turned
"death is a message to whoever linked" into a reliability philosophy. danos has
already done it once without naming it: a child's death arrives as a notification
badge on the supervisor's endpoint — the microkernel's SIGCHLD, the IRQ-as-IPC
pattern reused. Signals are the same pattern reused a third time.
## The mechanism
- **`signal_bind(endpoint)`** — a process nominates the endpoint its signals arrive
on, exactly as `irq_bind` nominates where a device's interrupts land. The runtime
does this at startup for any program that opts in.
- **`process_signal(id, signal)`** — posts the signal as an asynchronous
notification to the target's bound endpoint: badge = `notify_badge_bit |
notify_signal_bit | pending signals`. Non-blocking for the sender, always.
Signals address the *process*: `id` may name any member of a threaded process
and resolves to its leader — whose endpoint the harness binds — with authority
mirroring `process_kill` ([shared-fate-plan.md](shared-fate-plan.md)).
- **Pending signals coalesce** in a per-process bitmask while the target has no
signal endpoint bound, and the whole mask arrives as one notification at bind —
POSIX's own semantics for non-realtime signals (two pending SIGTERMs are one
SIGTERM). Once bound, each `process_signal` flushes the mask straight into the
endpoint's notification ring, so a busy receiver drains separate posts as
separate notifications — harmless, because the badge is a set of bits, never a
count. The bitmask *is* the design: signals carry no payload. Anything with a
payload is a protocol message.
- **Authority**: the supervisor may signal its children — the same link that is
already the kill authority. A process may signal itself. Anything broader waits
for transferable process handles.
- **No binding, no problem**: a process that never calls `signal_bind` is not
broken — its signals pend unread and only `process_kill` works on it. Simple
programs stay simple; the vocabulary is opt-in, the kill authority is not.
Because delivery is a message into the process's own event loop, there is no
async-signal-safe list in danos: a handler is ordinary code running at a point the
process chose. The bug class is gone by construction, not by discipline.
## The vocabulary: POSIX.1-1990, sorted honestly
The full 1990 set, and what each becomes. Two intrinsically problematic cases get a
defense below the table.
| POSIX.1-1990 | danos disposition | Notes |
|---|---|---|
| SIGTERM | signal `terminate` | finish up and exit; the supervisor's polite half |
| SIGHUP | signal `reload` | re-read configuration / re-scan |
| SIGINT | signal `interrupt` | interactive interrupt; meaningful once a console can send it, in the vocabulary now so numbering is stable |
| SIGQUIT | signal `quit` | as SIGINT, without the core-dump baggage |
| SIGALRM | signal `alarm` | timer expiry as a message; the Unix SIGALRM+`longjmp` timeout hacks are impossible here. In the vocabulary, unbuilt: no consumer yet, and when one appears it is runtime sugar over the existing timer — zero kernel work |
| SIGUSR1, SIGUSR2 | signals `user_1`, `user_2` | service-defined |
| SIGCHLD | **already exists** — the exit notification | the badge carries the child id, dodging the classic coalescing bug (Unix code must loop `waitpid`) |
| SIGKILL | `process_kill` — kernel mechanism | its definition is "cannot be handled"; it was never really a signal |
| SIGABRT | exit reason `aborted` | synchronous self-termination is an exit, not an event; recorded for any nonzero exit code |
| SIGSEGV, SIGILL, SIGFPE | exit reasons, **never delivered** | see below |
| SIGPIPE | **an error return**, not a signal | see below |
| SIGSTOP, SIGTSTP, SIGTTIN, SIGTTOU, SIGCONT | deferred | job control needs terminals, sessions, and process groups; stop/continue is scheduler territory |
**The fault signals (SIGSEGV, SIGILL, SIGFPE) are intrinsically wrong for messages.**
They are *synchronous* — raised at a specific faulting instruction, not "sometime
soon". A message cannot be delivered to a process whose next instruction re-faults;
it never reaches its event loop to read it. POSIX only makes fault handlers "work"
via the async hijack (run the handler *instead of* the instruction), and even there,
returning from a SIGSEGV handler without curing the cause is undefined behavior.
danos's architecture already has the better answer: fault → the kernel kills the
process ([resilience.md](resilience.md) step 2, built) → the supervisor reads the
reason → restart. Recovery is restart, not a handler. This is also truer to the 1990
standard than handling is: the standard's default action for all three was
"terminate the process".
**SIGPIPE deserves special contempt.** Its default kills a process that writes to a
closed pipe — which is why "the whole server died because one client disconnected"
is roughly every network daemon's first production bug, and why every mature codebase
contains the same fix: ignore SIGPIPE, handle the `EPIPE` error return. danos made
the right choice natively already — a reply owed to a dead peer fails with `-EPEER`.
Errors from operations are error returns from those operations. The posix layer can
synthesize SIGPIPE for ported code that expects it.
### Statements, not questions
A signal and a protocol message both travel over IPC — the difference is the
**contract**, not the transport. danos IPC has two primitives, both already in
daily use: the **asynchronous notification** (a badge — bits that coalesce into a
pending mask; the sender never blocks; no payload, *no reply path*; how IRQs and
exit events arrive) and the **synchronous call** (a rendezvous — payload both
ways, the caller waits for the reply; how VFS requests work). A signal is the
first kind: a *statement*. `terminate` wants no reply — the exit notification is
its acknowledgement.
A health probe is the second kind: a *question*, worthless without its answer —
and the answer's absence within a deadline is the very thing being measured.
Asked as a signal it has no reply channel (a coalescing bit can't carry an answer,
and the authority rule forbids a child signalling its supervisor back); asked as a
call, the timeout-is-the-diagnosis semantics come free. So there is no `health`
signal. Liveness is the common **`ping`**: a reserved request every harness-run
service answers automatically on its main endpoint — still free for the service
author, still one obvious way — and a supervisor's probe is a `ping` call with a
deadline.
## The two iron rules
1. **Cleanup is the kernel's job.** A process can die with no warning — fault,
kill, power. Correctness must never depend on a `terminate` handler running. On
any death the kernel releases the address space, IPC handles, IRQ bindings,
owed replies, and **device, I/O-port, and interrupt claims and MSI vectors** —
the last of these was once the known gap in
[process-management.md](process-management.md), closed by increment 1
(`releaseTaskResourcesLocked`, on every death path). A signal handler is
for *graceful* work — flushing, deregistering, saving — never for *necessary*
work.
2. **Kill is not a signal, and exit reasons are load-bearing.** The standard stop
sequence is *terminate → deadline → `process_kill`*; the unhandleable kill stays
a kernel mechanism. And a supervisor deciding whether to restart must know *how*
the child died: clean exit (meant to — don't restart), fault (restart with
backoff), killed (the supervisor did it). The exit notification carries only the
id; the reason is recorded before the notification posts and read with the
supervisor-gated `process_exit_reason` query. Restart policy cannot be
written without it.
## Who learns of a death
A death has three audiences, and conflating them is how systems end up with either
zombie state or privileged snooping:
1. **The supervisor** — gets the exit notification on the endpoint it gave at spawn
(built), then reads the `ExitReason` with the `process_exit_reason` query
(increment 2). The supervisor is the only
audience that needs the *reason*, because it is the only one deciding whether to
restart.
2. **The peer owed a reply** — already built: a client that dies mid-request fails
the server's reply with `-EPEER`; a server that dies fails its waiting clients
the same way. This covers the *synchronous* case only.
3. **The subscribers** — the new piece, and it is the input service's
publish/subscribe shape ([input.md](../device-driver-development-guide/input.md)) applied to exits. A stateful
service accumulates per-client state across many requests: a filesystem server
(FAT today) holds a dead client's open file handles, the input service holds
its subscriptions, a future network stack holds its sockets. None of these
are the client's supervisor, and none learn anything from a failed reply if
the client simply never calls again.
So the kernel **publishes every exit** to whoever subscribed:
`process_subscribe(endpoint)` adds a subscriber, and each death posts a
notification to every subscriber (badge = `notify_exit_bit | process id` — the
same encoding supervisors already decode, the IRQ-as-IPC pattern once more). The
subscriber filters for ids it holds state for and releases what the dead client
held. Correlating is free of bookkeeping: an IPC sender's badge already *is* its
task id (`ipc.Received`), so the id a service has been keying client
state by all along is the id the exit event carries.
Subscription, not broadcast-to-everyone: only processes that asked receive
events, the kernel keeps a bounded subscriber table, and delivery is the same
non-blocking coalescing notification as everything else — a dying process never
waits on its mourners. Subscribing is ungated, like `process_enumerate`: what is
running (and dying) is not a secret between cooperating processes. Subscribers
do not receive the exit reason — the filesystem server does not care *why*
the client died.
This is the service-side mirror of iron rule 1: **a service must never depend on
its clients cleaning up after themselves.** Handle release on client death is the
service's job, triggered by the published exit event — never by a courtesy
"closing now" message that a crashed client will never send.
## The stable interface: `process`
`process` already owns what a process receives at birth (`Init`, the
argv contract). It grows to own the other end of life.
**The runtime is the stable interface; the numbers are not.** danos applications do
not make system calls — they call the runtime library, and the system-call numbers,
notification bits, and signal bit positions beneath it are a **private kernel ↔
runtime contract** that may change at any time (settled 2026-07-12). This is why
the runtime exists. Today kernel and runtime ship from one tree in one image, so
"stability" is simply building them together. When driver binaries start shipping
as separately-versioned applications — the whole point of the restart design — the
binary's embedded runtime version becomes compatibility metadata (the same idea as
the protocol version in the device manager's `hello`), and the kernel refuses what
it cannot serve. Signals therefore need no reserved numbering scheme: the enum
below is vocabulary, not ABI.
```zig
/// The signal vocabulary. The value is the bit position in the pending mask — a
/// private kernel/runtime detail, free to change while they ship together.
pub const Signal = enum(u5) {
terminate = 0, // SIGTERM: finish up and exit
reload = 1, // SIGHUP: re-read configuration
interrupt = 2, // SIGINT
quit = 3, // SIGQUIT
alarm = 4, // SIGALRM
user_1 = 5, // SIGUSR1
user_2 = 6, // SIGUSR2
};
/// A decoded pending mask: the coalesced set of signals a notification delivered.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool { ... }
};
/// Nominate `endpoint` as this process's signal endpoint (signal_bind). The
/// runtime's service harness calls this; a bare program may call it directly and
/// fold signals into its own replyWait loop.
pub fn bindSignals(endpoint: usize) bool { ... }
/// Decode a received badge into signals, or null if the badge is not a signal
/// notification (mirrors ipc.Received.isChildExit).
pub fn signalsFrom(badge: u64) ?SignalSet { ... }
/// Send `signal` to process `id`. Supervisor-gated, like kill; non-blocking.
pub fn sendSignal(id: u32, signal: Signal) bool { ... }
/// The standard stop sequence: terminate, wait up to `deadline_ms` for the exit
/// notification on `exit_endpoint` (the endpoint the child was spawned with),
/// then process_kill. The one call a supervisor needs.
pub fn stop(id: u32, deadline_ms: u64, exit_endpoint: usize) void { ... }
/// Subscribe `endpoint` to published exit events (process_subscribe). Every
/// process death posts an asynchronous notification: badge = notify_exit_bit |
/// process id — the same encoding a supervisor's exit notification uses, decoded
/// by the same ipc.Received helpers. For stateful services: release what the dead
/// client held (file handles, subscriptions, sockets). Ungated, like
/// process_enumerate.
pub fn subscribeExits(endpoint: usize) bool { ... }
/// How a process ended — queried after the exit notification (the kernel records
/// it first, so the two never race). What restart policy reads. (Built in M17.2.)
pub const ExitReason = enum(u8) {
exited, // returned from main / clean exit
aborted, // deliberate failure exit — any nonzero exit code (SIGABRT's ghost)
segmentation_fault, // SIGSEGV's ghost
illegal_instruction, // SIGILL's ghost
arithmetic_fault, // SIGFPE's ghost
protection_fault, // general protection fault
fault, // any other CPU exception
killed, // process_kill
};
```
Two deliberate absences. There is no `mask`/`block` API — a process that is not
ready for a signal simply has not waited on its endpoint yet; the pending mask *is*
the blocked set. And there is no per-signal handler registration at this layer —
dispatch is the process's own `switch` over `SignalSet`, or the service harness's
callbacks (`on_terminate`, `on_reload`) for programs that want defaults.
### The service harness
`service` owns the `replyWait` loop and folds every event source — signals,
child exits, protocol messages — into callbacks, with the vocabulary's defaults:
`terminate` returns from the loop (clean exit), the common `ping` is answered automatically,
`reload` is ignored unless overridden. One loop, no locking, nothing reentrant. A
service author writes domain logic; the lifecycle contract is satisfied by the
harness. A process that bypasses the harness and ignores its signals meets the
deadline-then-kill escalation — you cannot force a process to implement an
interface, but you can make compliance free and non-compliance fatal.
### The musl layer later
The POSIX C layer is a **musl port**: musl's arch/syscall layer retargeted so that
what musl believes are kernel syscalls become danos runtime calls and IPC — `open`
and `read` onto the VFS protocol, `kill`/`sigaction`/`waitpid` onto this document's
vocabulary, `exit` onto the runtime's exit path. `sigaction` handlers registered
through it are invoked by the runtime's loop when the signal message arrives —
synchronous underneath, async-looking to ported code, delivered at wait boundaries
the way most Unix programs already experience signals (at syscalls). No stack hijack
ever happens, `SA_RESTART` semantics come free because nothing was interrupted, and
SIGPIPE can be synthesized from `-EPEER` for the programs that expect it. C programs
get POSIX; danos-native programs never pay for it.
## Increments
1. **Kernel: release device/port/IRQ claims and MSI vectors on death** — the
cleanup half of iron rule 1, and the prerequisite for any restart story. Test:
kill a claiming driver, spawn it again, the claim succeeds.
2. **Exit reason in the death notification** (`ExitReason` above).
3. **Exit events**: `process_subscribe` in the kernel (bounded subscriber table,
publishes on every death), `process.subscribeExits`; the userspace VFS
router was the first subscriber — releasing a dead client's handles was its
proof test — and the FAT server inherited the role when the router moved into
the kernel (clients now hold the filesystem server's node ids directly).
4. **Signals**: `signal_bind` + `process_signal` + the pending mask in the kernel;
`process` grows the interface above; the service harness handles
`terminate` and answers the common `ping`; `stop()` for supervisors.
[device-manager.md](../device-driver-development-guide/device-manager.md) builds directly on all four.
## Settled questions (2026-07-12)
- **Signal numbering is not ABI**: the runtime is the stable interface; the numbers
beneath it are a private kernel ↔ runtime contract (see "The stable interface").
- **Liveness is a `ping` call, not a signal**: signals are statements, questions
are synchronous calls (see "Statements, not questions"). A service wanting *deep*
health ("can I reach my hardware?") defines its own protocol message on top.
- **Process handles: deferred.** Pids + the supervisor gate cover everything
planned; transferable handles (Fuchsia-style, delegating signalling without
delegating kill) wait for the capability table to grow types beyond endpoints.
- **`alarm`: in the vocabulary, unbuilt.** No consumer yet; when one appears it is
runtime sugar over the existing timer (arm a timer that posts your own signal) —
zero kernel work, so deferring costs nothing.
- **Subscription granularity: all exits**, subscriber-side filtering — one
subscription per service, a bounded kernel table. Per-id subscriptions only if
event volume ever matters (hundreds of processes, not before).
- **Client identity across the exit boundary: no convention needed** — an IPC
sender's badge already is its task id (see "Who learns of a death").
@@ -0,0 +1,130 @@
# Process Management
How danos lists, supervises, and kills processes — the microkernel answer to
`ps`, `kill`, and `SIGCHLD`/`wait`.
## Why system calls, not `/proc`
Unix systems sit on a spectrum. Classic BSD/macOS list processes through
syscalls (`sysctl(KERN_PROC)`) and kill through `kill(2)`; Linux renders the
process table as `/proc` for *reading* but still kills through a syscall; Plan 9
made the file tree the whole interface (`echo kill > /proc/n/ctl`). Microkernels
mostly abandon ambient PIDs: Minix and QNX route everything through a user-space
process-manager server, and Fuchsia/seL4 control processes only through handles.
danos rules out `/proc` **as the primitive**: the path router lives in the
kernel (`fs_resolve`), but what is mounted under a path is served by a
user-process filesystem server (the way FAT serves `/mnt/usb`) — a `/proc`
would be one more such server, which would put a user process in the path of
process control. If that server (or anything under it) hangs, nothing could be
listed or killed, *including the hung server*. The control plane for processes
must not depend on a process. So the primitives are kernel system calls; a
read-only `/proc` rendering can be layered on later, and a POSIX-style
process-manager server can be built *from* these primitives when one is needed.
## The three primitives
### `process_enumerate(buffer, maximum) -> total`
A snapshot of the task table into a caller buffer of `abi.ProcessDescriptor`
(id, supervisor, state, priority, name) — the exact shape of
`device_enumerate`, so `ps` is a user program over a snapshot, not a kernel
service. The total may exceed what fit; call again with a larger buffer. Kernel
tasks are included with an empty name — an honest listing shows the idle tasks
too. Ungated and read-only: what is running is not a secret between cooperating
bring-up processes.
### `system_spawn(..., exit_endpoint) -> child id`, and the supervision link
`system_spawn` records the caller as the child's **supervisor** and returns the
child's process id (ids are monotonic, never reused — a stale id can only miss).
That link is the kill authority: it answers "who may kill process 7?" without
inventing users or permissions, the same way a device *claim* is the capability
for `mmio_map`. It composes with the supervision hierarchy the device manager
already forms: init supervises the services it starts, the device manager
supervises the drivers it matches. (A transferable process *handle* — Fuchsia
style — can replace the id once the handle table grows types beyond endpoints.)
`exit_endpoint` (a handle, or `abi.no_cap`) is the supervisor's death-watch: when
the child ends — clean exit, CPU fault, or `process_kill` — the kernel posts an
asynchronous notification to that endpoint, exactly like a bound IRQ. The badge
carries `abi.notify_badge_bit | abi.notify_exit_bit | child_id`, so one endpoint
supervises many children and can even share with IRQ notifications. This is the
microkernel's SIGCHLD: no new mechanism, just the IRQ-as-IPC pattern reused, and
a supervisor's event loop (`ipc.replyWait`) already knows how to receive it. The
child holds a reference to the endpoint from birth, so the notification cannot
dangle even if the supervisor dies first.
### `process_kill(id) -> 0 / -ESRCH / -EPERM`
Only the supervisor may kill; kernel tasks are not killable processes. The kill
is a **whole-process** kill ([shared-fate-plan.md](shared-fate-plan.md)): `id`
may name any member of a threaded process — it resolves to the group's leader,
authorization is checked against the *leader's* supervisor, and every thread
dies. Like a signal, delivery is prompt but asynchronous — 0 means the kill is
accepted and irrevocable; the exit notification (badged with the leader, posted
once the last member is gone) confirms completion.
## How a kill lands (the kernel mechanics)
Everything below runs under the big kernel lock, where task states cannot move.
- **Target ready or blocked** (not on any core): reaped on the killer's own
call. The reap releases what death always releases (IRQ bindings first, then
a client the target still owed a reply to is failed with `-EPEER`, IPC handles
closed, the exit notification posted last) — plus the unlinking only a
*remote* death needs: out of the ready queue, out of an endpoint's sender FIFO
(`Task.ipc_wait_endpoint`), out of a receive wait queue (`Task.wait_queue`),
and out of any server's owed-reply slot, so nothing ever dequeues a dangling
pointer. Destroying the address space is safe because no core can have it
loaded: every switch away from a task loads the next task's tables.
- **Target running on another core**: it cannot be torn down mid-instruction,
so it is condemned (`Task.kill_pending`) and dies at whichever comes first:
- its next **system_call entry** — checked before dispatch, so a condemned
process cannot spawn, claim, or message anything on its way out;
- its core's next **timer tick** — but only when the task is not inside one
of its own system calls (`Task.in_system_call`): the tick may have
interrupted kernel code mid-operation, where teardown would leak whatever
the operation held. User-mode execution is always a safe kill point. The
tick-time terminate abandons the interrupt frame exactly like the fault
path (the LAPIC is acknowledged before the tick hook runs);
- any core's tick finding it **blocked or ready** (it entered a syscall and
parked after being condemned) — reaped by the same remote-reap path.
A pure user-mode spin loop that never makes a system call therefore dies
within one tick; nothing a process does can outrun the kill.
The scheduler stays below the process layer: finishing a kill (IRQ bindings,
handles, the notification) is called *up* through two hooks process.zig
registers at boot (`terminate_current_hook`, `reap_task_hook`), mirroring how
the architecture layer calls up into `tick`.
## Known gaps (bring-up honesty)
- ~~Device claims are not released on death~~ Closed (M17.1): every path out of a
process releases its device claims alongside its IRQ and MSI bindings
(`releaseTaskResourcesLocked`), so a restarted driver can claim its hardware
again — the cleanup half of [process-lifecycle.md](process-lifecycle.md)'s iron
rule 1. The `claim-release` test proves the kill → release → re-claim cycle.
- ~~Kernel stacks of dead tasks are leaked~~ Closed (threading-plan M8): a task
exiting on its own core queues on the core's reap list in `.reaping` state, and
the next switch away (or tick) frees its kernel stack; one killed while off-CPU
has its stack freed synchronously by the reap itself. Both paths are accounted
by `live_stack_bytes`, which returns to baseline when no extra tasks are live.
- ~~There is no exit status in the notification~~ Closed (M17.2): the kernel
records how every process ends — exited, a fault class, or killed — before it
posts the exit notification, and the supervisor reads it with
`process_exit_reason` (`process.exitReason`). This is the input to
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
for the clean case can still ride alongside later.
- Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
isolation break).
## Tests
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
notifications), `supervision` (the whole user-side surface via the process-test
service: spawn supervised → enumerate → kill blocked and spinning children →
notifications → gone), `claim-release` (a killed claim-holder's device is
claimable again). See test/qemu_test.py.
+70
View File
@@ -0,0 +1,70 @@
# The release ISO — flashable boot media
`zig build release-x86-64` produces **`zig-out/danos-x86-64.iso`**, the file you
hand to someone who wants to try danos on a real machine: point
[balenaEtcher](https://etcher.balena.io) (or Raspberry Pi Imager, or plain `dd`)
at it, flash a USB stick, and boot the stick. The same file also burns to
optical media. `zig build check-iso-image` validates it without booting.
```
zig build release-x86-64
# Etcher: select danos-x86-64.iso → select the stick → Flash
# or: sudo dd if=zig-out/danos-x86-64.iso of=/dev/rdiskN bs=4m (macOS; triple-check N)
```
## Why an ISO when danos-usb.img already boots
`danos-usb.img` is a raw FAT32 **superfloppy** — a filesystem starting at
sector 0, no partition table. UEFI firmware accepts that from a USB stick (it
probes whole-disk FAT before giving up), which is why `dd`-ing the .img works
and why QEMU and the test harness boot it directly. But it is a
developer-shaped artifact: flashing apps expect an ISO, and a superfloppy
can't be burned to a CD/DVD or carry a partition table for pickier firmware.
The ISO wraps that same FAT image — bit-identical, built by the same
`tools/make-fat-image.py` — in a container that boots everywhere release media
gets consumed. One payload, two images: the .img stays the raw volume the QEMU
harness mounts and boots, the .iso is what leaves the building.
## How a hybrid ISO boots twice
The trick (the same one Linux distribution ISOs use, usually via `xorriso
-isohybrid…`) is that ISO9660 reserves its first 32 KiB as a **system area** it
never touches — exactly where an MBR lives on a disk. So one file can carry two
tables of contents, both pointing at the same embedded FAT image:
* **Flashed to USB (Etcher, dd):** firmware sees a disk whose sector 0 is an
MBR with one partition of type `0xEF` (EFI System Partition) covering the
embedded FAT image. It mounts that ESP and runs `\EFI\BOOT\BOOTX64.efi` —
the standard removable-media path ([efi.md](efi.md)).
* **Burned to optical media:** firmware reads the ISO9660 volume descriptors
at sector 16 and finds an **El Torito** boot record. Its catalog has one
entry, platform ID `0xEF` (EFI), whose start LBA is — again — the embedded
FAT image. The firmware exposes that image as a virtual disk and runs the
same `BOOTX64.efi` off it.
Neither path involves the legacy BIOS boot-sector machinery: danos is
UEFI-only ([system-requirements.md](../system-requirements.md)), so the MBR holds
no boot code, just the partition entry, and the El Torito entry is EFI-class,
not floppy emulation.
One El Torito wrinkle: the catalog's sector-count field is 16-bit (units of
512 bytes), so it can name at most 32 MiB — less than the 64 MiB FAT image.
That is fine in practice: firmware sizes the FAT filesystem from its own BPB,
and the boot files sit in the first few MiB of the image (clusters are
allocated from the front) either way. The USB path has no such cap.
## The builder
`tools/make-iso-image.py` follows the house rule of
[make-fat-image.py](../../tools/make-fat-image.py): pure Python 3 standard
library, no external tools (no xorriso, mkisofs, or isohybrid), with a
`--verify` mode the `check-iso-image` step runs — it checks that the MBR
partition and the El Torito catalog agree on where the FAT image lives and
that a FAT32 boot sector is actually there. Every timestamp field in the ISO
is zeroed, so the build is reproducible byte-for-byte.
The ISO9660 filesystem around the boot machinery is minimal but real: a root
directory listing `BOOT.CAT` (the catalog) and `EFI.IMG` (the FAT image), so
`file`, mount tools, and archive browsers can open the ISO and see what's in
it.
+165
View File
@@ -0,0 +1,165 @@
# Resilience: fault isolation and live restart
Steps 1–4 of the ordering below are **built** (M17–M18, 2026-07-13): user-mode
isolation; fault → kill the process → keep the core (`onException`; the
`fault-recovery` test); the supervisor notification **with exit reasons**
([process-lifecycle.md](process-lifecycle.md) — clean exit, fault class, or
killed, recorded before the notice posts); and the **restart policy itself**
([device-manager.md](../device-driver-development-guide/device-manager.md)): the device manager supervises every
driver, restarts crashes with backoff, caps crash loops, and re-claims work
because the kernel releases a dead process's claims. The `driver-restart` and
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
end to end. What remains of this document's ladder is scope, not mechanism:
more of the system moved into restartable processes (the discovery migration,
[discovery.md](discovery.md), is the next rung). This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](../vision.md) shape was chosen, and it's a *separate* goal
from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) —
one that's less pervasive to build (see [vision.md](../vision.md)).
## The idea: "let it crash" + supervision
The philosophy is older than microkernels and shows up across systems: don't try to
make every component perfect — make failures **contained and recoverable**. Isolate
each component, watch it, and when it dies, restart it from a known-good state. A
small trusted core supervises a fleet of restartable, untrusted parts.
Prior art worth studying (see Further reading): **MINIX 3's reincarnation server** (a
driver crashes, a supervisor restarts it live — the closest thing to your goal),
**QNX** (restartable drivers on a message-passing microkernel), **Erlang/OTP
supervision trees** ("let it crash", not a kernel but the canonical design), and the
historical **Tandem NonStop** (fault-tolerant by process pairs).
## Why a microkernel makes this possible
The blast radius of a fault is the address space it happens in. In a monolith, a
driver bug can corrupt anything — the kernel *is* the driver. In a microkernel,
drivers and services are **isolated user-space processes**, so a fault is trapped by
the kernel and confined to that one process. The kernel — the one thing that *can't*
be restarted, because it's the trusted base — stays tiny, which is precisely why a
small kernel is a *more recoverable* kernel: less code that can take the whole system
down. **Keeping the kernel minimal is a resilience strategy, not just an aesthetic.**
## The building blocks
1. **Address-space isolation.** A fault in one component can't corrupt another or the
kernel. This is the [user-mode milestone](../vision.md) (ring 3, per-process page
tables) — the shared prerequisite for *any* of this, and it's needed regardless.
2. **Fault detection** — how the system notices a component is dead or sick:
- **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps
to the kernel, which kills the process and notifies the supervisor. The clean,
easy case — and danos already reports CPU faults (see
[interrupts.md](interrupts.md)); user mode turns "halt on fault" into "kill the
process and tell the supervisor."
- **Hang**: a livelocked or infinite-looping component needs a **watchdog /
heartbeat** and the ability to **preempt and kill** it. The preemptive scheduler
already built ([scheduling.md](scheduling.md)) is what makes a runaway component
killable — a nice case of a scheduling mechanism serving resilience without any
real-time *guarantee*.
- **Misbehaviour**: IPC timeouts, failed health checks.
3. **A supervisor / reincarnation server.** A user-space server holding the *policy*:
what components exist, their dependencies, and each one's restart strategy. When a
component dies, it decides whether/how to restart it. (MINIX 3 calls this the
reincarnation server; Erlang calls it a supervisor.)
4. **A resource model that supports clean teardown.** When a component dies, its
resources — memory, MMIO grants, IPC channels, IRQ routes — must be **reclaimed**,
and a restarted replacement must be able to **re-acquire** them. This is where a
**capability** model shines (seL4's reference design): a component holds
capabilities to its resources; killing it **revokes** them, which frees everything
in one clean sweep, and the supervisor hands the replacement fresh caps. A simpler
grant/ownership table can work too — capabilities are the principled version.
5. **Re-initialisable drivers.** A driver must start from a known state and
re-establish its hardware. Some hardware is easy to reset; some holds state that's
hard to recover — a real limit on what "just restart it" can fix.
## The hard part: restarting *correctly*
Detecting and killing is the easy half. The genuinely tricky questions are about the
*rest of the system* when a component dies:
- **In-flight IPC**: messages sent to the dead component, or replies its clients are
blocked waiting for. The channel has to break cleanly and unblock the waiters with
an error rather than hang them forever (a design constraint that reaches back into
[ipc.md](../device-driver-development-guide/ipc.md) — channels need a "peer died" outcome).
- **Clients**: how does a client discover the service it was talking to is gone and
has been replaced? Options: capability revocation makes stale handles fail; or a
**name server** re-binds clients to the new instance; or clients retry through a
stable endpoint.
- **State**: the cheapest model is **stateless restart** — the replacement starts
fresh and clients re-establish whatever they need. Richer options (checkpointed
state, state handed to a standby) are more work and more failure modes. Start
stateless.
These are the constraints most worth *bumping into and researching* — they're where
resilience gets genuinely interesting.
## Kernel mechanism vs user-space policy
The microkernel split applies to fault management itself:
- **Kernel (mechanism):** isolation, trapping faults, enforcing capabilities/grants,
IPC, creating/destroying address spaces, granting/revoking resources, preempting a
runaway task.
- **User space (policy):** the supervisor decides *what* to restart, *when*, and
*how* — dependency order, retry limits, escalation. None of that belongs in the
kernel.
So the kernel gains a few primitives (kill an address space, reclaim its resources,
deliver a "child died" notification); everything smart lives in a user-space server.
## What's *not* recoverable this way
Honest boundaries:
- **The kernel itself.** It's the trusted base; if it faults, this mechanism can't
save it. The mitigation is to keep it tiny — the microkernel bet.
- **Corrupted hardware state.** Isolation limits the blast radius to one process, but
if a driver wedged the device itself, a restart may not un-wedge it.
- **Shared-resource corruption** that happened *before* the fault was detected. Clean
capability revocation limits this, but it's why fault *detection latency* matters.
## Suggested ordering
1. **User mode + address-space isolation** — the shared prerequisite (also on the
path for everything else). **Done.**
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
"confine to the process and report it." **Done** (the kill and reclaim; the
supervisor notification waits for step 3's supervisor). A killed server's
pending client is unblocked with `-EPEER` rather than hung.
3. **A minimal supervisor server** that can (re)start a process.
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
table.
5. **First restartable driver** — the keyboard — as the end-to-end proof: crash it on
purpose, watch it come back.
## Relationship to real-time
Resilience needs **structural** features (isolation + supervision + a resource
model); real-time needs a **pervasive** timing invariant. They're separable, and
resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](../vision.md)).
Note the overlap, though: **preemptive scheduling** and **priorities** — already
built — serve resilience too (you can preempt and kill a misbehaving component, and
run the supervisor at high priority). So danos keeps the useful *mechanisms* of the
real-time work without owing anyone a timing *guarantee*.
## Further reading
- Herder, Bos, Gras, Homburg, Tanenbaum — the **MINIX 3** papers, esp. *"Construction
of a Highly Dependable Operating System"* and *"Fault Isolation for Device
Drivers"* — the reincarnation server, the closest match to danos's goal.
- **QNX** architecture — a shipping microkernel with restartable drivers.
- **Erlang/OTP** supervision trees and the *"let it crash"* philosophy — the design
pattern, distilled.
- **seL4** capability model — the principled basis for clean resource teardown.
- **Tandem NonStop** (historical) — fault tolerance via process pairs.
## Related
- [vision.md](../vision.md) — the goals this serves (learning by doing; resilience over
hard real-time).
- [scheduling.md](scheduling.md) — preemption, which makes runaway components killable.
- [ipc.md](../device-driver-development-guide/ipc.md) — channels that need a "peer died" outcome for clean restart.
- [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and
restart" instead of "halt".
- [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context.
+149
View File
@@ -0,0 +1,149 @@
# Scheduling
The scheduler turns danos from a linear "boot then halt" kernel into a **running
multitasking system**. It's **fixed-priority preemptive**: the highest-priority
ready task always runs, and tasks at the same priority take turns. That model is
chosen for [real-time](../vision.md) — it's predictable (you can reason about which
task runs when) and its decisions are O(1), unlike a fair-share scheduler.
The scheduler proper (`system/kernel/scheduler.zig`) is generic; the context switch and new-task
stack setup are architecture-specific (`system/kernel/architecture/x86_64/`, see [architecture](architecture.md)).
## Tasks
A **task** is a kernel thread: ring-0 code with its own 16 KiB stack (allocated
from the [heap](heap.md)). A task struct holds its saved stack pointer, priority,
state, and a ready-queue link. The currently-running kernel context (kmain)
registers itself as task 0, so there's always something to switch *from*.
## The context switch
Switching tasks means swapping stacks. `switch_context(old, new)` (in `isr.s`)
saves the **callee-saved** registers on the current stack, stores the stack pointer
into the old task, loads the new task's stack pointer, restores *its* callee-saved
registers, and `ret`s — landing wherever the new task was last suspended. Only
callee-saved registers are handled explicitly: to the compiler this looks like a
normal function call, so it already preserves the caller-saved ones itself (this is
the [SysV](sysv.md) convention doing the work).
A **freshly spawned** task has never run, so there's nothing to restore. Its stack
is faked to look as if it had just called `switch_context`: `init_task_stack` lays
down a return address pointing at `task_trampoline` and zeroed callee-saved slots
(smuggling the entry function in via the `r15` slot). When first switched to, the
`ret` lands in the trampoline, which enables interrupts and calls the entry.
## Two ways to switch, one flag discipline
`schedule()` — pick the best task and switch — runs from two places:
- **`yield()`** — a task voluntarily gives up the CPU.
- **`tick()`** — the 1000 Hz [timer](../device-driver-development-guide/device-interrupts.md) preempts the running
task. This is what lets a task that never yields still share the CPU.
The subtlety in mixing them is the **interrupt flag (IF)**. The rule: `switch_context`
is always entered with interrupts *disabled* — naturally so inside the timer ISR,
and explicitly (`cli`) in `yield`. Then every task ends up with interrupts enabled
again through whichever path resumes it:
- a task suspended in `yield` re-enables them (`sti`) right after `schedule` returns;
- a task suspended mid-ISR resumes through the interrupt return (`iretq`), which
restores the `RFLAGS` it had when it was preempted (IF set);
- a brand-new task enables them in the trampoline.
One related detail: the timer interrupt is **acknowledged (EOI) before** its handler
runs, so a handler that switches tasks and doesn't return promptly can't stall the
LAPIC from delivering the next tick.
## Priority selection, in O(1)
Ready tasks live in a **FIFO queue per priority level** (8 levels), plus a
**bitmap** with one bit per non-empty level. Picking the next task is: find the
highest set bit (one instruction), take the front of that level's queue. No list
walking, no scanning — the decision cost is constant regardless of how many tasks
exist, which is what a real-time scheduler needs.
- **Highest priority wins.** A ready high-priority task always runs before a
lower-priority one.
- **Round-robin within a level.** When a task is descheduled it goes to the *back*
of its level's queue, so equal-priority tasks share the CPU fairly.
## Affinity: pinning a task to a core
By default a task runs on **any** core — the ready queue above is global, and any
idle core pulls the highest-priority task from it (work-conserving; see
[smp.md](smp.md)). A task can instead be **pinned** to one core with
`spawnOn(entry, priority, cpu)`, giving it an *affinity*: it will only ever run
there, never migrating.
Mechanically, each core has its **own** pinned queue (same 8-level FIFO + bitmap)
alongside the global one. A pinned task is enqueued only into its core's pinned
queue; selection compares the top of the global queue and the running core's pinned
queue and takes the higher priority (still O(1) — two bit-scans and a compare), with
a pinned task winning an equal-priority tie so it can't be starved by global work.
Because every queue is mutated under the [big kernel lock](smp.md), one core enqueuing
into another core's pinned queue is safe.
This is the *explicit-affinity* model (no surprise migration mid-deadline), which is
the more real-time-predictable direction. `spawnOn` refuses to pin to an offline or
out-of-range core — it creates the task unpinned instead, so it still runs somewhere
rather than stranding in a queue no core services, and returns whether the pin took.
## Sleeping and the idle task
A task can **block** — give up the CPU until an event, rather than busy-wait
(busy-waiting is the enemy of a real-time system: it wastes cycles a
higher-priority task should get). The first form is time-based: **`sleep(ms)`**
marks the task blocked with a wake deadline and switches away. On every tick the
timer wakes any task whose deadline has passed (a bounded scan, so it stays
deterministic), which makes it ready again; the scheduler then runs it when its
priority comes up. `sleep` measures its deadline on the [calibrated
clock](../device-driver-development-guide/device-interrupts.md), so it's real time.
When *every* task is blocked, something still has to run — so there's an **idle
task** at the lowest priority that just `hlt`s until the next interrupt (see
[halting.md](halting.md)). Because it's always runnable, the scheduler always has a
task to pick, and the "nothing to run" case never arises.
## Event-based blocking
The other form of blocking is waiting for an **event** rather than a duration. A
**wait queue** is a set of tasks parked until something happens: `wait(wq)` blocks
the caller on it, `wake(wq)` moves the highest-priority waiter back to ready
(preempting if it now outranks the running task). A task links into a wait queue
through the same field the ready queues use — it's in exactly one queue at a time.
These are the primitives locks, semaphores and [IPC](../device-driver-development-guide/ipc.md) are built on.
Blocking safely needs **composable critical sections**. A blanket `cli`/`sti` pair
doesn't nest: an IPC channel that `cli`s and then calls `wait` would have `wait`'s
`sti` re-enable interrupts too early, mid-operation. So the blocking primitives use
`saveInterrupts` / `restoreInterrupts` — capture the interrupt flag, disable, and
later restore *only if it was set* — which nests correctly. The invariant that
makes it all work: `schedule()` is always entered with interrupts disabled, so a
task always resumes from a switch with interrupts disabled and can restore its
caller's state.
## Verifying it
Three tests (see [testing.md](../testing.md)) prove the guarantees:
- **`sched`** spawns three tasks that busy-loop *without ever yielding*. They all
make progress — which can only happen if the timer is **preempting** between them
and the context switch is correct (nothing yields voluntarily).
- **`priority`** (with preemption off, for determinism) spawns tasks at three
priorities; they run and exit **highest-priority first** — `[6, 4, 2]`.
- **`sleep`** blocks a task for 50 ms and checks the elapsed time on the clock — a
real block (the idle task runs meanwhile), not a busy-wait.
- **`event`** blocks a task on a wait queue; waking it (from another task) resumes
it, and since it's higher priority it preempts immediately.
## What's next (partly done since)
- **Priority inheritance** — still open. Tasks now do block on shared resources
(IPC rendezvous, the big kernel lock), and nothing yet bounds priority
inversion — a [real-time](../vision.md) requirement.
- **Task exit / a reaper** — done. A dying task goes on its core's reap list in a
`.reaping` state; the timer tick drains the list, frees the stack back to the
heap, and recycles the task-table slot.
- **Per-address-space tasks** — done. User processes each own an address space,
and the context switch reloads CR3 when the target's tables differ (see
[paging.md](paging.md)).
@@ -0,0 +1,310 @@
# Shared fate: whole-process death (plan)
**Status: implemented 2026-07-22 (branch shared-fate), M1–M4 all landed; leader
`thread_exit` → `-EPERM` as decided. One scope addition forced by M4: the
per-task DMA/shared-memory arena cursors moved to the per-space object (the
`shm-mapping-ref` test could not distinguish corruption-by-remap from
corruption-by-free while sibling threads overlapped the arena) — the same move
the mmap/MMIO cursors made in threading M7.**
[threading.md](threading.md) promises that a process dies *whole* — a fault in any
thread, or a kill, takes down every thread. The kernel doesn't do that yet: every
death path (`exit`, `thread_exit`, a ring-3 fault, `process_kill`) tears down
exactly one `Task`, and the address-space refcount keeps the space alive for the
siblings — so a faulting worker orphans its threads, which keep running in the
possibly-corrupted address space (`process.zig` `killCurrentProcess`;
`scheduler.zig` `exitUserLocked`). This plan closes that gap the way Linux, Windows,
and Fuchsia all did: **the process is the unit of fate; only a voluntary
`thread_exit` is per-thread.** The supervisor contract — one exit notification,
then `process_exit_reason` — is deliberately unchanged.
## The contract
| Event | Who dies | Reason the supervisor reads (leader's record) |
|---|---|---|
| CPU fault in **any** thread (recoverable vector) | the whole group | the fault class (`segmentation_fault`, …) |
| `process_kill` on **any member id** | the whole group | `killed` |
| `exit(code)` from **any** thread | the whole group | `exited` (0) / `aborted` (≠0) |
| `thread_exit` from a worker | that worker only | — (per-task record: `exited`) |
| `thread_exit` from the **leader** | nobody — refused, `-EPERM` (decided, below) | — |
| NMI, double fault, machine check | the core halts (unchanged) | — |
`exit` gaining group semantics is the `exit_group` lesson from Linux: the runtime's
main-return path calls `exit`, and a process whose main returned must not leave
workers running. `thread_exit` (what the worker trampoline calls) keeps today's
per-thread behavior, refcount and all.
**The leader-`thread_exit` rule (decided: refuse).** The syscall is reachable from
the leader even though the runtime never does it. Options weighed: (i) **refuse
with `-EPERM`** — cheapest and honest; the group ends only through
`exit`/fault/kill; (ii) escalate to `exit(0)` — Linux-flavored, but silently turns
a (buggy) library call into process death; (iii) a Linux-style zombie leader whose
slot survives until the group ends — the most faithful, and by far the most
machinery. **(i) chosen at sign-off**; rows and tests below follow it.
## Group identity: a leader id
Nothing on `Task` names a process today — threads are tied to their process only by
an equal `address_space`, their name is `"thread"`, and their `supervisor` is
whichever *task* spawned them (possibly another worker), so supervision links form a
chain, not a group. Scanning by `address_space` is also fragile during teardown,
because both death paths zero it.
So: **add `leader: u32` to `Task`** — the Linux tgid, in danos clothes.
`spawnProcessSupervised` sets `leader = own id`; `spawnThreadSupervised` copies the
*caller's* leader; kernel tasks keep `leader = 0`, which is never followed. The
leader id is exactly the id `system_spawn` returned to the supervisor, so the
outside world already speaks it. Group membership = equal `leader`, where a *live
member* means `state` ∈ {`.ready`, `.blocked`, `.running`} — the same filter
`taskByIdLocked` applies; `.reaping` corpses are excluded. While the field is being
introduced, add `leader` to `ProcessDescriptor` too (the ABI is private, so this is
cheap now and lets `process_enumerate` consumers group threads).
`process_kill` re-derives its authority through the leader — with the existing
guard order preserved: a kernel task (`address_space == 0`) is `-ESRCH` *before*
any leader resolution (the kernel test asserts exactly that). Then: resolve the
target, follow `target.leader`, require `leader.supervisor == caller`. The kill
capability becomes per-*process*, aimed at any member id, and the odd
today-behavior where a worker can be individually killed by its spawning thread
disappears. (Audited: nothing in-tree kills a worker tid or relies on
thread-supervisor kill semantics.)
## The group-dying latch
`AddressSpaceRef` — one per space, refcounted by its member tasks, recycled with a
full struct re-init — gains the group-death state:
```zig
dying: bool = false, // set by the first trigger; never cleared
group_reason: abi.ExitReason, // what the leader's record will say
exit_endpoint: ?*ipc.Endpoint, // the leader's counted ref, moved here
leader: u32,
```
The latch answers three attacks the red team confirmed against a latch-free
design:
- **The spawn gate.** A member already *inside* `thread_spawn` on another core
when the fan-out runs (it passed the syscall-entry `kill_pending` check, then
spun on the BKL) would otherwise complete the spawn after the fan-out's lock
hold ends — a fresh, uncondemned member that escapes the kill and, worse, holds
a space reference that keeps the group-death hook from ever firing. Fix:
`retainAddressSpace` (equivalently `spawnUserLocked`) **refuses a dying
space**; the in-flight `thread_spawn` fails with `-ESRCH` under the same lock
that would have created the member.
- **Concurrent triggers.** A second member faulting (or exiting) on another core
while the first fan-out runs must not re-run the fan-out, double-bump
`fault_kill_count`, or re-stamp reasons. Every kill path checks the latch first:
already dying → skip straight to `terminateCurrentLocked`, no stamp, no count.
First trigger wins, deterministically. `exit_reason` and `fault_kill_count`
writes move under the BKL as part of this.
- **Notification ownership.** The leader's `exit_endpoint` is a counted
birth-to-death reference dropped at notify time. The stamp pass **moves** that
reference onto the `AddressSpaceRef` and nulls `Task.exit_endpoint` in the same
hold, so the leader's own `releaseTaskResourcesLocked` sees null (no early
notify, no double drop); the group-death hook notifies and drops exactly once.
## The fan-out: `killGroupLocked`
One new function in `process.zig`, running under a **single BKL hold** (built from
the `*Locked` primitives — the lock is non-recursive, and `terminateCurrentLocked`
never returns, which forces the shape):
```
killGroupLocked(leader: u32, reason: ExitReason, trigger: ?*Task)
0. Latch: AddressSpaceRef.dying = true, stash {reason, leader,
leader's exit_endpoint (moved)}.
1. Stamp pass: the LEADER's exit_reason = reason — the leader's
record is the one the supervisor can read, so it carries the
group reason even when the trigger is a worker. The trigger
also keeps `reason` (its own record tells the truth); every
other live member gets .killed. All members get kill_pending.
Stamping precedes any teardown, because recordExitLocked
snapshots the reason first thing.
2. Reap pass, to fixpoint: reap every member in .ready or .blocked
via reapTaskLocked, re-reading Task.state each iteration — a
member's teardown can wake another member (-EPEER wakes, joiner
wakes), flipping it .blocked → .ready behind the scan cursor.
Terminates in ≤ one pass per member: the scrub calls in
releaseTaskResourcesLocked (abandonSenderLocked,
removeFromWaitQueueLocked, forgetIpcClientLocked,
killOwnedEndpointsLocked) run before destroy, so no wake path
holds a pointer to a reaped member.
3. Members .running on other cores stay condemned (kill_pending);
a condemned member dies at its next syscall entry, at its own
core's next tick while in user mode, or — once it blocks or is
preempted — at any core's next tick reap. There is no kill IPI.
(The entry check reads kill_pending unlocked; benign on
x86-TSO — a missed read is caught by the next delivery point —
but make the field atomic when touching it.)
4. If the current task is a member (fault, exit, in-group kill):
terminateCurrentLocked, last, because it switches away and the
reap paths free the kernel stack being stood on.
If the caller is outside the group (supervisor kill): return.
```
The invariants this preserves, each load-bearing today:
- **Only `.ready`/`.blocked` tasks are reaped synchronously.** A member running on
another core can only be condemned — it tears itself down after switching CR3
off the dying page tables (the stack it stands on is freed later by the reap
list), and its address-space reference protects the page tables its CR3 still
points at. Force-destroying the space under a running sibling is the one
unrecoverable mistake available here.
- **The refcount decides when the space dies.** Reaping N members drops N
references; the last drop — possibly on a condemned sibling's core, a tick
later — destroys the space. No path forces it.
- **`fault_kill_count` bumps once per group**, not per member (`fault-recovery`
asserts `== 1` exactly); the latch is what enforces this under racing faults.
- **Per-tid resource sweeps stay per-tid.** Each member's
`releaseTaskResourcesLocked` releases what *that tid* owns — claims, GSI/MSI
bindings, registered endpoints, handles. That keying is correct under shared
fate (and is today's hazard: a lone worker death already yanks its claims out
from under live siblings). A worker that *does* carry an `exit_endpoint` (the
ABI allows it; the runtime passes `no_cap`) keeps today's per-task posting at
its own teardown — only the leader's notification moves.
## When is the group dead? The notification
Today each task posts its own exit notification as the *last* step of its release,
so a supervisor observes a fully-released child. For a group that guarantee must
hold for the **whole group**: if the leader's notification fires while a condemned
sibling still runs on another core, the device manager can respawn the driver into
a claim conflict with a not-yet-dead sibling.
The clean fix falls out of the refcount: **the group is dead exactly when the
address space is destroyed.** `releaseAddressSpace`'s last-drop path calls a new
`group_exit_hook` (the scheduler already calls up through hooks —
`terminate_current_hook` — precisely to keep this layering), which:
1. **re-stamps the leader's exit record** with the stashed `group_reason` — the
record is written (again) at group-death time, so "the reason is recorded
before the notification posts" stays true and a supervisor can never be
notified and then read `-ESRCH` because the burst evicted an old record;
2. posts the leader's exit notification (and subscriber broadcast) from the
stashed endpoint, and drops that reference — exactly once.
Both `releaseAddressSpace` call sites (`exitUserLocked`, `destroyTaskLocked`) run
under the BKL, so the hook does too; its wakes are safe at both (verified). For a
single-threaded process the behavior is *externally indistinguishable* from
today — the order of notify vs. destroy inverts, but both sit inside one lock
hold, so no other core can observe the space destroyed but the notification
unposted, or vice versa. That sentence is the correctness argument; it is also the
first invariant to re-examine if the BKL is ever split, along with
`killGroupLocked`'s single-hold atomicity. (Hand-built spaces that were never
retained take the immediate-destroy path and are out of the hook's scope.)
Workers' `exit_subscribers` broadcasts still fire per task — the FAT server's
dead-client sweep is keyed by tid and needs those.
**Signals.** `signal_bind` is per-task and the service harness binds on the main
thread, so signals address the leader in practice; that stays. During a group
death, `process_signal` may return `0` (accepted by a condemned member — never
delivered, every delivery point kills first) or `-ESRCH` (member already reaped);
init's stop sequence already tolerates both, and its timer escalation to
`process_kill` covers the gap. `process_signal` follows `process_kill`'s
leader re-key for consistency.
## The shared-memory frame hazard
`dropSharedMemoryReference` frees a region's physical frames when the last *handle*
reference drops, but mappings die only with the address space. If the last handle
lived in a torn-down member while any task still has the region mapped, that task
holds a live mapping onto freed frames — and the red team showed this is **not**
group-specific: a plain `thread_exit` of the handle-holding thread, or a last-ref
drop by a task *outside* the dying group during the condemned window, hits the same
use-after-free.
So the fix is a property of the **object**, not the dropper: give
`SharedMemoryObject` a per-*mapping* reference — `shared_memory_map` (and create's
self-map) retains; each space's destruction releases. "Last reference" then means
*no handles and no mappings*, both hazard paths collapse into the existing
refcount, and no group-kill special case is needed at all.
## Deliberately unchanged
- Worker `thread_exit`: per-thread, full per-tid resource sweep, refcount drop.
- The condemned-but-running window: a member on another core can finish its
in-flight syscall and run user code for up to a tick before dying — identical to
today's single-task `process_kill` semantics ("prompt but asynchronous, like a
Unix signal"). A kill IPI would shrink it; it is not part of this plan.
- `thread_join` returns 0 for a killed thread; joiners inside a dying group are
woken and then reaped like any member.
- The `.reaping` state, reap lists, and stack reaper.
## Accepted limits (documented, not fixed here)
- **Notify-ring overflow**: a group death posts one subscriber badge per member
into 8-slot rings; a >8-member group can drop badges. Group size is bounded by
the 48-task table; today's largest production group is 2 (display) and
thread-test already reaches 5.
- **Exit-record ring pressure**: one 64-entry ring, one record per member — made
harmless for the supervisor by the hook's group-death re-stamp.
- **Enumerate shows a partial group** mid-death: reaped members vanish at once,
condemned members linger up to a tick (audited: no in-tree consumer
misbehaves; the `leader` field in `ProcessDescriptor` lets future consumers
group correctly).
- **Pre-existing reap race, not widened**: a task preempted *mid-syscall* is
`.ready` with `in_system_call = true`, and the tick's reap loop will reap it —
an existing hazard the fan-out inherits but must not add new instances of.
Filed to investigate separately.
- **Per-task DMA/shm cursors** — *fixed during M4 after all*: the
`shm-mapping-ref` test tripped the overlap (the sibling's churn regions mapped
over the worker's region), so both cursors moved to the `AddressSpaceRef`
like the mmap/MMIO cursors before them. The post-implementation review then
found the other half: the shm/DMA page-table walks and their pmm/heap calls
ran *outside* the big kernel lock — pre-existing, but fatal once siblings
were invited to race them (and a plausible root for the long-standing
intermittent AP ring-3 fault at the shm base). All three paths now follow
the mmap discipline: metadata and allocation under one hold, the map itself
per-page under brief holds.
- **Mapping-record slots are never recycled**: 16 per space, one per
`shared_memory_create`/`map`, freed only at space destruction (there is no
shm unmap). A long-lived compositor that churns surfaces will hit the cap;
the failure is a clean refused create, and slot recycling can ride whatever
adds `shared_memory_unmap`.
- **Two properties lack direct tests**: the spawn gate (an in-flight
`thread_spawn` racing the fan-out — inherently nondeterministic to arrange;
covered by code inspection and the `-ESRCH` path) and the `process_signal`
leader re-key (exercised only implicitly by the signals case).
## Milestones
- **M1 — the leader id.** `Task.leader` (kernel tasks: 0, never followed), set on
both spawn paths; `leader` added to `ProcessDescriptor`; `process_kill` and
`process_signal` re-keyed (kernel-task `-ESRCH` guard *before* leader
resolution). No fan-out yet. Existing tests must pass untouched.
- **M2 — the latch + fan-out.** `AddressSpaceRef.dying` + stash;
`retainAddressSpace` refuses dying spaces; `killGroupLocked`; wire the fault
path, `exit`, and `process_kill` into it; leader `thread_exit` → `-EPERM`;
`exit_reason`/`fault_kill_count` writes under the BKL; `kill_pending` made
atomic. Group notification via the `group_exit_hook` re-stamp + post.
- **M3 — shared-memory mapping refs.** `SharedMemoryObject` counts mappings;
space destruction releases them; frames free only at zero handles *and* zero
mappings.
- **M4 — tests + docs.** New `-Dtest-case`s (all `smp: 4` where cross-core
matters), driving `thread-test` with new argv modes:
- `thread-fault-group`: a worker faults; assert both tasks gone from
`enumerate`, `fault_kill_count == 1`, `process_exit_reason(leader) ==
segmentation_fault`, address-space and stack-bytes counters return to base.
- `kill-threaded-group`: `process_kill(leader)` with a worker spinning on
another core; assert the worker dies by the deferred path, exactly one exit
badge, delivered only after both members are dead, and
`process_exit_reason(leader) == .killed`.
- `kill-via-worker-tid`: `process_kill(worker)` kills the whole group;
`-EPERM` for a non-supervisor aiming at the worker.
- `racing-triggers`: two members fault/exit simultaneously on different cores;
assert a deterministic leader reason and `fault_kill_count == 1`.
- `exit-group`: a *worker* calls `exit(3)`; group dies, leader reason
`.aborted`.
- `leader-thread-exit`: asserts the chosen rule (`-EPERM`, workers unaffected).
- `thread-exit-solo`: regression — worker `thread_exit` still leaves siblings
running.
- `group-claim-release`: a member claims a device; assert the claim is free and
the leader notification arrives only after every member is dead.
- `shm-mapping-ref`: last handle dropped by a dying thread; sibling's mapping
stays valid until space death (M3 regression).
Then update [threading.md](threading.md) (the shared-fate gap note),
[process-lifecycle.md](process-lifecycle.md),
[process-management.md](process-management.md), and
[ipc.md](../device-driver-development-guide/ipc.md)/[drivers.md](../device-driver-development-guide/drivers.md) mentions.
+275
View File
@@ -0,0 +1,275 @@
# SMP: multiple cores, the microkernel way
A design/research note that predates the build — danos now runs on **multiple
cores** by default (see [Implementation status](#implementation-status) and
[scheduling.md](scheduling.md)). This maps how microkernels — especially the L4
family and seL4 — handle **symmetric multiprocessing (SMP)**, the plan and
reading list the port followed. It also flags where those choices depend on
whether danos is chasing **real-time** or **resilience** (see the note at the end).
## First, the vocabulary
Three independent things people conflate (see [scheduling.md](scheduling.md) for the
danos specifics):
- **Task capacity** (`max_tasks`) — how many tasks can *exist*. A table size.
- **Cores** — how many tasks *run at the same instant*. One running task per core.
- **Real-time** — whether timing is *predictable*. Comes from bounded operations
(our O(1) scheduler), not from core count.
danos now runs on multiple cores. The firmware starts only the **bootstrap processor
(BSP)**; the kernel wakes the other cores (**application processors**, APs) with
INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor tables,
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
the `smp` self-test confirms worker tasks executing on all four cores at once under
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
and thread-to-core affinity (see [Implementation status](#implementation-status)).
## The common microkernel instinct: don't share kernel state
Monolithic kernels (Linux) share large amounts of state across cores behind many
fine-grained locks. Microkernels lean the other way — the kernel does *little* (IPC,
scheduling, capabilities), so the pressure is to make kernel state **per-core** and
coordinate cores with **inter-processor interrupts (IPIs) or messages** rather than
shared, locked data structures. Two hallmarks follow:
- **Thread-to-core affinity.** Threads are usually *bound* to a core; migration is an
explicit operation, not automatic load-balancing.
- **Policy in user space.** *Which* core a thread runs on tends to be a user-level
decision (a scheduler/manager server); the kernel just provides the mechanism to
run it there and to signal across cores. This matches the microkernel creed:
mechanism in the kernel, policy outside.
Within that instinct, the L4 family split on **how much to lock**.
## seL4: the "big kernel lock" — and why it's not a hack
seL4's choice is striking: a **single big kernel lock (BKL)**. Only one core runs
*kernel* code at a time; **user code runs fully in parallel** on all cores. A core
that traps into the kernel takes the global lock, does its (short) work, releases it.
Why so coarse? **Formal verification.** seL4's whole value is a machine-checked
correctness proof, built for a *uniprocessor* kernel — reasoning about one thread of
kernel execution. Fine-grained SMP locking explodes the interleavings you'd have to
reason about. The big lock **serialises kernel execution so the single-core reasoning
still holds**. It trades kernel scalability for verifiability.
And it works better than it sounds, *because the kernel does so little*: the lock is
held for short, bounded intervals, while the real work (drivers, services) runs in
user space in parallel, outside the lock. Scheduling is otherwise **per-core** (each
core its own ready queues), threads carry an **affinity**, and cross-core IPC costs an
IPI.
> Nuance: the fully *verified* seL4 configuration is the uniprocessor one. The
> SMP/big-lock version isn't covered by the same end-to-end proof — extending
> verification to multicore has been ongoing research. So the big lock is partly
> "stay close to the thing we proved."
seL4 also layers **MCS** (mixed-criticality scheduling) on top: **scheduling
contexts** carrying a time *budget* and *period*, so a thread can't overrun its share
— temporal isolation, reasoned about per core. This is the seriously real-time part.
## Fiasco.OC / NOVA: per-CPU, finer-grained
Not all L4s took the big lock. **Fiasco.OC** (TU Dresden L4, part of L4Re) is
**per-CPU**: per-CPU run queues, CPU-local kernel objects, IPIs for the rare
cross-CPU operations. Threads bind to a CPU; moving one is explicit. Scales better
than a big lock, more complex, and without seL4's verification constraint forcing the
issue. **NOVA** (a microhypervisor) is similarly per-CPU. The shared pattern: make
everything CPU-local you can, and when cores must interact, **send a message/IPI**
instead of touching shared data.
## The extreme: the "multikernel"
Taken to its logical end you get **Barrelfish** (ETH Zurich): treat a multicore
machine as a *network of cores*, each running its **own kernel instance**, sharing
**no** kernel memory, communicating **only by message passing** — the microkernel's
IPC philosophy applied to the kernel's own structure. seL4's "clustered multikernel"
explorations use the same idea: groups of cores, each cluster a big-lock domain,
clusters talking by messages. The insight: if you're already committed to messages
for user-space isolation, structure the kernel across cores the same way and sidestep
shared-memory locking entirely.
## Does the right choice depend on real-time vs resilience?
Yes — and this is the branch that matters for danos right now.
- **If the goal is hard real-time:** favour **per-core scheduling with fixed
affinity**. A thread never gets surprise-migrated mid-deadline, and each core's
timeline can be reasoned about in isolation. Global load-balancing (Linux's default)
is great for throughput and *bad* for determinism, which is why RT microkernels
mostly pin threads. seL4's MCS scheduling contexts are the reference model.
- **If the goal is resilience / restartability:** the SMP priority shifts to **fault
isolation and recovery**, not timing. What matters is that a failed component (a
driver, a service) on any core can be **killed and restarted** without taking the
system down — which is a property of address-space isolation + a supervising
restart server (below), *largely orthogonal to how cores are scheduled*. A big lock
is perfectly fine here; you're optimising for "a crash is contained and
recoverable," not "latency is bounded to N µs."
- **If the goal is throughput:** you'd care about lock contention and per-core
queues — the least microkernel-flavoured of the three.
These pull in different directions, so **picking the primary goal comes before
picking the SMP design.** (danos's founding assumption was real-time; that's under
active reconsideration in favour of resilience — see [vision.md](../vision.md).)
## What this would mean for danos
Whatever the top goal, the *sequence* is the same and seL4 validates starting simple:
1. **Enumerate cores** — needs [device discovery](discovery.md) (ACPI MADT on x86,
device tree on ARM). SMP is a concrete consumer of that work. **Done on x86:** the
MADT parse records every usable Local APIC — with the `apic_id` an AP wake targets —
and `platform.cpus()` returns the list (see [discovery.md](discovery.md)). The boot
log reports the count; the ARM (device-tree) path still needs it.
2. **Wake the APs** — INIT–SIPI–SIPI on x86; PSCI/spin-tables on ARM. Each core brings
up its own tables, timer, and idle task. **Done on x86** — cores climb to long mode,
set up their own GDT/TSS, and enter the scheduler; tasks run in parallel across all
cores ([status](#implementation-status)).
3. **Start with a big kernel lock.** It's a legitimate first design, not a shortcut —
philosophically aligned with a tiny kernel, and it lets the single-core correctness
model you already have (the interrupt-flag discipline in
[scheduling.md](scheduling.md)) stay largely intact: one lock around kernel entry
instead of rethinking every critical section. **Done** — see
`system/kernel/sync.zig`.
4. **Later, if contention bites,** evolve toward **per-core run queues + explicit
affinity** (the Fiasco.OC direction) — also the more real-time-predictable model.
5. **Placement stays a user-space policy** — the kernel runs a thread on the core it's
told to, a user-level manager decides which.
Big-lock-first → per-core-later. The affinity/MCS depth is only worth it if real-time
turns out to be the actual goal.
## Implementation status
The "wake + schedule" build (real parallel task execution) is going in as a sequence
of green checkpoints — each step keeps the single-core test suite passing before the
next lands.
**Done:**
- **Core enumeration** — the MADT parse records every usable Local APIC (with its
`apic_id`, which an AP wake targets); `platform.cpus()` returns the list. See
[discovery.md](discovery.md).
- **The big kernel lock** (`system/kernel/sync.zig`) — one coarse spinlock guarding the
scheduler queues and IPC, always held with local interrupts disabled. It is held
*across* a context switch and released by whichever task resumes (the hand-off
rule); `task_trampoline` releases it for a freshly-spawned task. `scheduler.zig` and
`ipc.zig` run every critical section under it. Uncontended on one core, so behaviour
is identical to the old interrupt-flag model.
- **Per-CPU state** — a `PerCpu` struct (running task, idle task, APIC id) per core,
its pointer kept in the x86 **GS base** (`IA32_GS_BASE`). No `swapgs` was needed at
the time — there was no user mode yet; with ring 3 in place, every ring transition
now swaps it against the user's own GS base under the `swapgs` discipline (see
`system/kernel/architecture/x86_64/per-cpu.zig`), and each AP enables the fast
system-call path (`initSystemCall`) for itself at bring-up. The old global
`current` is now `thisCpu().current`. The ready
queues stay **global** under the lock — work-conserving, so any idle core will pull
the highest-priority ready task; per-core queues are a later optimisation.
- **AP wake to long mode** — `architecture.startSecondary` drives INIT–SIPI–SIPI (via the
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
real mode at a low page and runs the [trampoline](../../system/kernel/architecture/x86_64/trampoline.s)
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
all four cores report `online`.
The trampoline earns its complexity from four hardware facts:
- a STARTUP IPI vectors a core to physical `vector << 12` (a *byte* vector), so the
trampoline must live **below 1 MiB** — the kernel reserves that page from the frame
allocator at boot, before paging/heap draw down the scarce low frames;
- the blanket RAM identity map is **NX** (W^X), but the AP fetches the trampoline
from it under paging, so that one page is made executable for bring-up;
- the blob is copied to a page whose address isn't known at link time, so it is
**position-independent**: it derives its own base from `CS` and, crucially,
addresses data *segment-relative in real mode* (where the segment base already
supplies the page base) but *base-register-relative in protected/long mode* (flat
segments, base 0). Getting that distinction wrong was the first bug found;
- an AP starts with a bare `CR0`/`CR4`, but the kernel is built **with SSE** (the
x86_64 baseline) and the compiler emits SSE for things as ordinary as a struct
copy — so the trampoline must set `CR4.OSFXSR`/`OSXMMEXCPT` and fix `CR0.EM`/`MP`,
or the first SSE instruction on the AP `#UD`s. The BSP inherited those bits from
UEFI; the AP has to set them itself. This was the second bug — it masqueraded as a
fault in `lgdt` (the first kernel code after entry that the compiler vectorised).
- **Per-core tables + scheduler entry** — each AP loads **its own GDT** (with its own
TSS descriptor) and **its own TSS** (its own IST/`rsp0` stack), loads the shared
IDT, enables its LAPIC and timer, then calls the generic `secondaryMain`: it turns
its bring-up context into the core's idle task (as task 0 is for the BSP), marks the
core online, and enters the run loop. With interrupts on, each core's own timer tick
preempts its idle context into whatever the global ready queue offers — so all cores
pull real work in parallel. The `smp` test spawns CPU-bound workers and confirms they
execute on all four cores at once, and `fault-ap-df` pins a #DF to an AP and checks
that core catches it on **its own** IST (a broken per-core TSS would triple-fault) —
reported as "core N: …", so a fault is always attributed to the core it happened on,
and is contained to that core (the rest of the system keeps running).
- **Thread affinity** — `spawnOn(entry, priority, cpu)` pins a task to a core (its own
per-core pinned queue, merged with the global queue at selection; see
[scheduling.md](scheduling.md#affinity-pinning-a-task-to-a-core)). The `affinity`
test confirms a pinned task never migrates. This is the mechanism the fault-on-AP
test rides on, and the *explicit-affinity* real-time-predictable model.
- **Right-sized footprint** — the per-CPU ceiling (`parameters.maximum_cpus`, one
constant shared by discovery, the scheduler, and the per-core GDT/TSS) is generous
(128), but
the *large* per-core resources — the kernel and IST (double-fault) stacks — are
**heap-allocated at bring-up**, only for cores that actually come online. Only the
BSP's IST stack is static, because it must exist before the frame allocator does.
This kept the kernel image small (a static `[128][16 KiB]` IST array would have been
2 MiB of `.bss`); it's a few tens of KiB instead.
- **`single_threaded` off** — the kernel was built `single_threaded = true`, which
compiles `std.atomic` down to plain non-atomic ops. Harmless on one core, but it
quietly breaks the big kernel lock across cores; it's now `false`.
- **Re-armable wake + retry** — the trampoline frame is reserved for the system's
life, but kept **inert between wakes**: zeroed and non-executable, armed (blob
copied in, page made executable) only for the moment a core is actually climbing,
then disarmed again. So there's never a dormant executable page, and a core can be
(re)woken at any time — `architecture.startSecondary` is one self-contained attempt (arm →
INIT–SIPI–SIPI → disarm), and its `INIT` resets a wedged core, so retrying just
works. Boot retries a non-responding core up to three times; the same primitive is
the groundwork a future **power manager** would drive to bring cores up (and,
eventually, its counterpart to take them offline — which additionally needs the
core's tasks migrated off first).
**Next (refinement, not first-light):**
- **IPIs** — cross-core wake/preempt. Not needed for correctness: an idle core wakes
on its own timer tick and pulls ready work then; IPIs only cut that latency from
≤1 ms to near-instant.
- **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock
contention ever bites. (Thread *affinity* already exists — see above; this is the
further step of giving each core its own primary run queue for load distribution.)
- **Fault recovery** — today a fault halts (only) the faulting core. Turning that into
"kill the task, keep the core running" is the [resilience](resilience.md) track (it
needs the task's lock/resource state handled), and for taking a core fully offline,
its tasks migrated first.
## Further reading
**Microkernel SMP & scheduling**
- Klein et al., *"seL4: Formal Verification of an OS Kernel"* (SOSP 2009) — the
verification that shapes seL4's whole SMP stance.
- Lyons et al., *"Scheduling-Context Capabilities: A Principled, Light-Weight OS
Mechanism for Managing Time"* (EuroSys 2018) — seL4 MCS, the real-time model.
- The **seL4 whitepaper** and "towards a verified multiprocessor seL4" material — the
big-lock / clustered-multikernel reasoning.
- **Fiasco.OC / L4Re** documentation (TU Dresden) — the per-CPU alternative.
- Baumann et al., *"The Multikernel: A New OS Architecture for Scalable Multicore
Systems"* (SOSP 2009) — Barrelfish, the share-nothing extreme.
**Resilience / self-healing (if that's the real goal)**
- Herder et al., *"Fault Isolation for Device Drivers"* and the **MINIX 3**
*reincarnation server* — a driver crashes, a supervisor restarts it live. The
closest existing system to "re-initialise parts of the OS."
- **QNX** — commercial microkernel RTOS built on message passing and restartable
drivers; good study of the combination.
- **Erlang/OTP** *supervision trees* and the *"let it crash"* philosophy — not a
kernel, but the canonical design for "isolate failures and restart the failed
part," directly relevant to danos's restartability motivation.
## Related
- [scheduling.md](scheduling.md) — the single-core scheduler SMP would extend.
- [discovery.md](discovery.md) — enumerating cores is a device-discovery problem.
- [ipc.md](../device-driver-development-guide/ipc.md) — the message passing cross-core coordination rides on.
- [vision.md](../vision.md) — the goals question (real-time vs resilience) this note
keeps bumping into.
+95
View File
@@ -0,0 +1,95 @@
# System Calls
System calls (syscalls) are the bridge between your programs and the operating system's restricted core (kernel).
> **Status:** danos has real user processes. User programs enter the kernel
> via the `syscall` instruction (STAR/LSTAR/SFMASK set per core; the entry stub in
> `isr.s` does the `swapgs` + kernel-stack switch and reuses the interrupt
> dispatcher); the `int 0x80` gate is kept alongside as a minimal test path.
> The live table is `system/abi.zig` (private, renumberable — see
> [vdso.md](vdso.md) for the public boundary): process lifecycle + threads,
> memory (mmap/dma/shared-memory), synchronous + async IPC with capability passing,
> device access, time, the tagged-log diagnostics (`debug_write` with a level,
> `klog_read`/`klog_status`), and filesystem NAMING (`fs_resolve`/`fs_node`/
> `fs_mount`/`fs_unmount` — the kernel VFS root routes paths and serves the
> read-only /system initrd mount; file DATA stays with userspace filesystem
> servers over the vfs-protocol, docs/vfs-protocol.md).
## The Mechanism of a Syscall
A system call follows a highly orchestrated, 6-step sequence to ensure hardware safety and process security:
1. Setting the Arguments: The program places a unique ID corresponding to the requested service (e.g., \(sys\_write\)) into a specific CPU register (like RAX in x86_64) along with its required parameters.
2. Executing the Trap: The program triggers a special CPU instruction, such as syscall or int 0x80. This acts as a software interrupt.
3. Mode Switch: The CPU atomically flips its execution privilege from unprivileged User Mode (Ring 3) to the highly privileged Kernel Mode (Ring 0).
4. Lookup and Execution: The kernel looks up the syscall ID in a dispatch table and runs the designated internal routine to perform the actual work (like fetching data from the hard drive).
5. Returning the Status: The kernel places the result of the operation—or an error code—back into the RAX register.
6. Return to User Mode: The kernel executes a return instruction (such as sysret), the CPU shifts back to User Mode, and your program continues executing.
## Approaches
There are two approaches to system calls, stable public ABI and unstable private ABI. A public ABI uses a standard defined set of numbers. It makes it easy to guess, easy to write compilers and tools for. private ABI tend to change the numbers to obscure the numbers, preventing attackers from bypassing libc using CPU instructions directly. Private ABIs force users to use a library like libSystem, which the kernel can inject the syscall numbers. An alternative solution is to just have 1 system number but make it generic by the address to a struct in memory.
Practical minimum primitives (Monolithic kernel)
| Category | Primitive | Detail |
|----------|--------------|---------------------------------------------------------------------------------------|
| lifecyle | spawn / exec | Loads a binary from storage into memory and begins execution. |
| lifecyle | exit | Terminates the current process and frees its memory back to the kernel. |
| I/O | read | Requests data from a hardware device or file descriptor into user memory. |
| I/O | write | Pushes data from user memory out to a device or file descriptor (like a screen). |
| Memory | brk / mmap | Requests the kernel to allocate more physical or virtual memory pages to the process. |
| Control | ioctl | A catch-all "escape hatch" call to send hardware-specific commands to device drivers. |
The Microkernel Minimum Set
Everything else---including`read()`,`write()`,`malloc()`, and`fork()`---will run in user space as servers (e.g., a VFS server, a memory manager server) that threads communicate with using these three calls:
1. **`IPC_Call(endpoint, message_buffer)`(Synchronous Send + Receive)**
- **What it does:**The calling thread sends a message block to a service endpoint and immediately blocks (sleeps) until that service processes the request and sends a reply back.
- **Why it's minimal:**Combining*Send*and*Receive*into a single atomic atomic system call eliminates the need for separate tracking and prevents a massive amount of context-switching overhead. This is the foundation of high-performance microkernels like[seL4](https://sel4.systems/)and L4.[[1](https://en.wikipedia.org/wiki/L4_microkernel_family),[2](https://www.microchip.com/en-us/products/microprocessors/64-bit-mpus/pic64hx/ecosystem)]
2. **`IPC_ReplyWait(endpoint, reply_buffer)`(Respond + Wait for Next)**
- **What it does:**Used strictly by your background user-space servers (like your disk driver or filesystem). It sends a reply to the last client that called it, and immediately puts the server to sleep until the next request arrives.[[1](https://news.ycombinator.com/item?id=33078441)]
3. **`Yield()`/`Thread_Ctrl()`**
- **What it does:**Allows a thread to voluntarily give up its CPU time slice, or allows a root task to spawn/kill threads.
4. **`ipc_send(endpoint, message_buffer)`(Asynchronous Send)**
- **What it does:**Posts a small payload to an endpoint's bounded queue and returns *without* blocking — no rendezvous, no reply. The receiver picks it up through the same `IPC_ReplyWait`, as a buffered message. It is the async counterpart of `IPC_Call`, for one-to-many broadcasts where a synchronous rendezvous would let one dead or slow receiver hang the sender. The [input service](../device-driver-development-guide/input.md) — keyboard-event fan-out — is its first user. A full queue drops the oldest message (a buffered message is discrete data, unlike a coalescing interrupt notification).
* * * * *
Hardware Implementation: x86_64 vs. aarch64
Because you are targeting both platforms, you must design a clean**Architecture Abstraction Layer (AAL)**. Each architecture uses completely different assembly instructions, CPU registers, and privilege levels to jump from user space (Ring 3 / EL0) to kernel space (Ring 0 / EL1).[[1](https://android.googlesource.com/kernel/common/+/84d3e59750bbd/arch/arm64/Kconfig),[2](https://blog.codingconfessions.com/p/making-system-calls-in-x86-64-assembly),[3](https://alex.dzyoba.com/blog/os-segmentation/),[4](https://dev.to/ripan030/linux-kernel-interrupt-handling-part-2-fe1)]
Here is how you will map your bare minimum system calls on both architectures:
1\. x86_64 Implementation
On 64-bit Intel and AMD processors, you ignore the old`int 0x80`software interrupts. Instead, you use the high-performance`syscall`and`sysret`instructions.[[1](https://alex.dzyoba.com/blog/os-segmentation/)]
- **The Trap:**The user-space program executes the`syscall`instruction.[[1](https://namastedev.com/blog/kernel-vs-user-space-2/)]
- **The Registers:**The hardware automatically moves the instruction pointer, but you must pass your arguments in specific registers. A common microkernel convention mimics the System V AMD64 ABI:
- `rax`: System Call ID (e.g.,`0`for IPC_Call,`1`for IPC_ReplyWait)
- `rdi`: Argument 1 (Endpoint ID / Destination)
- `rsi`: Argument 2 (Pointer to the message payload buffer)
- `rdx`: Argument 3 (Size of the message)[[1](https://dev.to/kaamkiya/hello-world-in-assembly-x86-64-2kb8)]
- **Kernel Setup:**Your kernel must configure the Model Specific Registers (MSRs)---specifically`IA32_STAR`and`IA32_LSTAR`---during boot to point to your kernel's system call entry assembly code.[[1](https://nfil.dev/kernel/rust/coding/rust-kernel-to-userspace-and-back/),[2](https://johannst.github.io/notes/arch/x86_64.html)]
2\. aarch64 (ARM 64-bit) Implementation
On ARMv8-A and ARMv9-A architectures, privilege levels are called Exception Levels. User space runs at**EL0**, and your microkernel runs at**EL1**.[[1](https://community.nxp.com/pwmxy87654/attachments/pwmxy87654/imx-processors/183079/1/AN12212.pdf),[2](https://developer.arm.com/-/media/Arm%20Developer%20Community/PDF/Learn%20the%20Architecture/Exception%20model.pdf?revision=a62f2bf2-b08a-4a4f-8cbe-38c67ddf4434),[3](https://people.kernel.org/linusw/),[4](https://drewdevault.com/blog/Helios-aarch64/)]
- **The Trap:**The user-space program executes the`svc #0`(Supervisor Call) instruction.[[1](https://hackmd.io/@xlYUTygoRkyuQQlwXuWDWQ/SJWQuIsIZe)]
- **The Registers:**ARM provides a clean, plentiful register set. You typically pass your arguments in the standard parameter registers:
- `x0`: System Call ID
- `x1`: Argument 1 (Endpoint ID / Destination)
- `x2`: Argument 2 (Pointer to message payload buffer)
- `x3`: Argument 3 (Size of the message)[[1](https://github.com/lelegard/arm-cpusysregs),[2](https://medium.com/@vincentcorbee/http-server-in-arm64-assembly-apple-silicon-m1-077a55bbe9ca)]
- **Kernel Setup:**Your kernel must set up an Exception Vector Table and write its base address to the`VBAR_EL1`register. When`svc`is executed, the CPU jumps to the synchronous exception offset in that table.[[1](https://dev.to/ripan030/linux-kernel-interrupt-handling-part-2-fe1),[2](https://eastrivervillage.com/blog/archive/2018/06/),[3](https://www.wadixtech.com/blog/armv8-a-exception-levels-el0-to-el3)]
* * * * *
Managing the Payload Challenge
Because it is a microkernel, performance lives or dies by how fast your`IPC_Call`can move data from Client to Server. You have two minimal choices for handling the`message_buffer`pointer:[[1](https://anazimzada2020.medium.com/microkernel-architectural-pattern-5e4e9184170e)]
- **The Copy Method (Simplest to start):**Your kernel pauses the client, reads the data from the client's memory space, switches page tables to the server, and copies the data into the server's buffer.
- **The Shared Memory Method (Fastest):**The kernel sets up a temporary, shared virtual memory page between the client and server. The client writes to it, calls`syscall`/`svc`, and the server reads it instantly without the kernel copying any bytes
+131
View File
@@ -0,0 +1,131 @@
# system.img — the boot capsule
## What it is
`boot/system.img` is the **boot capsule**: every bundled user binary — init, the
services, the drivers, the test programs — packed into **one file** on the boot
volume. It is not a filesystem image and it is not compressed; it is exactly the
kernel's **initial-ramdisk wire format** (`system/initial-ramdisk.zig`, format
v2), written to disk ahead of time. The EFI loader reads it in a single
sequential pass and hands the bytes to the kernel unmodified.
The capsule is a *performance artifact*, not a source of truth. The boot
volume's `/system` and `/test` file trees remain the canonical layout (see
[danos-file-system-hierarchy-FSH.md](../file-system-development/danos-file-system-hierarchy-FSH.md));
the capsule is a pre-baked snapshot of the same binaries, derived from the same
build graph, so the running system is identical whether the loader read the
capsule or walked the tree.
## Why it exists
Firmware file I/O has exactly one fast shape: **one open + one sequential
read**. Everything else is a lottery. Loading the system per-file — dozens of
opens, seeks, and short reads through the firmware's FAT driver — measured
**minutes** on real hardware, against milliseconds in QEMU/OVMF. Packing the
binaries into a single file turns the whole of user space into the shape
firmware is good at.
Because the capsule already *is* the ramdisk wire format, the loader doesn't
even repack it: `loadCapsule` (`boot/efi.zig`) validates the magic and passes
the buffer straight through as `BootInformation.initial_ramdisk_base`/`len`.
## The format
The container is deliberately trivial — danos owns both producer and consumer,
so it need be no fancier. Little-endian throughout:
```
Header magic: u32 = "DNR2" (0x32524E44), count: u32
Entry × count name: [64]u8 (NUL-padded FHS path), offset: u64, len: u64
blobs... each entry's file bytes, at its offset within the image
```
- **Names are full FHS paths** (`/system/services/init`), not basenames — that
is what "v2" means. The 64-byte capacity matches `abi.maximum_process_name`,
so a task named after its binary path is never truncated. Paths longer than
63 bytes are a build error (`pack-system-image.py` rejects them).
- **The v1 magic (`"DNRD"`, basename entries) is rejected**, not tolerated: a
stale image should fail loudly at `Reader.init`, not misparse names.
- `initial_ramdisk.Reader` is the one validated view over the bytes — magic
check, table bounds, per-blob bounds — used by the kernel and shared with the
loader. `Reader.find` resolves a binary by exact path first, then by unique
basename, ASCII case-insensitively (the entries come from a FAT volume, whose
name lookups are case-insensitive by definition).
## How it is built
`build.zig` maintains one `bundled` list — every user binary and its FHS home.
Three artifacts are derived from that same list, in the same build graph, so
they cannot drift apart:
1. **The tree**: each binary installed at its FHS path (`zig-out/system/...`
and `zig-out/test/...`, mirrored onto the FAT boot volume by
`tools/make-fat-image.py`).
2. **The manifest** (`system/manifest`): the FHS path of every bundled binary,
one per line — the loader's per-file fallback input.
3. **The capsule**: `tools/pack-system-image.py` packs the same binaries into
the v2 container, installed at `zig-out/boot/system.img` and placed on the
boot volume at `boot/system.img`.
Note what the capsule does *not* contain: the kernel (`system/kernel` is loaded
separately by `loadKernel`, as an ELF) and the EFI loader itself. It is user
space only.
## How it is loaded
`loadSystemTree` (`boot/efi.zig`) tries three strategies, most portable first —
the running system cannot tell which one ran, because all three produce the
same in-RAM ramdisk image:
1. **The capsule** — open `boot\system.img`, read it whole, check the magic,
hand it over as-is. The normal path on any build-produced volume.
2. **The manifest** — read `system\manifest` and open each listed path *by
name*. FAT name lookup is case-insensitive and firmware-portable, unlike
directory enumeration. The loader assembles the v2 image in RAM itself.
3. **The tree walk** — enumerate `/system` and `/test` recursively (`/test`
is optional: a stick without fixtures still boots). Last resort for
hand-assembled sticks with neither file: some firmware FAT drivers return
bare 8.3 names uppercase from enumeration, which is why this is the
fallback and not the primary path.
All three are best-effort: a **kernel-only volume still boots** — the kernel
just has no user binaries to spawn and reports the absence.
One operational consequence of the ordering: the capsule *shadows* the tree.
If you hand-edit binaries on a stick that also carries a `boot/system.img`,
your edits are invisible — the loader boots the capsule's snapshot. Delete
`boot/system.img` from the volume to force the manifest/tree path.
## What the kernel does with it
The loader records the image's physical base and length in `BootInformation`;
the kernel (`kernel.zig`) then publishes the same bytes twice, to two
consumers:
- **The process layer** (`process.zig`): `system_spawn` looks binaries up in
the ramdisk via `Reader.find` — exact FHS path, or unique basename for
pre-path callers — and loads them as fresh ring-3 processes. The stored path
becomes the task's name.
- **The VFS root** (`vfs.zig`, `setInitialRamdisk`): the image is mounted as
kernel-backed, read-only mounts — one per top-level tree named by the entry
paths, so `/system` and, when the fixtures are bundled, `/test`. Directory
nodes are derived from the entry paths (the unique parents), so the trees
are listable and their files readable over the normal VFS protocol — the
FHS boot tree every process sees comes straight out of the capsule bytes.
The image is never copied after the handoff and never mutated: the initrd is
immutable, which is what makes the VFS's node serving lock-free.
## What it is not
- **Not `danos-usb.img`.** That is the 64 MiB FAT32 *boot volume* built by
`tools/make-fat-image.py` — the thing a machine actually boots, which
*contains* `boot/system.img` alongside the loader, kernel, manifest, and
tree. See [efi.md](efi.md) and [release-iso.md](release-iso.md).
- **Not a mountable filesystem.** No FAT, no block device, no driver — just a
header, a table, and concatenated blobs, parsed by ~90 lines of
`initial-ramdisk.zig`.
- **Not required.** It is the fast path, with two slower equivalents behind
it.
- **Not a place where state lives.** It is regenerated on every build from the
bundled binaries; nothing writes to it, at build time or runtime.
+122
View File
@@ -0,0 +1,122 @@
# SysV: the kernel's calling convention
Several places in danos say "the kernel is SysV" — most visibly `system/boot-handoff.zig`:
```zig
pub const kernel_abi: std.builtin.CallingConvention = .{ .x86_64_sysv = .{} };
```
**SysV** is short for the **System V AMD64 ABI**, the calling convention that
Unix-like systems (Linux, the BSDs, macOS) use on x86-64. This page explains what
that means and why danos has to pin it explicitly.
## What a calling convention is
At the machine level there's no language keeping two functions honest when one
calls the other — just registers and a stack. So there has to be a shared
agreement on the mechanics:
- which registers carry the **arguments**, and in what order,
- where the **return value** goes,
- which registers the callee must **preserve** versus may freely clobber,
- the **stack alignment** required at a call,
- how larger things (structs, floats, varargs) are passed.
That agreement is the calling convention. Both sides of a call must be compiled to
the *same* one, or they read arguments out of the wrong registers and get garbage.
("ABI" — Application Binary Interface — is the broader term, also covering type
sizes and object-file format; here we mean the calling-convention part.)
## What SysV specifies (the parts that matter here)
Integer and pointer arguments go in this register sequence:
| arg | 1 | 2 | 3 | 4 | 5 | 6 |
|------|---------|-----|-----|-----|----|----|
| reg | **RDI** | RSI | RDX | RCX | R8 | R9 |
The return value comes back in **RAX**. RBX, RBP and R12–R15 are **callee-saved**
(a function must restore them before returning); the rest are caller-saved. The
stack must be 16-byte aligned at a `call`. And there's a **red zone** — 128 bytes
below RSP that a function may use as scratch without adjusting RSP.
The name is historical: it descends from AT&T's *System V* Unix, whose ABI
documents this lineage comes from. The modern spec is the "System V Application
Binary Interface, AMD64 Architecture Processor Supplement."
## Why danos pins it: RDI vs RCX
The reason this is called out explicitly is a clash with the *other* common x86-64
convention, **Microsoft x64** — used by Windows **and UEFI** — where the first
argument arrives in **RCX**, not RDI.
danos's two binaries default to different conventions:
- `boot/efi.zig` is built for the UEFI target, so its default C convention is
Microsoft x64 (first argument → RCX).
- The kernel is freestanding, so its convention is SysV (first argument → RDI).
When the loader jumps to the kernel passing the `BootInformation` pointer, both sides have
to agree *which register that pointer lands in*. Left to their defaults, the loader
would place it in RCX while the kernel looked in RDI — and the kernel would read
garbage. So both sides reference the same `boot_handoff.kernel_abi` (SysV): the loader's
function-pointer type and the kernel's `_start` both carry
`callconv(boot_handoff.kernel_abi)`, and the pointer reliably arrives in RDI. That is the
whole reason `kernel_abi` lives in the shared contract — see [efi.md](efi.md) for
the handoff it governs.
## The process-entry stack (argc/argv)
The SysV ABI also fixes what a *fresh process* finds on its stack — and danos
follows it, so its own runtime and any future C libc read arguments the same way.
At the first user instruction, `rsp` is 16-byte aligned and points at (addresses
growing upward):
```
rsp → argc u64
argv[0] … argv[argc-1] pointers into the strings area below
NULL argv terminator
NULL envp terminator (no environment yet)
{AT_PAGESZ, page size} auxiliary vector
{AT_NULL, 0} auxiliary-vector terminator
argv string bytes NUL-terminated
───────────────────────── stack top (stack_top_virtual)
```
The kernel builds this block at the top of the process's stack — 8 pages (32 KiB,
`parameters.user_stack_pages`) mapped RW+NX below a fixed top, with the page below
them left unmapped as a **guard**, so a stack overflow faults (killing only that
process) instead of silently corrupting the image
(`buildEntryStack` in `system/kernel/process.zig`); `argv[0]` is always the path
or initial-ramdisk name the process was spawned as, and `system_spawn`'s optional
argument blob becomes `argv[1..]`. The runtime's `_start`
(`library/kernel/start.zig`) hands the block to `rt_start`, which builds a
`process.Init` from it and passes that to the program's `main`
(`pub fn main(init: process.Init)`; a parameterless `main()` is also
accepted). A C runtime's `crt0` would walk
the identical layout unmodified — that's the compatibility being bought. The
`args` test proves the round trip.
## Where else it surfaces
- **The red zone → `red_zone = false`.** `build.zig` disables the red zone for the
kernel. Interrupts push their frame onto the current stack; if the interrupted
code was using its 128-byte red zone, that push would stomp it. Turning the red
zone off is the standard fix for kernel code — a direct consequence of SysV
*having* a red zone. (See [interrupts.md](interrupts.md).)
- **`callconv(.c)` == SysV here.** The exception/interrupt dispatcher
(`interruptDispatch`) is declared `callconv(.c)`, which resolves to SysV on this
target. That's why the assembly stub in `isr.s` moves the `CpuState` pointer into
**RDI** before `call`ing it — the same first-argument rule.
So "the kernel is SysV" means: its functions pass arguments in RDI/RSI/RDX/…,
return in RAX, preserve the SysV callee-saved registers, and assume a red zone —
and every boundary that calls into the kernel (the loader, the interrupt stubs)
has to speak that same convention at the point of the call.
## A note on other architectures
This is x86-64-specific. An AArch64 port ([architecture.md](architecture.md)) has its own calling
convention (arguments in X0–X7, and so on) — a different ABI entirely. `kernel_abi`
would be set per-architecture, but the *principle* is the same: the loader/entry
boundary and the kernel must agree on how arguments are passed.
+471
View File
@@ -0,0 +1,471 @@
# Threading — build plan (`runtime.Thread` over a private thread ABI)
The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
[display-v2-plan.md](../device-driver-development-guide/display-v2-plan.md). Read threading.md first for the *why*.
## Locked decisions (do not relitigate)
- **`runtime.Thread` mirrors `std.Thread`'s API; the implementation is danos-native.**
Not literal `std.Thread` — that would break the [private ABI](syscall.md).
- **Threads are a narrow, per-binary opt-in.** Default concurrency stays process + IPC
([resilience.md](resilience.md)); only a service that asks is built
`single_threaded = false`.
- **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an
idle core still halts ([halting.md](halting.md)).
- **New syscalls are private**: extend [abi.zig](../../system/abi.zig) `SystemCall` after
`shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`,
`futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number.
- **Restart granularity stays the process** — a faulting thread kills its process; the
supervisor restarts the process, which respawns its threads.
## Conventions
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations,
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
`addUserBinary` (with the new `threaded` flag where a binary spawns threads) and get
packed into the initial-ramdisk; new syscalls extend [abi.zig](../../system/abi.zig)
`SystemCall` + a `library/runtime` wrapper; test services live beside the code they
exercise and register a `ServiceId` if they must be looked up.
## How to verify along the way
**Every gate is serial-checkable — no screenshots** (this plan runs unattended). A
thread proves it ran by writing to **shared memory** the parent reads back, and proves
parallelism by stamping the **core index** it ran on (like the `smp`/`affinity` cases).
- `zig build test` — host unit tests (closure packing, mutex state machine, futex
wrapper encodings).
- `python3 test/qemu_test.py <case>` — boots the kernel in QEMU; asserts on serial
markers. Thread cases set `smp: true` (real parallelism) and bump `mem` (they boot
the process/scheduler stack); each milestone **adds its case to `CASES`** so its gate
is runnable.
- **Guardrail every milestone:** the concurrency-sensitive existing cases stay green —
`smoke`, `sched`, `priority`, `smp`, `affinity`, `process`, `process-kill`,
`supervision`, `fault-recovery`, `vfs-client-death`, `ipc`/`ipc-cap`,
`display-service`. A threading change that regresses those is rejected.
## Unattended execution (the loop contract)
This plan runs to completion **without human input**. Every design choice is already
fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration must:
1. **Resume** at the first milestone that still has an unchecked `- [ ]`. (All earlier
milestones are done — do not revisit them.)
2. **Work on a branch.** On the first iteration, branch off the current `main` into a new
branch (e.g. `threading-phase2` — Phase 1's `threading` is already merged); never
commit to `main` directly. Push that **branch** to `origin` after each milestone (step
5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into
`main` stays a human step.
3. **Implement** every unchecked item in that milestone, including adding its
`-Dtest-case` to `CASES` in [test/qemu_test.py](../../test/qemu_test.py) (with
`smp: true` / a `mem` bump where noted) so the gate is runnable.
4. **Run the gate**: `python3 test/qemu_test.py <case>`, then the full **guardrail
set**, then `zig build` (clean) and `zig build test` (green).
5. **Decide, do not ask:**
- **Green** = the milestone's case prints its stated marker(s) and reports `PASS`,
the whole guardrail set passes, `zig build` is clean, and host tests are green.
→ tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit`
(`threads(M<n>): <summary>`, no `Co-Authored-By` trailer per
[coding-standards.md](../coding-standards.md)), then **`git push` the working branch to
`origin`** (use `-u` on the first push to set upstream). Continue to the next
milestone in the same iteration if budget remains; otherwise let the loop re-fire.
- **Red** = anything above fails. Diagnose from the captured serial log
(`zig-out/qemu-test/<case>-failed-serial.log`) and fix in place, then re-run — up to
**3 fix attempts** for that gate. A concurrency case that fails then passes on a
bare re-run is **flaky, not green**: re-run it **twice more** and treat green only
if it passes all; otherwise fix the race (a real threading bug), don't paper over
it.
6. **A genuinely ambiguous fork is not a stop.** Pick the option most consistent with
[threading.md](threading.md)'s *Locked decisions*, note the choice in the commit
message, and continue. Do not pause for confirmation on in-scope, reversible work —
this plan is that authorization.
**The only stop conditions:**
- **Done** — every milestone box **in this plan** is checked (M1 through M11), `zig build`
clean, the whole `thread-*` suite + guardrail green. Phase 1 (M1–M6) is *already*
checked, so do **not** read that as Done: the loop's real work is the first plan section
that still has unchecked boxes — Phase 2 (M7–M11). Only stop when M7–M11 are all checked
too. Update threading.md's status line, push the final branch state to `origin`, and
stop. The branch is on `origin` for review; **merging Phase 2 into `main` is the user's
step**, not the loop's.
- **Blocked** — a gate is still red after 3 fix attempts, or a step needs something
outside the repo (a toolchain change, new hardware, a decision no locked decision
covers). Append `> **BLOCKED (M<n>):** <what failed, what was tried, the serial
marker missing>` under that milestone, commit **and push** the WIP on the branch, and
stop. Do not thrash further and do not silently skip the milestone.
Nothing else warrants stopping — not "should I proceed?", not "is this right?". The
checkboxes + git history are the resumable record; the next iteration picks up from the
first unchecked box.
---
## M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅
The one invariant change threads require, landed and proven **before** anything shares
an address space. Today address space is 1:1 with a task and teardown destroys it on any user
task's exit; make destruction happen on the **last** exit.
- [x] A refcount keyed by the address-space root, held in `scheduler.zig`
(`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the
success path, after the slot + stack are secured), all under the big kernel lock.
- [x] Both task-teardown paths ([scheduler.zig](../../system/kernel/scheduler.zig):
`exitUserLocked` and `destroyTaskLocked`) call `releaseAddressSpace`, which decrements
and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test
spaces) is destroyed directly, preserving prior behaviour.
- [x] `-Dtest-case=address-space-refcount`: spawn and reap several ring-3 processes in sequence
and assert (via test-observable `liveAddressSpaceCount`/`addressSpaceDestroyCount`) that the
live-space count returns to **baseline** and destructions advance by exactly that
many — each space destroyed exactly once, no leak, no double-free. (Refcount
observables, not raw frame counts, since kernel stacks are still leaked on exit.)
**Gate (met):** `python3 test/qemu_test.py address-space-refcount` passes
(`address-space-refcount: spaces released to baseline ok` → `DANOS-TEST-RESULT: PASS`), and the
full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `smp`,
`affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`,
`vfs-client-death`, `ipc`, `ipc-cap`, `display-service`); default `zig build` clean,
`zig build test` green. The reframing is invisible until an address space is actually shared.
## M2 — `thread_spawn` + `thread_exit`: a thread runs in the shared address space ✅
Spawn only — no join yet. Prove a second task executes in the **caller's** address
space and exits cleanly.
- [x] [abi.zig](../../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in
process.zig; `thread_spawn` calls `scheduler.spawnThread` (today, after M3, the
handler goes `spawnThreadSupervised` → `scheduler.spawnUserLocked`; shares the caller's
address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)`
(`terminateCurrent` → `releaseAddressSpace`). The closure pointer is delivered in the new
thread's **rdi** via a new `jump_to_user_arg` asm path (`t.user_arg`, 0 for a
process) — no naked runtime asm.
- [x] `library/runtime/thread.zig` (barrel-exported as `runtime.Thread`): `spawn` maps a
stack (`mmap`), heap-allocates the `{args}` closure, and calls
`thread_spawn(&Closure.entry, stack_top, closure)`; `Closure.entry` (a plain C-ABI
Zig fn, closure in rdi) runs the function and calls `thread_exit`. Stack top is
16-aligned-minus-8 for the C entry.
- [x] A `threaded` flag on the user-binary recipe (`addThreadedUserBinary` →
`single_threaded = false`); `thread-test` is the first opt-in binary.
- [x] `-Dtest-case=thread-spawn`: `thread-test` spawns a worker that writes a sentinel to
a **shared** global and release-stores `done`; the main thread acquire-polls `done`
and asserts the shared global holds the sentinel — proof the worker ran in the same
address space.
**Gate (met):** `python3 test/qemu_test.py thread-spawn` passes
(`thread-test: child ran in shared address space ok` → `DANOS-TEST-RESULT: PASS`); guardrail set
16/16 green (incl. `args`/`init`/`process`, which exercise the new `jump_to_user_arg`
process path with arg 0) plus `address-space-refcount`; `zig build` clean, `zig build test`
green.
> **Note (deferred to M3+):** the mmap arena is per-*task* (`heap_next`), so two threads
> in one address space that both `mmap` would collide. Fine for M2 (only the parent maps, for the
> child's stack); make the arena per-address-space and the runtime heap thread-safe alongside the
> `Mutex` work (M5).
## M3 — `join` + `detach` + real parallelism ✅
- [x] `join` over the existing exit-notification path
([process-lifecycle.md](process-lifecycle.md)): `thread_spawn` gained a 4th arg, an
`exit_endpoint` handle (resolved + refcounted like `spawnProcessSupervised`, via
`spawnThreadSupervised`); `join` blocks in `ipc_reply_wait` on that endpoint until
the child-exit notice for its `tid`, then `munmap`s the stack. `detach` relinquishes
the join right (its stack is reclaimed at process exit — kernel-reaper reclaim for
detached threads is deferred; see note).
- [x] `runtime.Thread.join` / `detach`, plus `Thread.currentCore()` (a new `current_core`
= 39 syscall) for the parallelism proof. `getCurrentId` deferred to M6 (TLS), where
a lighter self-id fits. The closure now rides the **thread's own stack** (not the
heap) — private per thread, so spawn/join touch no shared heap.
- [x] `-Dtest-case=thread-join` (`smp: 4`): `thread-test` join mode spawns N=4 workers
that each do K=100k `@atomicRmw`-increments on a shared counter and stamp the core
they ran on; the main thread joins all N and asserts `counter == N*K` **and**
`@popCount(cores_seen) > 1` (genuine cross-core parallelism), then a detached worker
proves `detach` runs without a join.
**Gate (met):** `python3 test/qemu_test.py thread-join` passes (`thread-test: join ok` →
`DANOS-TEST-RESULT: PASS`), robust across 4 runs; guardrail 17/17 green (incl. `smp`,
`affinity`, `process-kill`, and `args`/`init`/`process` on the exit-endpoint spawn path)
plus `address-space-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green.
> **Note (deferred):** a detached thread's stack is freed only at process exit (not by the
> reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And
> the runtime heap is still not thread-safe: threads that both allocate concurrently would
> race (the thread *machinery* avoids the heap, but worker code sharing an allocator does
> not). Both fold into the M5 `Mutex`/allocator work.
## M4 — Futex: the one blocking primitive ✅
- [x] [abi.zig](../../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a
`.blocked` task tagged with `Task.futex_addr` (no queue linkage);
`futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock,
parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr,
count)` scans the task table and readies up to `count` matching waiters (same
address space). No spinning — a parked waiter leaves its core free to `hlt`. A
timed wait also sets `wake_at`, so the timer's `wakeExpired` wakes it; `futex_addr`
staying non-zero (only `futex_wake` clears it) is how the waiter tells timeout from
a real wake.
- [x] `runtime.Thread.Futex` (`wait` / `timedWait` / `wake`) over the syscall wrappers.
- [x] `-Dtest-case=thread-futex` (`smp: 4`): a waiter thread prints `waiting` and
`futex_wait`s on a word; the main thread publishes it, prints `waking`, and
`futex_wake`s; the waiter prints `woke`. Then a `timedWait` on an unwoken word
reports `error.Timeout`.
**Gate (met):** `python3 test/qemu_test.py thread-futex` passes, robust across 3 runs —
the case's **ordered** regex asserts `waiting → waking → woke → PASS` on the serial
stream (the handoff proof), and `thread-futex: timeout ok` confirms the timeout.
Guardrail 18/18 green (incl. `sleep`/`event`/`ipc` blocking paths) + `address-space-refcount`,
`thread-spawn`, `thread-join`; `zig build` clean, `zig build test` green.
> **Note:** the kernel test checks only the freshest verdict marker via `bufferHas` (the
> in-memory log ring buffer evicts older lines); ordering is asserted against the full
> serial stream by the qemu regex instead.
## M5 — `Mutex` + `Condition` + `Semaphore` ✅
- [x] `runtime.Thread.Mutex` (three-state futex mutex: CAS fast path, `futex_wait`/`wake`
slow path), `Condition` (`wait`/`timedWait`/`signal`/`broadcast`, a futex sequence
counter), `Semaphore` (permits over `Mutex`+`Condition`) — the same state machines
`std.Thread` uses, ported onto our `Futex`.
- [x] `-Dtest-case=thread-mutex` (`smp: 4`): a bounded producer/consumer — 2 producers +
2 consumers over one `Mutex` and two `Condition`s move N=2000 unique items through
an 8-slot ring; the consumed checksum and tally match exactly (no lost/duplicated
item, no overrun) under real cross-core contention. The small ring forces producers
to block on full and consumers on empty, exercising `Condition.wait`.
**Gate (met):** `python3 test/qemu_test.py thread-mutex` passes (`thread-mutex: ok` →
`DANOS-TEST-RESULT: PASS`), robust across 3 runs; guardrail 17/17 green (incl.
`sleep`/`event`/`ipc`) + all M1–M4 thread cases; `zig build` clean, `zig build test`
green.
> **Deferred (with rationale):**
> - **`join` → futex completion word** — the exit-endpoint join (M3) is correct and
> tested. A futex-completion join needs the *kernel* to clear+wake a word after the
> thread is fully off its stack (a CLONE_CHILD_CLEARTID-style mechanism); doing it in
> the thread's own trampoline would let `join` `munmap` the stack while the thread still
> runs on it (use-after-free). Left on the exit-endpoint path; the kernel clear-on-exit
> is a later, separate refinement.
> - **Host unit tests for the state machines** — `Mutex`/`Condition` bottom out in the
> `futex_*` syscalls, unavailable on the host without a mockable `Futex` seam. The QEMU
> `thread-mutex` gate exercises them under real concurrency instead; a host-side mock is
> future work.
## M6 — `getCurrentId`, docs, and CI wiring ✅
- [x] `getCurrentId` via a small `thread_self = 42` syscall (`runtime.Thread.getCurrentId`
returns the kernel task id). **Per-thread `threadlocal` TLS is deferred** — no
consumer needs it, and it would require context-switching the thread pointer per task
(real kernel + per-switch cost) for an unused feature; threaded binaries have run fine
without it through M2–M5. threading.md's TLS reasoning already scoped it as
deferred-unless-needed. When a consumer appears, the shape is: `thread_spawn`
allocates a per-thread TLS block, sets the thread pointer, and the context switch saves/
restores it.
- [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same
`Futex`/`Mutex`/`Condition` when wanted.
- [x] All `thread-*` cases wired into [test/qemu_test.py](../../test/qemu_test.py)
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md
status updated to **built**; the worked example is threading.md's win-condition.
- [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main
thread confirms all three ids are non-zero and distinct — each thread has its own
kernel identity. (Renamed from `thread-tls`, which implied `threadlocal`.)
**Gate (met):** `python3 test/qemu_test.py thread-id` passes; the whole `thread-*` suite
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`) plus the full guardrail set pass; default
`zig build` clean, `zig build test` green.
---
## Status
**Phase 1 (M1–M6): built.** danos has `runtime.Thread` — `spawn`/`join`/`detach`,
cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private
thread ABI behind the runtime.
**Phase 2 (M7–M11): built.** Thread-safe allocation (M7), a task reaper that reclaims dead
tasks' kernel stacks (M8), endpoint-free `thread_join` (M9), the per-thread thread pointer (M10),
and `RwLock`/`WaitGroup` + host-testable sync (M11). Two things stay deferred by design
(no consumer): the Zig `threadlocal` *compiler* layer (M10) and detached-thread user-stack
reclaim (M9) — both noted in place.
---
## Phase 2 — hardening (M7–M11)
The organising principle, so Phase 2 reinforces danos's goals rather than eroding them:
- **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks,
and futex words live in the process's **address space**, and the kernel's per-process
state is keyed by the address-space root — so the M1 refcount + `destroyAddressSpace` already
free all of it when the last thread exits. A crashed or killed threaded process leaves
**nothing** behind. Phase 2 closes the one thing that is *not* address-space-owned — the
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
[resilience](resilience.md) restart guarantee, extended to threads.
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
restores the thread pointer, and reaps dead tasks; the runtime decides allocation, TLS layout,
and lock algorithms. Every new kernel entry stays a private syscall behind the runtime
([syscall.md](syscall.md)) — the ABI stays renumberable.
- **The process is still the isolation and restart boundary.** Threads share fate within
one process; Phase 2 never adds a way for one process to reach into another (the
cross-process futex stays explicitly out of scope, below).
### M7 — Thread-safe allocation (the correctness gap) ✅
Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two
threads in one process that both allocate corrupt each other. The thread *machinery*
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
multi-threaded code would hit it. Closed it:
- [x] **Kernel — per-address-space mmap arena.** Grew M1's `address_space_refs` entry into the
per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`);
`scheduler.addressSpaceMmapNextPtr`/`addressSpaceDeviceMapNextPtr` expose them. `systemMmap`
reserves a disjoint range under a *brief* lock, then maps **per page** under a
short-held lock — not the whole grant — because the big lock is held with interrupts
disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed
the `affinity` scenario out mid-bring-up). Freed at refcount zero, so the cursors
vanish with the process.
- [x] **Runtime — thread-safe heap.** The allocator's two free-list mutators
(`rawAlloc`/`rawFree`) take a `Thread.Mutex`, gated on
`!@import("builtin").single_threaded` so single-threaded binaries compile it out and
pay nothing. Uncontended acquisition is a single CAS (no syscall).
- [x] `-Dtest-case=thread-alloc` (`smp: 4`): 4 threads each do 500 `alloc`/fill/verify/
`free` cycles of varied sizes; each block is filled with a per-thread pattern and
verified before free, so any overlap between concurrent allocations is caught.
**Gate (met):** `thread-alloc` passes (3× non-flaky); full guardrail 23/23 green,
`zig build`/`zig build test` clean.
> **Also fixed here:** the `affinity` guardrail's fixed-count busy-loop (`while (spins <
> 3e9)`) had codegen-dependent wall-time — adding a function to `tests.zig` flipped how
> the optimiser compiled it, swinging affinity from ~4 s to ~63 s and timing it out.
> Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is
> independent of unrelated code changes.
### M8 — The task reaper (cleanup + resilience) ✅
A dead task's **kernel** stack was leaked ("no reaper yet") — every process *and* thread
death lost one, so a crash loop bled kernel memory. The reaper fixes it and serves the
[resilience](resilience.md) restart goal directly:
- [x] A dying task cannot free the kernel stack it runs on, so `exit()`/`exitUserLocked`
record it in a **per-core `reap_after_switch` slot** and switch away; the task that
resumes on that core frees the stack in `switchTo`'s tail (it's on its own stack, the
big lock is still held so the slot can't have been reused). A **tick-time drain**
(`reapKillPendingLocked`) is the safety net for the case where the next task is
*fresh* (enters via the trampoline, bypassing `switchTo`'s tail). A task killed while
*not* running is freed immediately in `destroyTaskLocked`. A `live_stack_bytes`
counter is the observable. *(Detached-thread user-stack reclaim moves to M9, which
adds the joinable/detached flag.)*
- [x] `-Dtest-case=task-reap` (`smp: 4`): spawn and kill 12 processes; poll the
test-observable `scheduler.liveStackBytes()` until it returns to **baseline** (a
correct reaper gets there in a few ms; a genuine leak times out) — every kernel
stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered.
**Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`,
`supervision`, `process-kill`, `address-space-refcount`, `smp`, `affinity` all still green (24/24
full guardrail); `zig build`/`zig build test` clean.
> **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap
> first read the `pc` **parameter**, but a task that migrated cores carries a *stale* `pc`
> in its saved `switchTo` frame — so it read the wrong core's slot and freed a live stack
> (a #GP under SMP). Fixed to re-fetch `thisCpu()` after the switch (the switch only swaps
> stacks on the current core).
### M9 — Futex-completion join (retire the per-thread endpoint)
With the reaper (M8) able to act *after* a thread is fully off its stack, migrate `join`
to the std shape and drop M3's per-thread exit endpoint:
- [x] A **`thread_join(tid)` syscall** (not a user futex word): it blocks the caller until
the task with id `tid` exits, and the exit paths call `wakeJoinersLocked`. `join`
only reclaims the joined thread's **user** stack, which the thread vacates the moment
it enters the kernel to exit — so waking at *exit* time (not reap time) is safe, and
no reaper/address-space juggling or user-memory write is needed. This is equally
std-shaped (like `pthread_join`) and much simpler/safer than the planned
reaper-written completion word. The runtime no longer passes `thread_spawn` an exit
endpoint (it passes `no_cap`; the kernel's 4th `exit_endpoint` arg remains and is
still honored); the runtime's per-thread IPC endpoint is gone.
- [x] `thread-join` passes on the new path, and its join mode now runs **40 spawn+join
cycles** — under the old per-thread-endpoint scheme those leaked handles would
exhaust the 16-slot handle table; here they all succeed, proving join is endpoint-free.
**Gate (met):** `thread-join` passes (3× isolated) on the `thread_join` path; full
guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-reap`);
`zig build`/`zig build test` clean.
> **Reaper hardened here (fixes an M8 flake).** M8's single per-core reap slot could be
> *overwritten* by a second death on that core before the first drained (a fresh-task/SMP
> timing window) — an intermittent one-stack leak (`task-reap` flaked ~20%). Replaced it
> with a per-core reap **list** plus a `.reaping` task state so a pending slot can't be
> reused before its stack is freed. `task-reap` now 11/11 isolated + 2× in the batch.
> **Deferred:** detached-thread **user-stack** reclaim (still freed at process exit, as in
> M3). Doing it in the reaper needs the saved address space + stack range and a
> translate/unmap in a not-currently-loaded address space — real complexity for a bounded leak.
> A follow-up when a consumer needs it.
### M10 — Per-thread TLS: the thread-pointer mechanism ✅
Give each thread its own thread pointer and private TLS storage — the foundation
self-hosting Zig ([zig-self-hosting.md](../zig-self-hosting.md)) will build `threadlocal` on.
- [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch
**only when it changes** (the same conditional-load discipline as CR3;
`architecture.setThreadPointer` → `wrmsr IA32_FS_BASE` on x86_64). A
`set_thread_pointer(addr)` = 44 syscall sets the caller's `thread_pointer` and loads it
now. The kernel never touches FS, so there is no swapgs complication.
- [x] **Runtime** lays a small per-thread TLS block at the top of each thread's stack
(self-pointer at `%fs:0` + scratch slots) and the thread trampoline calls
`set_thread_pointer` before any user code — so every spawned thread has a private,
switch-stable thread pointer. Reclaimed with the stack.
- [x] `-Dtest-case=thread-tls` (`smp: 4`): two threads each write a unique marker to their
own `%fs:8` slot and — after both have written — read it back; a shared (non-per-thread)
FS base would clobber one and cause cross-talk. Both read their own marker → pass.
**Gate (met):** `thread-tls` passes (3×); full guardrail 25/25 (the switch-time thread-pointer
restore touches every context switch); `zig build`/`zig build test` clean.
> **Deferred: the Zig `threadlocal` *compiler* layer.** Real `threadlocal` variables need
> the ELF **variant-II TLS** surface — `.tdata`/`.tbss` sections + a `PT_TLS` program header
> in `user.ld`, a runtime that copies the template with exact negative-offset layout, and
> the `.large`-code-model TLS section names — a high-uncertainty lift for a feature with
> **no consumer today** (threading.md scopes it "only if a consumer needs it"). What lands
> here is the load-bearing piece — the per-thread thread pointer, context-switched — so adding the
> compiler layer later is purely runtime+linker work on top, no kernel change. `getCurrentId`
> stays the `thread_self` syscall (M6) rather than an fs self-slot (which would need the
> main thread's TLS set up in `_start` too).
**Gate:** `thread-tls` passes; full `thread-*` suite + guardrail green.
### M11 — `RwLock`, `WaitGroup`, and host-testable sync ✅
- [x] `runtime.Thread.RwLock` (reader-preferring: `>0` readers / `-1` writer / `0` free,
with `lock`/`tryLock`/`unlock` + `lockShared`/`tryLockShared`/`unlockShared`) and
`WaitGroup` (`start`/`finish`/`wait`), both on the existing `Mutex`/`Condition`.
- [x] A compile-time `Futex` seam gated on `builtin.os.tag == .freestanding`: the futex
syscalls on danos, a spin+yield mock off-target (Zig 0.16 has no `std.Thread.Futex`;
`wake` is a no-op since the state machines re-check). `thread.zig` is wired into
`zig build test`, so `Mutex`/`RwLock`/`WaitGroup` run as **host unit tests** with real
`std.Thread` threads (`test` blocks only compile under test).
- [x] `-Dtest-case=thread-rwlock` (`smp: 4`): 2 writers set both halves of a value under
the exclusive lock while 3 readers check the halves match under the shared lock —
zero half-write observations across ~150k reads. Host tests cover the Mutex,
RwLock, and WaitGroup state machines.
**Gate (met):** `zig build test` covers the sync primitives (host threads); `thread-rwlock`
passes (3×); full Done gate **26/26** (whole `thread-*` suite + guardrail); `zig build`
clean.
---
## Deferred (explicitly not in this plan)
- **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a
physical-address key so two processes share a futex through a [shared-memory](../device-driver-development-guide/display-v2.md)
region. Not needed for intra-process threads.
- **Per-thread priorities / affinity distinct from the process** — threads inherit the
process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep.
- **Per-thread signal delivery** — signals stay process-scoped
([process-lifecycle.md](process-lifecycle.md)).
- **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more.
- **A real `std.Thread` backend** — arrives with self-hosting
([zig-self-hosting.md](../zig-self-hosting.md)); it sits on these same primitives, so it
swaps the impl under `runtime.Thread`, not the call sites.
+349
View File
@@ -0,0 +1,349 @@
# Threading: `Thread`, a std-shaped API over a private thread ABI
A note on danos **threads** — several tasks sharing one address space — provided by a
`Thread` type that mirrors the shape of Zig's `std.Thread` while keeping every
kernel entry behind the [runtime](../../library/kernel). **Built** (M1–M11, see
[threading-plan.md](threading-plan.md)): `spawn`/`join`/`detach`, cross-core parallelism,
a futex, `Mutex`/`Condition`/`Semaphore`/`RwLock`/`WaitGroup`, `getCurrentId`/`currentCore`,
per-thread thread-pointer TLS, thread-safe allocation, and a task reaper that reclaims dead
tasks' kernel stacks. Deferred by design (no consumer yet): the Zig `threadlocal`
*compiler* layer (the per-thread thread pointer is in place, so it's runtime+linker work on top) and
detached-thread user-stack reclaim — see the plan's M9/M10 notes. The analysis is against
**Zig 0.16** (the pinned toolchain); `std.Thread`'s internals move between releases, so
treat upstream shapes as "0.16.x."
## The win condition
A danos service can write
```zig
const t = try Thread.spawn(.{}, worker, .{ctx});
// ... do other work concurrently ...
t.join();
```
and get real parallelism across cores — with `Thread.Mutex`,
`Thread.Condition`, and `Thread.Semaphore` available for
coordination — **without any code path reaching the kernel except through the
runtime**. The call sites read exactly like `std.Thread`, so the day danos becomes a
real Zig target (see [self-hosting](#the-self-hosting-endgame)) we swap the
implementation underneath, not the API above.
## Locked decisions (do not relitigate)
- **We build `Thread`, not literal `std.Thread`.** It mirrors std's *API and
features*; the implementation underneath is danos-native. See
[Why not literal std.Thread](#why-not-literal-stdthread).
- **Threads are a narrow, opt-in capability — not the default concurrency tool.** The
default for resilience stays **process + IPC** ([resilience.md](resilience.md),
[ipc.md](../device-driver-development-guide/ipc.md)). See [Where threads fit](#where-threads-fit-the-resilience-tension).
- **Blocking synchronization is futex-backed, never spin-backed.** Waiters sleep in
the kernel so an idle core still halts ([halting.md](halting.md)).
- **Per-binary opt-in to multi-threaded codegen.** Only a service that asks for
threads is built `single_threaded = false`; the rest stay lean and single-threaded.
- **The thread ABI is private.** New syscalls extend [abi.zig](../../system/abi.zig)
`SystemCall` and are reached only through `library/kernel` wrappers, exactly like
every other danos syscall ([syscall.md](syscall.md)) — numbers stay renumberable.
## Why not literal `std.Thread`
danos's ABI invariant is that the **runtime is the sole holder of the syscall ABI**,
and that ABI is private and renumberable ([syscall.md](syscall.md) — "unstable
private ABI"). That is a security and evolvability asset: no compiled binary can
hardcode a syscall number, and the kernel can renumber freely because only the
runtime — rebuilt in lockstep — knows the mapping.
`std.Thread` is incompatible with that invariant on two counts:
1. **It selects its backend from `builtin.os.tag`, and issues syscalls directly.**
danos targets `.os_tag = .freestanding` ([build.zig](../../build.zig)), for which
`std.Thread` resolves to an unsupported stub that `@compileError`s. Adding a real
backend would either bake danos syscall numbers into std (breaking ABI privacy and
renumbering) or fork std to route back through the runtime — a permanent rebase
cost that buys nothing the native type doesn't.
2. **Our user binaries are built `single_threaded = true`** ([build.zig](../../build.zig)
`addUserBinary`), which compiles threading out entirely and makes atomics and TLS
single-threaded. Threads need this flipped per binary regardless.
So we take the *shape* of `std.Thread`, not the *type*. The cost of replicating the
surface (spawn/join/Mutex/Condition) is small; the cost of the std type is the ABI
invariant.
## Where threads fit: the resilience tension
Threads are in genuine tension with a resilience-first microkernel, and it is worth
being explicit so we do not reach for them by reflex.
The reason danos pays for a microkernel is **fault isolation**
([resilience.md](resilience.md)): a component corrupts its own address space, faults,
and is **restarted** without touching anyone else — because the boundary *is* the
address space. Threads deliberately remove that boundary *within* a process:
- Threads share one address space, so one thread's stray write corrupts them all —
there is no isolation **between** threads.
- Threads share fate — by contract: a fault in any thread, or a "kill the process"
decision, takes down **all** of them, so restartability lives at the process level,
not the thread level. (The kernel does not yet enforce this fan-out — see the
Lifecycle note under
[Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).)
- Shared mutable state reintroduces data races — the failure class the
isolate-and-message model was chosen to avoid.
**Therefore:** the default answer to "make X concurrent" stays *another process over
IPC* (isolated, independently restartable) or a single event loop with several
message sources. Reach for a thread only inside **one** service that needs genuine
**shared-memory, low-latency parallelism** and can accept intra-service fate-sharing —
e.g. a compositor splitting tile compositing across cores, where per-tile IPC would be
too chatty. "Input on one thread, display on another" is *not* that case; it wants two
processes. The isolation boundary stays at process granularity.
## The API surface (mirrors `std.Thread`)
Lives in `library/kernel/thread.zig`, re-exported as `Thread`.
```zig
pub const Thread = struct {
pub const Id = u32; // the kernel task id
pub const SpawnConfig = struct {
stack_size: usize = default_stack_size, // no allocator: the closure lives at the top of the thread's own stack
};
pub const SpawnError = error{SystemResources};
pub fn spawn(config: SpawnConfig, comptime function: anytype, args: anytype) SpawnError!Thread;
pub fn join(self: Thread) void; // block until the thread ends, reclaim its stack
pub fn detach(self: Thread) void; // give up the right to join; stack reclaimed at process exit
pub fn getCurrentId() Id;
pub fn currentCore() Id; // danos extension: the calling core's dense index
pub const Mutex = struct { pub fn lock(*Mutex) void; pub fn tryLock(*Mutex) bool; pub fn unlock(*Mutex) void; };
pub const Condition = struct { pub fn wait(*Condition, *Mutex) void; pub fn timedWait(*Condition, *Mutex, u64) error{Timeout}!void; pub fn signal(*Condition) void; pub fn broadcast(*Condition) void; };
pub const Semaphore = struct { pub fn wait(*Semaphore) void; pub fn post(*Semaphore) void; };
pub const Futex = struct { pub fn wait(*const atomic.Value(u32), u32) void; pub fn timedWait(...) error{Timeout}!void; pub fn wake(*const atomic.Value(u32), u32) void; };
// RwLock / ResetEvent / WaitGroup follow the same pattern, added as needed.
};
```
Deviations from `std.Thread`, called out honestly:
- **The thread function's return value is discarded** (as `std.Thread.join` returns
`void`). Return data through shared state or a `Semaphore`/`Condition`, not the
return.
- No `getCpuCount()` (a service rarely needs it) and no `Thread.yield()` — `yield`
lives in the `process` module. Instead `currentCore()` exposes the calling core's dense
index ([smp.md](smp.md)), used to observe genuine cross-core parallelism.
## Kernel primitives (new private syscalls)
Five core entries extend [abi.zig](../../system/abi.zig) `SystemCall` after
`shared_memory_physical = 36` (plus small helpers `current_core`, `thread_self`, and
`set_thread_pointer`), each with a `library/kernel` wrapper:
| Syscall | Signature | Purpose |
|---|---|---|
| `thread_spawn` | `(entry, stack_top, arg, exit_endpoint) -> tid` | create a task sharing the **caller's** address space; the runtime passes `no_cap` for `exit_endpoint` (join is a syscall, not an endpoint) |
| `thread_exit` | `()` | end the calling thread; its stack is reclaimed by the joiner's `munmap`, not the kernel |
| `thread_join` | `(tid) -> 0` | block until the thread with id `tid` has exited |
| `futex_wait` | `(addr, expected, timeout_ns) -> status` | block if `*addr == expected`, until woken or timeout |
| `futex_wake` | `(addr, count) -> woken` | wake up to `count` waiters on `addr` |
Plus one invariant change with no new syscall: **address-space reference counting**.
## Mechanics
### Address-space reference counting
Before this work an address space was 1:1 with a task: `spawnUserLocked` records
`address_space` on the Task (as it still does), and teardown did
`destroyAddressSpace(t.address_space)` when **any** user task exited
([scheduler.zig](../../system/kernel/scheduler.zig)). With threads, several tasks share
one `address_space`, so the first to exit would rip the address space out from under its
siblings.
Fix: a small refcount keyed by the address-space root, kept in
[scheduler.zig](../../system/kernel/scheduler.zig): `retainAddressSpace` takes a
reference for every user task `spawnUserLocked` starts (count 1 on the first take, so
a thread sharing the caller's space increments it); task teardown calls
`releaseAddressSpace`, which only calls `destroyAddressSpace` at **zero**. All of
this is already under the big kernel lock, so no new locking. This is the one piece
that must land and be proven before anything shares an address space.
### `thread_spawn` and the trampoline
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle
values through scratch registers — `startUserTask` reads the entry/stack (and the
thread's closure arg, delivered in `rdi` via `jumpToUserArg`) from the Task
([scheduler.zig](../../system/kernel/scheduler.zig)). That makes the thread path clean:
1. The runtime's `spawn` `mmap`s a stack (syscall `4`) and writes the closure —
`{ tls_base, args }`, the std "Instance" pattern — at the **top of the new stack
itself** (no heap allocation), with a small per-thread TLS block just below it.
2. It calls `thread_spawn(entry = &Closure.entry, stack_top, arg = closure_ptr,
exit_endpoint = no_cap)`. The kernel calls the same `spawnUserLocked` path with the
**caller's address space** (refcount++), `entry`, and `user_sp = stack_top`.
3. `Closure.entry` (a small runtime shim) receives the closure pointer in `rdi` — the
kernel delivers `arg` as the entry's first C-ABI argument — sets the thread
pointer, calls the user function, then calls `thread_exit`.
Unlike a process start, there is **no** System V argc/argv/auxv block
([sysv.md](sysv.md)) — a thread stack carries only the closure and its TLS block.
### Lifetime: exit, join, detach, stack reclaim
- **`thread_exit`** (no arguments) marks the task dead. The kernel releases the
task's resources, decrements the address-space refcount, and frees the task slot —
the user stack is not the kernel's to unmap; the joiner reclaims it.
- **`join`** is a dedicated `thread_join(tid)` syscall: the caller blocks in the
kernel (`joinThreadLocked`, woken by `wakeJoinersLocked` when the thread exits),
then `munmap`s the stack. (The plan staged join over a per-thread `exit_endpoint`
first, with a futex `completion` word as a Stage-2 refinement; neither shipped — the
dedicated syscall replaced both. `thread_spawn` still accepts an `exit_endpoint`
argument, which the runtime passes as `no_cap`.)
- **`detach`** relinquishes the join right: no one waits for the thread, and its
stack is reclaimed at process exit — kernel-side reclaim of a detached thread's
user stack stays deferred (as the intro notes), since `thread_exit` passes no stack
range.
### Futex, and the sync primitives on top
`futex_wait`/`futex_wake` are the one blocking primitive; `Mutex`, `Condition`, and
`Semaphore` are ordinary user-space state machines over an `atomic.Value(u32)` that
call the futex wrappers on the slow path — the same construction `std.Thread` uses,
so the algorithms port directly.
Keying: threads share an address space, so a **virtual address within that address space**
identifies a futex uniquely; the kernel keys its wait queue by `(address_space_root, virtual_address)`.
Keying by the **physical** address instead (translate `virtual_address -> physical_address` on entry) is a
deliberate forward door: it lets two *processes* share a futex through an
[shared-memory](../device-driver-development-guide/display-v2.md) region later, without changing the API. We start with the
private-per-address-space key and note the physical-key upgrade.
No spinning: a contended lock parks the task in the kernel and the core is free to run
other work or `hlt` ([halting.md](halting.md)). This is why futex is a locked
decision, not a "maybe later."
### TLS and `getCurrentId`
Per-thread thread-pointer TLS is in place (the `threadlocal` *compiler* layer is not —
see the intro). Two scoped pieces, as built:
- **`getCurrentId`** returns the kernel task id via the trivial `thread_self` syscall.
- **The thread pointer** is per-thread: `spawn` carves a small TLS block (an `fs:0`
self-pointer plus scratch) from the top of the thread's own stack, the trampoline
calls `set_thread_pointer` before any user code runs, and the scheduler saves and
restores the pointer per task across context switches. Full `threadlocal` support is
runtime+linker work on top of this, only if a consumer needs it. Nothing in the core
spawn/join/mutex path requires `threadlocal`.
### Build: multi-threaded codegen, opt-in
A binary opts in by being added with `addThreadedUserBinary` — as `addUserBinary`,
but the shared implementation builds it `single_threaded = false` — so atomics and
(later) TLS are real. Threads and atomics are unsound in a `single_threaded` image,
so a binary must opt in **before** it may call `Thread.spawn`. Everyone else
stays single-threaded and lean.
## Interaction with the rest of the kernel
- **Scheduler / SMP** ([scheduling.md](scheduling.md), [smp.md](smp.md)): a thread is
just another `Task` with an `address_space` shared with its siblings; the existing
per-core ready queues, priorities, and affinity apply unchanged. Threads of one
process can run on different cores simultaneously — that is the point.
- **Halting** ([halting.md](halting.md)): futex-parked waiters keep the "idle core
halts" property intact under lock contention — no busy-wait.
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): the contract is that
killing a process kills *all* its threads and only then drops the last address-space
ref — and the kernel now implements exactly that
([shared-fate-plan.md](shared-fate-plan.md)): every death path (`exit` from any
thread, a fault, `process_kill` aimed at any member id) fans out through the whole
group via a `dying` latch on the address space; the supervisor's one exit
notification — badged with the leader — fires only when the last member is gone.
A worker's voluntary `thread_exit` stays per-thread; the leader's is refused
(`-EPERM`).
- **Resilience** ([resilience.md](resilience.md)): by the same contract, a faulting
thread kills its whole process (shared fate); the supervisor restarts the
**process**, which respawns its threads from a known-good state — restart
granularity stays the process. The leader's recorded exit reason carries the fault
class even when a worker faulted, so restart policy is unchanged.
- **IPC — two consequences threads forced ([ipc.md](../device-driver-development-guide/ipc.md)):**
- *Handles do not cross threads.* The handle table lives on the `Task`
([scheduler.zig](../../system/kernel/scheduler.zig)), so a handle number is meaningful
only to the thread that created it — thread A's endpoint handle `3` is not thread B's.
A thread that needs to reach an endpoint another thread owns looks it up
(`ipc.lookup(service)`) to install its **own** handle to the same underlying endpoint.
This is how the display's mouse-listener thread reaches the compositor loop's endpoint
to poke it awake (docs/display.md).
- *IPC syscalls that touch shared kernel state now serialize under the big kernel lock.*
`create_ipc_endpoint`/`ipc_register`/`ipc_lookup` allocate from the kernel heap and
mutate the global service registry, endpoint refcounts, and handle tables. Those paths
were unlocked because a single-threaded process could not race itself; a multi-threaded
one can, from two cores at once. They now take `sync.enter()` like `call`/`reply_wait`/
`send` already did — the kernel heap has no lock of its own (heap.zig: "every kernel
entry takes the big kernel lock"), so the big lock is what keeps its callers serialized.
## Build-out plan (staged, each gate serial-checkable)
The ordered, `/loop`-runnable milestones live in
**[threading-plan.md](threading-plan.md)** (shaped like
[display-v2-plan.md](../device-driver-development-guide/display-v2-plan.md)): every milestone lands on its own and ends in
a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
`zig build test` for host unit tests). The stages below are the shape it expands.
- **Stage 0 — address-space refcount.** Refcount on the address-space root; teardown destroys
at zero. No API yet; nothing shares an address space, so refcount is 1 everywhere.
*Gate:* the full QEMU suite stays green (no regression) — proves the reframing is
invisible until used.
- **Stage 1 — spawn / join / detach.** `thread_spawn` + `thread_exit`, the trampoline,
stacks via `mmap`, join over the exit-endpoint (as built, join became the dedicated
`thread_join` syscall instead), the `addThreadedUserBinary` build opt-in.
*Gate:* two cases as built — `-Dtest-case=thread-spawn`, where a worker thread runs
in the caller's address space (a shared-memory write, observed by the main thread),
and `-Dtest-case=thread-join`, where N workers each atomically increment a shared
counter K times, the parent joins all N and asserts the total is exactly N × K —
the join case running multi-core (`smp` 4) to prove real parallelism.
- **Stage 2 — blocking synchronization.** `futex_wait`/`futex_wake` + `Futex`,
`Mutex`, `Condition`, `Semaphore`; optionally migrate join to a futex completion
word. *Gate:* `-Dtest-case=thread-mutex` — a bounded producer/consumer over a
`Mutex` + `Condition` moves K items with no lost wakeups and no busy-wait (assert
the consumer blocked, e.g. via a low idle tick count).
- **Stage 3 — polish.** Per-thread TLS / thread pointer and `threadlocal` (only if a
consumer needs it), `RwLock`/`WaitGroup` as demanded, and this doc's cases wired
into [test/qemu_test.py](../../test/qemu_test.py).
## Conventions
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym
abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New syscalls
extend [abi.zig](../../system/abi.zig) `SystemCall` + a `library/kernel` wrapper
([syscall.md](syscall.md)). `Thread` is a first-class runtime module, the same
way `process` ([process-lifecycle.md](process-lifecycle.md)) and `ipc`
are — user code never names a syscall.
## Non-goals
- **No preemptive user-space signals delivered to a specific thread.** Signals stay
process-scoped ([process-lifecycle.md](process-lifecycle.md)).
- **No thread priorities distinct from the process.** Threads inherit the process
priority; per-thread priority is a later question if it ever earns its keep.
- **No cross-process shared-memory futex yet** — the physical-address key leaves the
door open, but the first cut is private-per-address-space.
- **No `pthread`/POSIX surface.** The API is `std.Thread`-shaped Zig, nothing more.
## The self-hosting endgame
When danos becomes a real Zig target and we (eventually) add a danos backend to std
([zig-self-hosting.md](../zig-self-hosting.md)), `std.Thread` can sit *on top of* these
same kernel primitives — the danos `std.Thread.Impl` would call the very
`thread_spawn`/`futex_*` wrappers `Thread` already uses. Because
`Thread` was built API-compatible from day one, that transition swaps the
implementation, not a single call site. Designing to the std shape now is what makes
the later self-hosting lift cheap.
## Further reading
- [scheduling.md](scheduling.md), [smp.md](smp.md) — the task model these threads join.
- [resilience.md](resilience.md), [vision.md](../vision.md) — why isolation is the default
and threads are the exception.
- [syscall.md](syscall.md), [ipc.md](../device-driver-development-guide/ipc.md) — the private ABI and the messaging model
threads sit beside.
- [halting.md](halting.md) — the idle/halt property futex-backed blocking preserves.
- [zig-self-hosting.md](../zig-self-hosting.md) — the target this bends toward.
+118
View File
@@ -0,0 +1,118 @@
# Timers and time
Two different needs hide under the word "timer", and danos keeps them apart:
- **Reading the clock** — *what time is it?* A read of a free-running counter.
- **Waiting** — *wake me in N milliseconds*, or *notify me when a deadline passes.*
Both are answered by the **kernel**, because the kernel already owns a timer: it has
to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this
are built in [device-interrupts.md](../device-driver-development-guide/device-interrupts.md); the scheduler's blocking and
wait queues are in [scheduling.md](scheduling.md). This page is about the surface a
ring-3 program actually uses, and one deliberate absence: **there is no user-space time
service.**
## Why time is a syscall, not a service
The tempting microkernel move is to put a timer *driver* in user space and have
applications ask it for the time over IPC. For a **monotonic clock that is wrong** —
reading `now()` should never cost an IPC round trip. The kernel is already holding the
answer: it computes the current time every time it schedules, from the TSC, in a couple
of instructions. Surfacing that as a system call is pure mechanism; routing it through a
message to another process would be slower *and* redundant, and a device like the HPET
(uncacheable MMIO reads) is a particularly bad thing to read on every `now()`.
This is the same conclusion every serious system reaches: Linux and Zircon read the
counter in the vDSO, L4 exposes a clock field in a shared kernel page, seL4 reads the
cycle counter directly. None of them make a clock read an IPC. danos makes it a syscall.
That "from the TSC" hides a portability question, because the TSC is only a valid clock
when the CPU guarantees it is *invariant* and when every core's TSC is *synchronized*.
danos checks both — the invariant-TSC CPUID bit (`0x80000007` EDX[8], set on Intel and
AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET
counter when either fails. So `now()` stays accurate on a real Intel box, a real AMD box,
and inside a VM alike; only the source behind it differs. The mechanism is in
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md).
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
model; that role now lives in [drivers.md](../device-driver-development-guide/drivers.md), as documentation.) The one place
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
at the end; it is deliberately not built yet.
## The three system calls
Time and waiting are three entries in the small syscall table ([syscall.md](syscall.md)):
- **`clock` (#23)** → monotonic nanoseconds since boot. It only moves forward. Not
wall-clock: no date, no timezone. Backed by `architecture.nanos()` (TSC, scaled with a
128-bit intermediate so a long uptime can't overflow) — a few nanoseconds of
resolution, and just an `rdtsc` plus a multiply.
- **`sleep` (#3)** → block the caller for N milliseconds. The scheduler records a wake
deadline and the tick sweep wakes it (`scheduler.sleep`).
- **`timer_bind` (#31)** → arm a one-shot timer that, after N milliseconds, posts a
**timer notification** to an IPC endpoint. Unlike `sleep` it does **not** block: a
service can keep answering messages on the same endpoint while a deadline is pending.
This is the timed wait that stop-sequence escalation, hello deadlines, and restart
backoff are built from ([process-lifecycle.md](process-lifecycle.md),
[device-manager.md](../device-driver-development-guide/device-manager.md)).
The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space;
programs read the TSC through `clock` and get timed wakeups through `sleep`/`timer_bind`,
both riding the scheduler tick.
## `time` — the generic interface
Applications don't call the syscalls directly; they use `time`
(`library/kernel/time.zig`), a thin `Instant`/`Duration` layer over them — an ergonomic
front door, not new mechanism.
```zig
const time = @import("time");
const start = time.now(); // Instant — monotonic
doWork();
const took = start.elapsed(); // Duration
time.sleep(time.Duration.fromMillis(5)); // block ~5 ms
// A deadline delivered as a notification, so a service keeps serving meanwhile:
_ = time.after(endpoint, time.Duration.fromMillis(200));
```
- `Duration` is nanoseconds under the hood, with `fromNanos/fromMicros/fromMillis/
fromSeconds` and `asNanos/asMillis`. `ceilMillis` rounds *up* to the kernel's
millisecond granularity, so a sub-millisecond `sleep` never rounds down to zero and
returns early. All arithmetic saturates rather than wraps.
- `Instant` is a point on the monotonic clock: `since`, `elapsed`, `plus`, `reached` —
built for deadline loops (`while (!deadline.reached()) …`).
- `now()` / `monotonicNanos()` wrap `clock`. `available()` reports whether the clock is
calibrated at all (the kernel returns 0 until the TSC frequency is known, so a caller
that needs real time can treat 0 as "unavailable" rather than assume it advances).
- `sleep(d)` wraps `sleep`; `spin(d)` busy-polls `now()` for the sub-millisecond delays
the millisecond tick can't express; `after(endpoint, d)` wraps `timer_bind`.
The raw wrappers (`clock`, `sleepMillis`, `timerOnce`) and the ergonomic
`Instant`/`Duration` layer both live in the `time` module
(`library/kernel/time.zig`); the latter is what everyday code uses.
## Wall-clock time (not built)
Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and
measurement, useless for "what is the date?" Calendar time — a real-time clock, time
zones, leap seconds — is genuinely a **user-space** concern, and it *is* the case a time
service is for. It would be backed by an **RTC** driver (the CMOS real-time clock), not
the HPET, and exposed as a `CLOCK_REALTIME`-style service alongside the monotonic
syscall. It is deferred until something needs it; the monotonic clock the kernel already
owns covers every current use.
## Verifying it
`time`'s `Instant`/`Duration` arithmetic has unit tests that run on the host:
```
$ zig build test # includes library/kernel/time.zig
```
End to end, the proof the clock is real is that it *advances*: read `now()`, `sleep` a
`Duration`, read `now()` again, and the second reading is later — the kernel's timer
driving a ring-3 program with no service in between.
+216
View File
@@ -0,0 +1,216 @@
# The vDSO — the public system-call boundary
> **Status:** design note, not built. The runtime today issues raw `syscall`
> instructions from `library/kernel/system-call.zig` using the numbers in
> `system/abi.zig`. This note designs the layer that replaces that arrangement:
> a **kernel-supplied, C-ABI entry library** mapped into every process — the
> only supported way into the kernel — so the raw numbers can stay private,
> be renumbered at will, and eventually be randomised per boot.
## Why: the ABI danos promises, and the one it doesn't
`system/abi.zig` is the **private** kernel ↔ runtime contract. Its header says
so: the numbers are an implementation detail the runtime hides and may
renumber, the same split as libSystem over the XNU syscalls on macOS or win32
over the NT syscalls on Windows. Linux — with its world-visible, frozen
syscall table — is the outlier, not the norm.
That stance has consequences the moment binaries exist that we don't rebuild
ourselves:
1. **Third-party binaries** (docs/zig-self-hosting.md) must keep working across
kernel updates. If they contain raw `syscall` instructions with today's
numbers baked in, every renumbering breaks the world — the ABI would be
*de facto* public no matter what the header says. Go on macOS made exactly
this mistake: it issued XNU syscalls directly instead of going through
libSystem, and macOS updates repeatedly broke every Go binary until Go
switched to the library like everyone else.
2. **Not everything is Zig.** A Rust or C program can't import the danos Zig
modules. The public boundary has to be expressible in the one calling
convention every language speaks: the C ABI.
3. **Randomised syscall numbers** — a hardening option we want open — only
work if no user binary anywhere knows a number at build time. The binding
must happen at *load time*, from something the kernel controls.
All three point at the same well-known shape: a **vDSO** (virtual dynamic
shared object). The kernel carries a small blob of user-mode code, maps it
into every process at spawn, and that blob — not the application — contains
the `syscall` instructions. Fuchsia works exactly this way: its vDSO is the
*only* kernel entry, version-matched by construction because the kernel itself
injects it. Because the kernel and the blob ship as one artifact, there is
**no version skew, no loader, no search path, and no shared file on disk** —
which is what makes this the resilient way to have a private ABI
(docs/resilience.md), where a conventional `ld.so` + `/lib/libdanos.so`
arrangement would add a loader to every spawn and a single shared point of
failure.
The public danos ABI then has exactly two layers, neither of which is
`abi.zig`:
| Layer | Contract | Spoken by |
|-------|----------|-----------|
| **vDSO** | C-ABI functions, this note | every language's thin shim (the `system-call` module for Zig, a `-sys` crate for Rust, a header for C) |
| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](../file-system-development/vfs-protocol.md) is the first one documented) | any client that can lay out bytes |
Everything above those — the heap, `file_system`, the service harness — is
per-language convenience, compiled into each binary from source, exactly as
today. Nothing about the Zig runtime's shape changes; it just stops being the
*only* door.
## The blob
A single copy of the vDSO code lives in the kernel image (built by
`build.zig` as a tiny freestanding object, embedded like the AP trampoline).
At boot the kernel finalises it once — this is where randomised numbers would
be patched in — and thereafter maps the **same physical pages** read-execute
into every process's address space. The blob is:
- **Position-independent.** It is mapped at a per-process randomised base, so
it must be PIC (rip-relative addressing only — no relocations to process).
- **Stateless and re-entrant.** No writable data. Anything stateful belongs to
the process, not the vDSO.
- **Architecture-specific.** The x86-64 blob wraps `syscall`; an aarch64 blob
wraps `svc #0`. It lives beside the other per-architecture kernel sources
(`system/kernel/architecture/<arch>/`), selected the same way the
`architecture` module is (docs/architecture.md).
### Shape: a function table, not an ELF
A real `.so` with a dynamic symbol table is the conventional vDSO shape, but
linking against one at load time needs a dynamic linker in every binary —
machinery danos deliberately doesn't have. Instead the v1 shape is the
simplest thing that is still a stable contract — a **function-pointer table**
at the vDSO base:
```
offset 0 u64 magic 'danosVDS' — a mapped-the-wrong-thing guard
offset 8 u64 api_level incremented when the table grows
offset 16 u64 count number of table entries that follow
offset 24 u64 table[count] function pointers into the vDSO's own code
```
Table *indices* are the public constants (published in a C header,
`danos.h`), assigned once and append-only — the same discipline the IPC
protocols use for operation values. The pointers point at stubs inside the
blob; what those stubs put in `rax` is nobody's business but the kernel's.
A language shim binds in one step: read the base from the init block, check
the magic, keep the table pointer. Feature detection for a binary built
against older headers is `count`/`api_level` — a kernel never removes or
reorders entries.
(If danos ever grows a real dynamic linker, the same blob can additionally
present an ELF `dynsym` without breaking the table — Fuchsia's vDSO is
likewise both a mappable blob and a linkable `.so`. That is a later
convenience, not a requirement.)
### Delivery: the auxiliary vector
The kernel already builds a System V entry block — argc, argv, envp
terminator, **auxiliary vector** — on every new process's stack
(`buildEntryStack`, read by the `start` module). The vDSO base rides in a new
auxv entry, exactly Linux's `AT_SYSINFO_EHDR` move. No new syscall, no magic
address, and a language shim finds it the same portable way on every
architecture.
## The function surface
One table entry per kernel call, C ABI (System V AMD64), names prefixed
`danos_`. The current `SystemCall` set maps directly; integer arguments and
returns are `u64`, errors return as negative values exactly as today.
The calls that return two values in `rax:rdx` today — `dma_alloc`
(virtual_address + physical_address), `msi_bind` (address + data), `shared_memory_create` (virtual_address + handle),
`fs_resolve` (route tag + node token / backend handle) —
become functions returning a two-`u64` struct. The System V ABI returns a
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
C-ABI spelling of the existing convention, at zero cost. The one call that
returns *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
`rdx`, received capability in `r8`) — exceeds the two-register return: its
function returns a three-`u64` struct, which the ABI passes via a hidden
result pointer, so that one stub stores `rax`/`rdx`/`r8` through the pointer
after the `syscall` — a few instructions rather than one.
Grouped as `abi.zig` groups them:
| Group | Functions |
|-------|-----------|
| process | `danos_exit`, `danos_yield`, `danos_sleep`, `danos_spawn`, `danos_process_enumerate`, `danos_process_kill`, `danos_process_exit_reason`, `danos_process_subscribe`, `danos_process_signal`, `danos_signal_bind` |
| threads | `danos_thread_spawn`, `danos_thread_exit`, `danos_current_core`, `danos_futex_wait`, `danos_futex_wake`, `danos_thread_self`, `danos_thread_join`, `danos_set_thread_pointer` |
| memory | `danos_mmap`, `danos_munmap`, `danos_dma_alloc`, `danos_dma_free`, `danos_shared_memory_create`, `danos_shared_memory_map`, `danos_shared_memory_physical` |
| ipc | `danos_endpoint_create`, `danos_ipc_register`, `danos_ipc_lookup`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` |
| devices | `danos_device_enumerate`, `danos_device_claim`, `danos_device_register`, `danos_mmio_map`, `danos_irq_bind`, `danos_irq_ack`, `danos_msi_bind`, `danos_io_read`, `danos_io_write` |
| time | `danos_clock`, `danos_wall_clock`, `danos_timer_bind` |
| diagnostics | `danos_debug_write` (leveled, kernel-stamped records), `danos_klog_read`, `danos_klog_status` |
| filesystem naming | `danos_fs_resolve`, `danos_fs_node`, `danos_fs_mount`, `danos_fs_unmount` (naming only — file DATA still crosses the vfs-protocol IPC, see below) |
The constants that ride alongside the calls — mmap protection bits, DMA
flags, notification badge bits, `ExitReason`, `Signal`, well-known service
ids, `page_size`, the IPC message maximum — move to the public header too:
they are wire values a Rust program needs verbatim. What stays private in
`abi.zig` is exactly the thing the vDSO exists to hide: the `SystemCall`
numbers and the trap convention.
## Enforcement, and an honest threat model
Renumbering only has teeth if the kernel **refuses syscalls that don't come
from the vDSO**. The check is cheap: on kernel entry, the saved user `rip`
must lie inside the calling process's vDSO mapping; otherwise the process is
killed with a fault-class exit reason (its supervisor restarts or gives up,
docs/process-lifecycle.md — a foreign-syscall attempt is a bug or an attack,
never something to limp past). Fuchsia enforces exactly this.
What this buys, precisely:
- **ABI freedom** — the real prize. The numbers can change per release or per
boot and nothing outside the kernel image cares. The private ABI stays
actually private, permanently.
- **A single audited chokepoint** for kernel entry, per process, at a
randomised address.
- **Raised bar for exploits**: shellcode can't issue a hard-coded `syscall`;
it must first discover the per-process vDSO base (ASLR) and call through
it.
What it does *not* buy: an attacker with arbitrary code execution in a
process can still *call* the vDSO functions — they are mapped executable in
that process, and return-oriented chains reach them. Syscall randomisation is
hardening, not a security boundary; the security boundary remains the
capability model (what the process's endpoints and device claims let it do).
It is worth building anyway — for the ABI freedom first and the hardening
second — but the design should never be sold as more than that.
## Migration
Phased so every step ships alone (the M-milestone discipline):
1. **The blob + the table.** Build the vDSO, map it at spawn, deliver the
base via auxv. `library/kernel/system-call.zig` binds through the table when the
auxv entry is present, falls back to raw `syscall` when absent — the whole
tree keeps booting during the transition.
2. **Cut the system library over.** Delete the raw stubs; the `system-call`
module no longer imports the `SystemCall` numbers at all (`abi.zig`'s enum becomes
kernel-internal). The QEMU suite passing proves the table carries the
whole system.
3. **Enforce + randomise.** Add the `rip`-range check, then per-boot number
randomisation patched into the blob at kernel init. A test boots with
randomisation on and runs the full suite.
4. **The other languages.** Publish `danos.h`; a Rust `danos-sys` crate wraps
the table. This is also the seam `std.os.danos` calls through when the Zig
self-hosting fork lands (docs/zig-self-hosting.md) — the vDSO is what
makes that seam stable across kernel versions.
## What deliberately stays out
- **No dynamic linker, no `/lib/*.so`.** The vDSO is kernel-injected precisely
so danos binaries can stay fully static above it. Sharing *library code*
across processes stays what it is today: a service behind IPC, or source
compiled into each binary.
- **No file/device I/O in the vDSO.** The microkernel line doesn't move: the
vDSO wraps the same deliberately tiny table (docs/syscall.md). The kernel
resolves file NAMES (`fs_resolve` — the mount table moved in-kernel), but
file data is still the filesystem server's business over the vfs-protocol
IPC; the kernel never blocks on a userspace filesystem.
- **No fast-path user-mode implementations yet.** Linux's vDSO exists mostly
to answer `gettimeofday` without a kernel entry. `danos_clock` could one
day read the calibrated TSC in user mode the same way — the blob is where
such an optimisation would live — but that is an optimisation, not part of
this design's contract.