re-org docs
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
# OS Developer Guide
|
||||
|
||||
This document is for those who need to understand the architectural decisions behind the OS.
|
||||
|
||||
## Written in Zig?
|
||||
|
||||
The os was initially written in zig because it has excellent support for EFI. With zig, we could forgo using a third party bootloader, reducing the time to boot up the kernel. Following the "Zen of Zig", helped to produce the most readable codebase for an operating system ever created. So those, new to OS development could quickly get up to speed.
|
||||
|
||||
|
||||
@@ -0,0 +1,178 @@
|
||||
# ACPI: finding the tables (RSDP → RSDT/XSDT → SDTs)
|
||||
|
||||
ACPI describes the hardware the kernel can't assume — the interrupt controllers, the
|
||||
PCIe config window, the timer, the power registers — in a set of **system
|
||||
description tables** (SDTs). But before it can read any of them, danos has to *find*
|
||||
them, and they aren't at a fixed address. Getting there is a short chain of pointers,
|
||||
and this note explains it — in particular the question it's easy to trip on: **how
|
||||
does the [platform / device module](architecture.md) know where the RSDT is?**
|
||||
|
||||
Short answer: it doesn't receive the RSDT. The firmware hands over the **RSDP**, and
|
||||
the RSDT's address is a field *inside* the RSDP. The platform follows that pointer.
|
||||
|
||||
## The locator chain
|
||||
|
||||
```
|
||||
UEFI configuration table
|
||||
│ the loader reads the RSDP's physical address
|
||||
▼
|
||||
BootInformation.acpi_rsdp (u64, in the loader↔kernel handoff) system/boot-handoff.zig
|
||||
│ the kernel forwards the whole BootInformation
|
||||
▼
|
||||
platform.discover(boot_information, …) system/kernel/platform.zig
|
||||
│ reads boot_information.acpi_rsdp, hands it to the ACPI backend
|
||||
▼
|
||||
acpi.discover(rsdp_phys, …) system/kernel/acpi.zig
|
||||
│ dereferences the RSDP, reads the pointer it contains
|
||||
▼
|
||||
RSDP ──(a field in the struct)──► RSDT / XSDT ──► SDTs (MADT, MCFG, FADT, HPET, DSDT…)
|
||||
```
|
||||
|
||||
The **RSDP** (Root System Description Pointer) is the root of the whole ACPI tree.
|
||||
Its only job is to point at the root *table* — the **RSDT** (ACPI 1.0) or its 64-bit
|
||||
successor the **XSDT** (ACPI 2.0+) — which in turn lists every other SDT.
|
||||
|
||||
## Step 1 — the loader finds the RSDP
|
||||
|
||||
Only the firmware knows where ACPI lives, so the RSDP must be grabbed while UEFI is
|
||||
still up. `acpiRootSystemDescriptorPointer()` in `boot/efi.zig` walks the UEFI
|
||||
**configuration table** for the ACPI GUID and returns the vendor pointer — the same
|
||||
"grab it before `ExitBootServices`" pattern as the [framebuffer](framebuffer.md) and
|
||||
the [memory map](memory-map.md).
|
||||
|
||||
## Step 2 — the handoff: a physical address in `BootInformation`
|
||||
|
||||
The loader can't just call the device module: the bootloader binary and the kernel
|
||||
binary are compiled separately, and **the loader isn't linked against the `platform`
|
||||
module at all** (it imports only the `boot-handoff` contract and the
|
||||
`initial-ramdisk` module). So instead of a call, it
|
||||
deposits a value in the handoff struct:
|
||||
|
||||
```zig
|
||||
// boot/efi.zig — while boot services are still up
|
||||
.acpi_rsdp = if (acpiRootSystemDescriptorPointer()) |p| @intFromPtr(p) else 0,
|
||||
```
|
||||
|
||||
Two things about what crosses the boundary:
|
||||
|
||||
- **It's a *physical* address, not a Zig pointer.** The loader and kernel don't share
|
||||
an address space at the moment of the jump, so a raw `u64` physical address is the
|
||||
only thing that survives the handoff. `BootInformation.acpi_rsdp` is `0` when the firmware
|
||||
exposed no ACPI (e.g. a future device-tree machine, which would fill a different
|
||||
field instead — the kernel never learns which firmware booted it).
|
||||
- **The kernel can dereference it through the physmap.** The RSDP lives in
|
||||
ACPI-reclaim memory, which [paging.zig](paging.md) maps — along with the rest of
|
||||
RAM — into the higher-half **physmap** (there is no identity mapping; the low half
|
||||
belongs to user space). So by the time discovery runs,
|
||||
`@ptrFromInt(physicalToVirtual(rsdp_phys))` is a valid pointer.
|
||||
|
||||
This is the concrete form of the "capture the description pointer" step sketched in
|
||||
[discovery.md](discovery.md) — a plain `acpi_rsdp: u64` rather than a tagged handle,
|
||||
since x86 is the only backend wired up so far.
|
||||
|
||||
## Step 3 — the platform derives the RSDT from the RSDP
|
||||
|
||||
`acpi.discover` reinterprets the physical address as the RSDP struct, validates it,
|
||||
and then reads the root-table pointer *out of it*. Which pointer depends on the ACPI
|
||||
version, because the RSDP carries **both**:
|
||||
|
||||
```zig
|
||||
const rsdp: *const RootSystemDescriptionPointer = @ptrFromInt(physicalToVirtual(rsdp_phys));
|
||||
if (!std.mem.eql(u8, &rsdp.signature, "RSD PTR ")) return error.BadRsdpSignature;
|
||||
if (!checksumOk(@ptrFromInt(physicalToVirtual(rsdp_phys)), 20)) return error.BadRsdpChecksum;
|
||||
|
||||
if (rsdp.revision >= 2) {
|
||||
// ACPI 2.0+: use the 64-bit XSDT pointer (the 32-bit RSDT is deprecated)
|
||||
const xsdp: *const ExtendedSystemDescriptorPointer = @ptrFromInt(physicalToVirtual(rsdp_phys));
|
||||
try walkRoot(u64, xsdp.extended_system_descriptor_table_address, …);
|
||||
} else {
|
||||
// ACPI 1.0: use the 32-bit RSDT pointer
|
||||
try walkRoot(u32, rsdp.root_system_description_table_address, …);
|
||||
}
|
||||
```
|
||||
|
||||
- The **`revision`** byte selects the root table. `root_system_description_table_address`
|
||||
(32-bit, → RSDT) and `extended_system_descriptor_table_address` (64-bit, → XSDT) are
|
||||
ordinary fields of the RSDP/XSDP structs — the platform never *receives* the RSDT
|
||||
address, it *reads* it here. On QEMU q35 the RSDP is revision 2, so the XSDT path is
|
||||
taken.
|
||||
- The signature (`"RSD PTR "`) and one-byte checksum guard against a bad pointer before
|
||||
anything downstream trusts it.
|
||||
|
||||
## After the root table
|
||||
|
||||
`walkRoot` treats the RSDT/XSDT as an array of physical pointers — 32-bit entries for
|
||||
the RSDT, 64-bit for the XSDT — and hands each SDT to `handleTable`, which dispatches
|
||||
on its 4-byte signature: **MADT** (CPUs + IOAPIC), **MCFG** (PCIe ECAM), **FADT**
|
||||
(power registers, and the pointer to the DSDT), **HPET** (timer). That's where the
|
||||
firmware-agnostic [device model](discovery.md) gets populated; this note stops at the
|
||||
part that answers "where are the tables?" — everything past the RSDP is just following
|
||||
more pointers the tables themselves provide.
|
||||
|
||||
## ACPI events: the SCI, the power button, and GPEs (M21)
|
||||
|
||||
The tables above are static description; ACPI is also a *live* channel. Hardware
|
||||
raises the **SCI** (System Control Interrupt) — one shared, level-triggered line
|
||||
whose vector the FADT names — and the OS reads status registers to learn what
|
||||
happened: a fixed event like the power button, or a **General-Purpose Event**
|
||||
(GPE) whose handler is an AML method. Since [discovery](discovery.md) moved AML
|
||||
to ring 3, the event side lives there too, in the same **acpi service** — the
|
||||
device discoverer and the event source are one process, because both need the
|
||||
namespace and the port grant.
|
||||
|
||||
**The kernel hands the service what it needs and no more.** Reading PM1 event
|
||||
blocks and GPE blocks requires the FADT, which the kernel already parses for its
|
||||
own power register map (feeding reboot), the PM timer, and the SCI line — the
|
||||
kernel itself has no S5/poweroff path. Rather than re-parse, the kernel appends the **FADT as one
|
||||
more memory resource** on the `acpi-tables` node; the service tells it apart
|
||||
from the AML blob resources by signature — the FADT keeps its intact `"FACP"`
|
||||
header, while the blob resources are header-stripped bytecode that starts with
|
||||
no signature. The kernel's own FADT parse is untouched; the service reads the
|
||||
PM1 *event* blocks (which the kernel never parsed — it extracts only the PM1
|
||||
*control* register, and it is the service, not the kernel, that writes it for
|
||||
`\_S5`) and the GPE0/GPE1 blocks straight from its copy. The **SCI itself**
|
||||
arrives as the node's one `len == 1` irq resource (distinct from the broad
|
||||
`[0, 256)` window that covers children's legacy lines), which is how the service
|
||||
finds the line to `irq_bind`.
|
||||
|
||||
With those in hand the service enables ACPI mode (only if `SCI_EN` is clear —
|
||||
some firmwares boot with it already set), sets `PWRBTN_EN`, and on each SCI:
|
||||
|
||||
- **The power button** is a *fixed* event: a set `PWRBTN_STS` bit in PM1 status.
|
||||
The handler clears it (write-1-to-clear), logs the press, and publishes a
|
||||
[`power`](power.md) `power_button` event to subscribers.
|
||||
- **GPEs** are the general path: for each set-and-enabled GPE bit `n`, the
|
||||
service evaluates its `\_GPE._L%02X` (level) or `_E%02X` (edge) handler
|
||||
method, drains the **Notify** queue that method produced, maps each notified
|
||||
device to an event (battery, AC, lid, or a generic `notify` with its code),
|
||||
and clears the status bit. A missing handler method is not an error: the
|
||||
status bit is cleared and the event silently dropped. Making GPEs work
|
||||
required teaching the interpreter one opcode it never
|
||||
handled — `Notify` (`0x86`) — which it now folds into a bounded queue drained
|
||||
per evaluation; everything else a handler needs (field access, control flow,
|
||||
method calls) was already proven by the ring-3 `_STA`/`_CRS` work.
|
||||
|
||||
**How this is tested.** QEMU cannot raise GPEs deterministically on this config,
|
||||
so GPE/Notify correctness is proven by **host unit tests** — hand-encoded AML
|
||||
with a `Notify` inside a method body, run under `zig build test`. The QEMU
|
||||
`power-button` scenario proves the fixed-event path end to end: a QMP
|
||||
`system_powerdown` injects a real ACPI power-button press, and the service's SCI
|
||||
handler must log it. Battery/AC/lid mapping is interface-complete but validated
|
||||
on real hardware later; the embedded controller's `_Qxx` queries are out of
|
||||
scope.
|
||||
|
||||
The service surface these events are *published on* — subscription, the event
|
||||
vocabulary, and orderly shutdown — is the power service, [power.md](power.md).
|
||||
|
||||
## Related
|
||||
|
||||
- [efi.md](efi.md) — the loader that captures the RSDP before `ExitBootServices`.
|
||||
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam, and
|
||||
the ACPI-reclaim memory the RSDP lives in.
|
||||
- [discovery.md](discovery.md) — the broader (still-evolving) plan for turning these
|
||||
tables into one neutral device model shared with the ARM device-tree path, and how
|
||||
ACPI enumeration and events moved to the ring-3 acpi service.
|
||||
- [power.md](power.md) — the domain-named power service the ACPI event side publishes
|
||||
to (button, lid, battery) and its orderly-shutdown path into S5.
|
||||
- [architecture.md](architecture.md) — why the kernel reaches the device code through a `platform`
|
||||
module and never names ACPI directly.
|
||||
@@ -0,0 +1,109 @@
|
||||
# Architecture split
|
||||
|
||||
danos targets x86_64 today, but is meant to grow onto other systems later — a
|
||||
Raspberry Pi, say, which is AArch64 and has no UEFI. To keep that possible without
|
||||
a rewrite, CPU-specific kernel code lives behind a boundary: the generic kernel
|
||||
never names an architecture, and each architecture plugs in behind it.
|
||||
|
||||
## The seam is a build-time module named `architecture`
|
||||
|
||||
The mechanism is deliberately boring — no vtables, no function-pointer tables, no
|
||||
runtime dispatch. `build.zig` exposes one architecture's code as a module called
|
||||
`architecture`:
|
||||
|
||||
```zig
|
||||
const architecture_module = b.addModule("architecture", .{
|
||||
.root_source_file = b.path("system/kernel/architecture/x86_64/cpu.zig"),
|
||||
});
|
||||
```
|
||||
|
||||
and the generic kernel imports it by that name:
|
||||
|
||||
```zig
|
||||
const architecture = @import("architecture");
|
||||
// ...
|
||||
architecture.halt(); // never says "x86_64"
|
||||
```
|
||||
|
||||
Adding a second architecture is then a build-time choice: create
|
||||
`system/kernel/architecture/aarch64/`, and point the `architecture` module at it when the target CPU is
|
||||
AArch64. `kernel.zig` and `console.zig` don't change. **That compiler-checked module
|
||||
boundary _is_ the architecture interface** — when a new architecture is missing a function
|
||||
the generic kernel calls, the build fails and names exactly what's missing.
|
||||
|
||||
## What's arch-specific vs generic
|
||||
|
||||
The split follows a simple test: does it name a CPU instruction, a hardware
|
||||
register, or a memory-management structure? If so, it's arch-specific.
|
||||
|
||||
| Arch-specific — `system/kernel/architecture/x86_64/` | Generic — kernel core |
|
||||
|---|---|
|
||||
| `cpu.zig`: CPU state, trap-frame accessors, paging, SMP | `console.zig` — pure pixel math, framebuffer drawing |
|
||||
| `gdt.zig`, `idt.zig`, `tss.zig` — descriptor tables | `kernel.zig` — kernel orchestration, scheduler, IPC |
|
||||
| `paging.zig` — page-table setup and management | `process.zig` — process lifecycle, address spaces |
|
||||
| `apic.zig`, `ioapic.zig` — interrupt controllers | `scheduler.zig` — task scheduling and context switch |
|
||||
| `serial.zig`, `io.zig` — UART, I/O primitives | `vfs.zig` — filesystem abstraction |
|
||||
| `isr.s`, `smp.zig` — exceptions, AP bring-up, context switch | `irq.zig`, `ipc*.zig` — interrupt dispatch, messaging |
|
||||
| `linker.ld` — kernel link layout, load address | |
|
||||
|
||||
Notice the framebuffer console is *generic*: it just writes pixels into whatever
|
||||
framebuffer it's handed, so it needs no per-arch version. Most of the kernel
|
||||
should end up on the generic side; the architecture module stays small.
|
||||
|
||||
## Two axes, kept separate
|
||||
|
||||
There are really two independent questions, and it's worth not conflating them:
|
||||
|
||||
- **CPU architecture** (x86_64 vs AArch64): instructions, MMU, interrupts →
|
||||
`system/kernel/architecture/<cpu>/`.
|
||||
- **Boot protocol** (UEFI vs Raspberry Pi firmware + device tree): handled
|
||||
*separately*, because loaders are their own binaries. `boot/efi.zig` builds
|
||||
`BOOTX64.efi`, a distinct executable from the kernel ELF. On a Pi there is no
|
||||
separate loader at all — the firmware jumps straight into the kernel with a
|
||||
device-tree pointer, so that entry work would live in the AArch64 architecture code.
|
||||
Either path converges on the same neutral [`BootInformation`](memory-map.md).
|
||||
|
||||
## Current x86_64 contents
|
||||
|
||||
- **`system/kernel/architecture/x86_64/cpu.zig`** — the `architecture` module root. Exposes the trap-frame
|
||||
`CpuState` and accessors, `init()` (bring up the descriptor tables), `enablePaging()`,
|
||||
`enterUser()`/`userExit()` for ring-0 ↔ ring-3 transitions, address-space management, and SMP
|
||||
entry points (see [halting.md](halting.md), [interrupts.md](interrupts.md), [paging.md](paging.md),
|
||||
[scheduling.md](scheduling.md)).
|
||||
- **`system/kernel/architecture/x86_64/gdt.zig`** / **`idt.zig`** / **`tss.zig`** — the GDT, IDT and
|
||||
TSS plus CPU-exception handling (see [interrupts.md](interrupts.md)).
|
||||
- **`system/kernel/architecture/x86_64/paging.zig`** — the kernel's page tables and address-space
|
||||
management (see [paging.md](paging.md)).
|
||||
- **`system/kernel/architecture/x86_64/apic.zig`** / **`ioapic.zig`** — the Local APIC, its timer,
|
||||
and the I/O APIC for device interrupts (see [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)).
|
||||
- **`system/kernel/architecture/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's
|
||||
machine-readable log channel, see [testing.md](../testing.md)) and the shared port-I/O + MSR primitives.
|
||||
- **`system/kernel/architecture/x86_64/smp.zig`** / **`per-cpu.zig`** — application-processor bring-up
|
||||
and per-CPU state (GS base, system-call entry point, see [scheduling.md](scheduling.md)).
|
||||
- **`system/kernel/architecture/x86_64/isr.s`** — the exception stubs, the `lgdt`/`lidt`/`ltr` load
|
||||
helpers, ring-0 ↔ ring-3 transitions, and the context switch — real assembly, since Zig inline asm can't
|
||||
express them (see [scheduling.md](scheduling.md)).
|
||||
- **`system/kernel/architecture/x86_64/linker.ld`** — the kernel link layout (fixed low load
|
||||
address, one PT_LOAD per permission set).
|
||||
|
||||
The kernel entry point `_start` lives in the architecture-specific `isr.s` (x86_64 here).
|
||||
On x86_64 it sets up the kernel stack in BSS and jumps to `kmain()` in `kernel.zig`.
|
||||
This is already per-architecture — an AArch64 port would have its own `isr.s` entry
|
||||
that parses the device-tree pointer from a register and jumps to the same `kmain()`.
|
||||
The entry interface is minimal and emerges naturally from the [boot-handoff](memory-map.md)
|
||||
contract both share.
|
||||
|
||||
## The discipline
|
||||
|
||||
The thing that makes this help rather than hurt: **only extract what's provably
|
||||
architecture-specific, and let the interface emerge with the second
|
||||
implementation.** With a single architecture you're guessing at the seam, and a
|
||||
wrong guess encoded as elaborate abstraction is expensive to undo. So:
|
||||
|
||||
- Move code into `architecture/` only when it genuinely names CPU-specific machinery.
|
||||
- Grow the `architecture` surface one function at a time, as steps need it.
|
||||
- Don't pre-design the interrupt or paging interfaces before writing them.
|
||||
|
||||
Directory hygiene is cheap and reversible; premature abstraction is neither. When
|
||||
architecture #2 lands and something doesn't fit, reshaping a few hundred lines is nothing —
|
||||
unwinding an abstraction empire is not.
|
||||
@@ -0,0 +1,125 @@
|
||||
# ARM targets (`arm` and `aarch64`)
|
||||
|
||||
danos aims to run on Raspberry Pi hardware eventually. "ARM" isn't one target,
|
||||
though — the Pis span **two different CPU architectures** (32-bit `arm` and 64-bit
|
||||
`aarch64`) and (stock) a different boot protocol from x86-64's UEFI. **danos targets
|
||||
`aarch64` only** (see the decision below); the `arm`/`aarch64` distinction still
|
||||
matters for understanding why. This page maps the landscape so the
|
||||
[architecture split](architecture.md) and build system can be planned for it.
|
||||
|
||||
## `arm` vs `aarch64` — 32-bit vs 64-bit
|
||||
|
||||
- **`arm`** = **32-bit** ARM (the *AArch32* state, A32/T32 instruction sets).
|
||||
ARMv7 and earlier, plus the 32-bit compatibility mode of newer cores. 16 × 32-bit
|
||||
registers.
|
||||
- **`aarch64`** = **64-bit** ARM (the *AArch64* state, A64 instruction set), from
|
||||
**ARMv8-A** on. Also called **arm64**. 31 × 64-bit registers, a fixed 32-bit
|
||||
instruction width, a redesigned exception model — *not* a widening of A32, a clean
|
||||
new ISA.
|
||||
|
||||
They are as different from each other as either is from x86-64: separate registers,
|
||||
page-table formats, and calling conventions. Each needs its own `system/kernel/architecture/<name>/`.
|
||||
|
||||
## The Raspberry Pi models
|
||||
|
||||
| Model | SoC | Core | Architecture | danos target |
|
||||
|-------|-----|------|--------------|--------------|
|
||||
| Pi Zero / Zero W | BCM2835 | ARM1176JZF-S | ARMv6, 32-bit only | `arm` (not planned) |
|
||||
| **Pi Zero 2 W** ← target | BCM2710 | Cortex-A53 | ARMv8-A, 64-bit | **`aarch64`** |
|
||||
| **Pi 3 / 3B+** | BCM2837 | Cortex-A53 | ARMv8-A, 64-bit | **`aarch64`** |
|
||||
| **Pi 4** | BCM2711 | Cortex-A72 | ARMv8-A, 64-bit | **`aarch64`** |
|
||||
| **Pi 5** | BCM2712 | Cortex-A76 | ARMv8.2-A, 64-bit | **`aarch64`** |
|
||||
|
||||
> **Decision: `aarch64` only.** The target small board is a **Pi Zero 2 W** (BCM2710,
|
||||
> Cortex-A53) — which is **`aarch64`**, *not* the original Zero W's 32-bit ARMv6. So
|
||||
> every ARM board danos targets (Zero 2 W and Pi 3-5) is `aarch64`, and the 32-bit
|
||||
> `arm`/ARMv6 backend is **not planned** — one ARM CPU port, not two. The original
|
||||
> Zero W (ARMv6) would only re-enter scope if that specific older board were ever
|
||||
> needed; the row above is kept only to explain the distinction.
|
||||
|
||||
## Booting: UEFI is not x86-only
|
||||
|
||||
The boot protocol is a **separate axis** from the CPU (see [architecture.md](architecture.md)):
|
||||
|
||||
- **UEFI** exists for ARM too — ARM servers require it (SBSA/SBBR), QEMU boots it
|
||||
with **AAVMF** (the AArch64 build of the same EDK2 firmware as x86's OVMF), and
|
||||
the Pi can even run it with community UEFI firmware. Under UEFI the handoff is the
|
||||
*same* as x86-64: system table, boot services, memory map, GOP framebuffer — so
|
||||
the loader logic largely carries over.
|
||||
- **Device tree / firmware boot** — the **stock** Raspberry Pi firmware (VideoCore
|
||||
bootloader) is *not* UEFI: it loads the kernel and jumps to it with a **device-tree
|
||||
blob (DTB)** pointer. Both the Zero W and stock Pi 3-5 boot this way.
|
||||
|
||||
Note that even under UEFI on ARM, the OS still gets its hardware description from
|
||||
**ACPI or a device tree** (often the DTB passed via a UEFI configuration table). So
|
||||
"UEFI on ARM" doesn't remove the device tree — UEFI gives you memory + framebuffer;
|
||||
the DTB/ACPI tells you what devices exist.
|
||||
|
||||
## What danos needs, layer by layer
|
||||
|
||||
- **One CPU arch module: `system/kernel/architecture/aarch64/`** — covering the Zero 2 W and Pi 3-5,
|
||||
providing the same `arch` interface as x86_64: `halt`, context switch,
|
||||
interrupt/exception vectors, page tables, a UART, a timer. No `system/kernel/architecture/arm/` is
|
||||
planned (see the decision above), so there's a single ARM backend to write.
|
||||
- **A device-tree boot path.** Since stock Pis boot via DTB, danos needs an entry
|
||||
that parses the DTB's `/memory` and `/reserved-memory` into the neutral
|
||||
[`MemoryMap`](memory-map.md) — the same neutral handoff `efi.zig` produces, just
|
||||
from a different source. This is where keeping boot-protocol knowledge on the
|
||||
loader side (as we did for the UEFI memory-map classification) pays off.
|
||||
- **The UEFI loader mostly carries over.** `boot/efi.zig` is largely
|
||||
boot-*protocol* code (`std.os.uefi` protocol calls), not x86 code. Its truly
|
||||
x86-specific bits are the ELF machine check (`.X86_64`), the SysV calling
|
||||
convention for the kernel jump, the `hlt` park on failure — and, the substantial
|
||||
one, the bootstrap page tables: `buildBootstrapTables` builds x86-64 4-level
|
||||
tables (PML4/PDPT/PD index shifts, x86 PTE bits, 2 MiB leaves) and hands the
|
||||
kernel a CR3. Page-table formats are per-architecture (see above), so an
|
||||
`aarch64` loader keeps the protocol code but rewrites that builder in the
|
||||
aarch64 translation-table format. Even so, an `aarch64`-UEFI target (QEMU
|
||||
`virt` + AAVMF) reuses most of it — which makes **aarch64-UEFI the easiest
|
||||
second target**, easier than the device-tree Pi.
|
||||
|
||||
## Pi hardware quirks (for when we port)
|
||||
|
||||
The Pi is not a "standard" ARM platform — expect Broadcom-specific peripherals:
|
||||
|
||||
- **Peripheral base moves per SoC**: `0x2000_0000` (BCM2835, Zero W),
|
||||
`0x3F00_0000` (BCM2837, Pi 3), `0xFE00_0000` (BCM2711, Pi 4), different again on
|
||||
Pi 5. Everything below is an offset from it.
|
||||
- **UART**: a **PL011** (at base + `0x20_1000`) plus a mini-UART; on some boards the
|
||||
PL011 is wired to Bluetooth, so which one is the console varies. This is the
|
||||
`aarch64`/`arm` equivalent of our x86 [COM1 serial](../testing.md).
|
||||
- **Interrupt controller**: *not* a standard ARM GIC on the older parts — the Zero W
|
||||
and Pi 3 use Broadcom's own ARMCTRL controller (Pi 3 adds a per-core "local"
|
||||
controller for timers/mailboxes). The **Pi 4 and 5 do have a GIC-400**. So the
|
||||
interrupt backend differs even within the `aarch64` Pis.
|
||||
- **Timer**: the ARM generic timer (`CNTPCT`/`CNTFRQ`) on ARMv8, or the BCM system
|
||||
timer — the counterpart to our calibrated LAPIC/TSC clock.
|
||||
|
||||
## Building each (intended)
|
||||
|
||||
`build.zig` currently pins the kernel to `x86_64`; supporting these means selecting
|
||||
the target and `arch` module together (e.g. a `-Darch=` option). The Zig target
|
||||
queries would be roughly:
|
||||
|
||||
- **Pi Zero W**: `.cpu_arch = .arm`, `.cpu_model = arm1176jzf_s`, `.os_tag = .freestanding`
|
||||
- **Pi 3**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a53`, `.os_tag = .freestanding`
|
||||
- **Pi 4**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a72`
|
||||
- **Pi 5**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a76`
|
||||
|
||||
## Testing in QEMU
|
||||
|
||||
Two routes, mirroring how we test x86-64 with OVMF:
|
||||
|
||||
- **Board emulation**: `qemu-system-aarch64 -machine raspi3b` (and `raspi4b` on
|
||||
recent QEMU) for the Pi 3/4; `qemu-system-arm -machine raspi0`/`raspi1ap` for the
|
||||
ARMv6 Zero-class board — closest to real hardware, device-tree boot.
|
||||
- **Generic aarch64-UEFI**: `qemu-system-aarch64 -machine virt` + AAVMF — the
|
||||
cleanest way to bring up the `aarch64` kernel via the reused UEFI loader before
|
||||
tackling Pi-specific boards. A future `run-aarch64` build step would use this.
|
||||
|
||||
## Related
|
||||
|
||||
- [architecture.md](architecture.md) — the arch-module boundary these targets plug into, and the
|
||||
CPU-arch vs boot-protocol "two axes".
|
||||
- [efi.md](efi.md) — the UEFI loader that carries over to aarch64-UEFI.
|
||||
- [vision.md](../vision.md) — why isolated, portable-across-architectures is the goal.
|
||||
@@ -0,0 +1,259 @@
|
||||
# Device discovery (ACPI / device tree), the agnostic way
|
||||
|
||||
"Device discovery" is how the kernel learns **what hardware exists and where** — the
|
||||
MMIO addresses, IRQ numbers, CPU count, and interrupt controller it can't just
|
||||
assume. On x86 that description comes from **ACPI** tables; on ARM from a **device
|
||||
tree** (DTB). This note is a design plan, not built yet: *when* danos should tackle
|
||||
it, and *how* to keep it architecture-agnostic — the same discipline the
|
||||
[memory map](memory-map.md) and [architecture split](architecture.md) already follow.
|
||||
|
||||
## What the kernel assumed when this plan was written
|
||||
|
||||
At the time danos discovered almost nothing — it coasted on legacy PC fixtures that
|
||||
are guaranteed to exist under QEMU + UEFI:
|
||||
|
||||
- `system/kernel/architecture/x86_64/apic.zig` assumed the **Local APIC** at the
|
||||
default `0xFEE0_0000` and calibrated its timer against the **PIT** (the legacy
|
||||
8254). (Today the PIT is the *last-resort* reference: calibration prefers the
|
||||
CPUID-reported TSC frequency, then the HPET, then the ACPI PM timer.)
|
||||
- `system/kernel/architecture/x86_64/serial.zig` hardcoded **COM1** at I/O port
|
||||
`0x3F8`. (Today `0x3F8` is only the default: the kernel loopback-probes the UART
|
||||
and parses ACPI's SPCR table to target the firmware's actual debug port.)
|
||||
- The framebuffer and memory map come from **UEFI** — that *is* discovery, just done
|
||||
by the firmware and handed over, not read from ACPI.
|
||||
|
||||
This worked only because PC-compatible hardware promises those legacy pieces exist at
|
||||
those addresses. It was a crutch, and it did not travel.
|
||||
|
||||
## The forcing functions: when to build it
|
||||
|
||||
Two things drive the need, and they set the timing:
|
||||
|
||||
1. **The second architecture makes it mandatory.** ARM has *no* legacy fixtures —
|
||||
no PIT, no fixed serial port, no standard interrupt controller address. You can't
|
||||
find the UART to print a character without reading the device tree. So on x86 we
|
||||
can defer discovery a long time (until we want the IOAPIC, PCIe, or SMP), but on
|
||||
**aarch64 it's required to boot at all**. The [aarch64 port](arm.md) is what
|
||||
forces the issue.
|
||||
|
||||
2. **Isolated user-space drivers need it.** In the [microkernel vision](../vision.md),
|
||||
drivers live in user space — but something has to enumerate the hardware and hand
|
||||
each driver its MMIO regions and IRQs. That enumeration *is* device discovery. So
|
||||
discovery is a prerequisite for real drivers, **not** for user mode itself.
|
||||
|
||||
The conclusion on timing: **don't build full discovery before user mode.** User mode
|
||||
+ address-space isolation needs none of it; the current assumptions are fine there.
|
||||
Build the agnostic discovery layer **when the aarch64 port starts** — because that's
|
||||
when a second, real implementation makes "agnostic" honest.
|
||||
|
||||
## Why build it *with* the second arch, not before
|
||||
|
||||
The same lesson as the arch split: an abstraction with only one implementation
|
||||
quietly bends to that implementation. Build "agnostic discovery" x86-only first and
|
||||
you'll get an ACPI-shaped interface with device tree bolted on afterward. Build it
|
||||
when aarch64 lands and the small, clean DTB parser pulls the abstraction toward the
|
||||
right neutral shape, which the x86 side then fills. Design it against two backends or
|
||||
it isn't really agnostic.
|
||||
|
||||
## The one cheap step to take sooner
|
||||
|
||||
Have the **loader capture the description pointer** into `BootInformation` — a
|
||||
neutral handle, no parsing:
|
||||
|
||||
```zig
|
||||
pub const HardwareInfo = extern struct {
|
||||
kind: enum(u32) { none, acpi, device_tree },
|
||||
addr: u64, // ACPI RSDP, or the DTB blob
|
||||
};
|
||||
```
|
||||
|
||||
On x86-UEFI that's the ACPI **RSDP**, read from the UEFI configuration table *before*
|
||||
`ExitBootServices` — the same "grab it before exit" pattern as the framebuffer and
|
||||
memory map (already flagged in [memory-map.md](memory-map.md)). On ARM it's the DTB
|
||||
pointer the firmware passes. This keeps the door open for near-zero cost without
|
||||
committing to the parser.
|
||||
|
||||
## What "agnostic" looks like
|
||||
|
||||
The happy accident: **the device-tree data model is already a good neutral
|
||||
representation.** A DTB is a tree of nodes, each with:
|
||||
|
||||
- a **`compatible`** string — what the device is,
|
||||
- **`reg`** — its MMIO base(s) and size(s),
|
||||
- **`interrupts`** — its IRQ number(s),
|
||||
- other properties.
|
||||
|
||||
Even OSes running on ACPI hardware normalize into a unified device model shaped like
|
||||
this. So the neutral layer is a **device model** — "here are the devices, each with a
|
||||
type, MMIO regions, and IRQs" — fed by two backends behind it:
|
||||
|
||||
- a **DTB parser** (ARM) — a few hundred lines against a well-specified binary format;
|
||||
- **static ACPI + PCI enumeration** (x86).
|
||||
|
||||
### The caveat that sets the effort: "ACPI" ≠ "AML interpreter"
|
||||
|
||||
A full ACPI namespace is **AML** (ACPI Machine Language) bytecode, and writing an AML
|
||||
interpreter is an enormous undertaking. **We skip it.** Interrupt routing and PCIe
|
||||
come entirely from the *static* tables:
|
||||
|
||||
- **MADT** — the APICs and the **IOAPIC** (interrupt routing),
|
||||
- **MCFG** — PCIe configuration space (then walk the PCI bus to enumerate devices),
|
||||
- **FADT**, **HPET** — power/reset and a precise timer.
|
||||
|
||||
Walking PCI config space finds most devices without any AML. So the x86 static-ACPI
|
||||
backend is roughly comparable in scope to the DTB parser; it's full AML that's the
|
||||
monster, and it isn't on the path.
|
||||
|
||||
## Where it sits in the seam
|
||||
|
||||
Discovery layers onto the existing loader↔kernel split cleanly:
|
||||
|
||||
| Layer | Responsibility | x86 | ARM |
|
||||
|-------|----------------|-----|-----|
|
||||
| **Capture** (loader) | grab the description pointer | RSDP from UEFI config table | DTB pointer from firmware |
|
||||
| **Parse** (kernel, per-mechanism) | pointer → neutral device model | static ACPI + PCI | DTB parser |
|
||||
| **Consume** (generic) | use the device model | — same code — | — same code — |
|
||||
|
||||
Boot-protocol-specific capture, mechanism-specific parse, generic consumption —
|
||||
exactly like the memory map, where the loader classifies and the kernel just sees
|
||||
neutral regions.
|
||||
|
||||
## Where it lives in a microkernel
|
||||
|
||||
For isolated drivers, discovery is not one lump — it splits by *who needs it and
|
||||
when*:
|
||||
|
||||
- **Minimal, in-kernel: the interrupt controller and the timer.** The IOAPIC/GIC and
|
||||
the timer are needed *before* user space exists (scheduling and preemption depend
|
||||
on them), so the kernel must parse at least these from ACPI/DTB itself. This is also
|
||||
the **first real consumer** of discovery — the first genuine reason to read MADT or
|
||||
the DTB is "where is the interrupt controller and how do I route an IRQ?" It's the
|
||||
point where x86 finally graduates from the PIT/legacy-LAPIC assumptions.
|
||||
- **User-space enumeration: a device-manager server.** Everything else — PCI devices,
|
||||
peripherals — is parsed (or queried from the kernel's parse) by a privileged
|
||||
user-space server that hands each driver process its MMIO regions and IRQ rights
|
||||
over [IPC](../device-driver-development-guide/ipc.md). Combined with **interrupts-as-messages** (an IRQ delivered to a
|
||||
driver as a message on a channel — a natural extension of the wait queues and
|
||||
channels already built), that's what makes drivers genuinely isolated.
|
||||
|
||||
So the device-manager server depends on user mode + IPC; only the interrupt-controller
|
||||
slice is unavoidably in-kernel.
|
||||
|
||||
## A nice ARM contrast
|
||||
|
||||
On ARMv8 the generic timer exposes its frequency directly via the `CNTFRQ` register —
|
||||
no calibration needed. That's cleaner than the x86 side, where we measure the LAPIC
|
||||
and TSC against the PIT because nothing tells us their frequency (see
|
||||
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). Discovery on ARM hands you more for
|
||||
free; discovery on x86 is partly about *finding* what ARM just tells you.
|
||||
|
||||
## Suggested ordering
|
||||
|
||||
1. **Now (cheap):** plumb the neutral `HardwareInfo` pointer through the loader into
|
||||
`BootInformation`. No parser yet.
|
||||
2. **Next milestone unchanged:** user mode + address-space isolation — needs no
|
||||
discovery.
|
||||
3. **With the aarch64 port:** build the agnostic discovery layer, **DTB first** (the
|
||||
forcing function), then x86 static-ACPI to fill the same model — starting with the
|
||||
**interrupt controller + timer**.
|
||||
4. **Then:** a user-space device-manager server + interrupts-as-messages → real
|
||||
isolated drivers (keyboard first).
|
||||
|
||||
## Related
|
||||
|
||||
- [acpi.md](acpi.md) — the built x86 side of the "capture" step: how the loader grabs
|
||||
the RSDP and the platform follows it to the RSDT/XSDT and the SDTs.
|
||||
- [arm.md](arm.md) — the aarch64 target that forces genuine discovery (DTB, GIC).
|
||||
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam,
|
||||
and the note about grabbing the RSDP before `ExitBootServices`.
|
||||
- [device-interrupts.md](../device-driver-development-guide/device-interrupts.md) — the LAPIC/timer bring-up that
|
||||
discovery will eventually feed (IOAPIC, real IRQ routing).
|
||||
- [ipc.md](../device-driver-development-guide/ipc.md) — the channels that interrupts-as-messages and the device manager
|
||||
will ride on.
|
||||
- [vision.md](../vision.md) — why drivers belong in isolated user space at all.
|
||||
|
||||
## Update (M19.3, 2026-07-13): PCI enumeration left the kernel
|
||||
|
||||
The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO
|
||||
apertures derived from the memory map's holes, bus range, and the 16-bit I/O
|
||||
window). The per-function walk moved to the ring-3 `pci-bus` driver
|
||||
([device-manager.md](../device-driver-development-guide/device-manager.md)): it claims the bridge, repeats the
|
||||
ECAM scan through its mmio grant, and `device_register`s what it finds, which
|
||||
the device manager mirrors and matches. The ACPI namespace walk follows in M20;
|
||||
the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side.
|
||||
|
||||
## Update (M20.3, 2026-07-13): ACPI enumeration left the kernel too
|
||||
|
||||
The kernel no longer folds the AML namespace's Device objects into the device
|
||||
tree. It still parses the *static* tables (MADT for SMP, HPET for the tick, MCFG
|
||||
for the host bridge, FADT); at this point it also still built the AML namespace —
|
||||
but only to read the `\\_S5` sleep type for poweroff. (That remnant is gone too:
|
||||
the kernel now runs no AML at all — soft-off belongs to the acpi service, and the
|
||||
kernel keeps only the AML-free reboot path.) Device discovery is the ring-3 **acpi
|
||||
service** ([device-manager.md](../device-driver-development-guide/device-manager.md)): it claims the `acpi-tables`
|
||||
node the kernel publishes (the AML blobs, a broad io_port grant, the SCI),
|
||||
re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`,
|
||||
and registers + reports each `_HID` device — the device manager matches drivers
|
||||
(ps2-bus) from those reports. With M19's pci-bus driver, discovery now runs
|
||||
entirely in user space, anchored on two kernel-seeded nodes: the host bridge and
|
||||
the acpi-tables node. (The kernel's static-table parse also seeds the processor,
|
||||
interrupt-controller, and HPET timer nodes, and it publishes the boot
|
||||
framebuffer as a claimable display node — but no *enumeration* happens in
|
||||
ring 0.)
|
||||
|
||||
## Discovery is a swappable process per firmware (M19–M20)
|
||||
|
||||
Moving PCI and ACPI enumeration out of ring 0 was not just a relocation — it
|
||||
made discovery **firmware-neutral by construction**, which is the whole reason
|
||||
to do it before the second architecture rather than after. Everything at and
|
||||
above the [device-manager](../device-driver-development-guide/device-manager.md) protocol — descriptors,
|
||||
containment, reports, matching, supervision — is generic and may never become
|
||||
x86-specific. Discovery is the single firmware-specific piece, and it is
|
||||
isolated as **one swappable process per firmware**:
|
||||
|
||||
- **x86** boots describe hardware with ACPI, so the discoverer is the **acpi
|
||||
service** ([acpi.md](acpi.md)): it claims the `acpi-tables` node and runs AML.
|
||||
- **The Raspberry Pis** hand over a flattened device tree, so the discoverer is
|
||||
an **fdt service**: it claims a `devicetree-blob` node and walks the tree —
|
||||
pure data, no bytecode, so it needs neither a port grant nor an interpreter,
|
||||
strictly simpler than ACPI. (A placeholder until the [aarch64](arm.md)
|
||||
bring-up fills it in.)
|
||||
|
||||
The device manager spawns the discoverer under the **neutral ramdisk name
|
||||
`discovery`** and never learns which firmware it is on; the build's
|
||||
`-Ddiscovery=acpi|fdt` option fills that slot (x86 defaults to `acpi`, the
|
||||
aarch64 target flips the default when it lands). The manager owns the device
|
||||
tree as *data* and touches no hardware, ever — firmware bytecode runs only
|
||||
inside the crashable, supervised discoverer, so an AML fault can never take
|
||||
down the supervisor.
|
||||
|
||||
Two consequences of neutrality bind on later work:
|
||||
|
||||
- **Cross-firmware surfaces are named by domain, not firmware.** System power is
|
||||
a [`power`](power.md) protocol, not an "ACPI events" protocol: on x86 the acpi
|
||||
service registers it, on ARM a PSCI/mailbox service registers the same
|
||||
`ServiceId.power`, and subscribers never learn the difference.
|
||||
- **Identity must widen before the fdt service exists.** `DeviceDescriptor`'s
|
||||
8-byte `hid` holds an EISA id but cannot hold an FDT `compatible` string
|
||||
(`"brcm,bcm2835-aux-uart"`); the identity field grows before the ARM path can
|
||||
report a real node.
|
||||
|
||||
Two supporting decisions keep the kernel's remaining slice honest:
|
||||
|
||||
- **The AML interpreter is a single build module**
|
||||
(`library/device/acpi/aml/aml.zig`) — one source, no fork. During the ring-3 move
|
||||
it was compiled into both the kernel (which linked it just for the `\_S5`
|
||||
poweroff evaluation) and the acpi service, with the `acpi-parse` test
|
||||
asserting the two produce the same device count. Since soft-off followed
|
||||
discovery out of the kernel, only the acpi service links the module — the
|
||||
kernel runs no AML — and the test now asserts a device-count *floor* for the
|
||||
ring-3 parse instead, there being no kernel count left to equal.
|
||||
- **Bridge apertures come from the firmware memory map, not AML.** Registered
|
||||
PCI functions carry BAR resources, and `device_register` containment demands
|
||||
the bridge own windows that cover them. Those apertures are derived
|
||||
kernel-side from the boot memory map's MMIO holes (regions that are neither
|
||||
RAM nor tables) — mechanical, AML-free, and available at boot regardless of
|
||||
what later moved to user space. The acpi service's authority is likewise
|
||||
exactly one node: the `acpi-tables` node, whose broad io_port grant is the
|
||||
documented trust boundary for the one process allowed to run firmware
|
||||
bytecode.
|
||||
@@ -0,0 +1,211 @@
|
||||
# EFI / The Boot Process
|
||||
|
||||
## What EFI is
|
||||
|
||||
**UEFI** (Unified Extensible Firmware Interface) is the software baked into your
|
||||
machine's flash chip that runs the instant it powers on — the modern successor
|
||||
to the legacy BIOS. Its job is to bring the hardware up to a sane state and then
|
||||
find and launch an operating system. From our point of view it's a small runtime
|
||||
that hands us a working CPU, a memory map, and a screen, and then gets out of the
|
||||
way.
|
||||
|
||||
The key thing to understand: **UEFI is not our OS, it's a stepping stone.** It
|
||||
exists to load *us*. Our `boot/efi.zig` is a UEFI *application* — a normal program
|
||||
that the firmware runs — and its entire purpose is to gather what the kernel needs
|
||||
and then jump into the kernel.
|
||||
|
||||
## How the firmware finds us
|
||||
|
||||
UEFI boots by looking for a FAT-formatted partition called the **EFI System
|
||||
Partition (ESP)** and running a file at a well-known fallback path:
|
||||
|
||||
```
|
||||
EFI/BOOT/BOOTX64.efi <- the "removable media" default for x86-64
|
||||
```
|
||||
|
||||
The boot volume is **FHS-shaped** (see the repository-layout note in
|
||||
[README.md](../README.md)): `build.zig` installs `boot/efi.zig` (built for the `uefi`
|
||||
target) at `EFI/BOOT/BOOTX64.efi` — the one path UEFI firmware fixes — and lays
|
||||
the rest out by FHS path: the kernel at `system/kernel`, init at
|
||||
`system/services/init`, the pre-packed boot capsule at `boot/system.img`
|
||||
([system-image.md](system-image.md)).
|
||||
`zig-out` mirrors that tree, but what a machine actually boots is the
|
||||
self-contained FAT32 image `tools/make-fat-image.py` builds from the same files
|
||||
(`danos-usb.img`). The `run-x86-64` step points QEMU at OVMF (UEFI firmware for
|
||||
virtual machines) and attaches that image (its serial-logging twin, built the
|
||||
same way) as a USB mass-storage device on the xHCI bus — the guest never sees
|
||||
`zig-out`. The firmware finds `BOOTX64.efi` on the image and runs it — that's
|
||||
our `main()`, which then loads the kernel and the system binaries from their
|
||||
FHS paths.
|
||||
|
||||
## Boot services: the firmware's API
|
||||
|
||||
While a UEFI app runs, it has access to **boot services** — a table of function
|
||||
pointers the firmware provides for allocating memory, reading files, locating
|
||||
hardware protocols, and so on. In `boot()` this is the very first thing we grab:
|
||||
|
||||
```zig
|
||||
const bs = uefi.system_table.boot_services orelse return error.NoBootServices;
|
||||
```
|
||||
|
||||
Everything the firmware offers hangs off tables reachable from
|
||||
`uefi.system_table`: `boot_services`, `con_out` (the text console we `log()` to),
|
||||
and the various *protocols* (GOP for graphics, SimpleFileSystem for disk access).
|
||||
|
||||
**The critical rule:** boot services are *temporary*. They stop existing the
|
||||
moment we call `ExitBootServices`. So the loader's structure is dictated by one
|
||||
constraint — **gather everything the kernel could ever need first, then exit.**
|
||||
The comment in `boot()` says exactly this:
|
||||
|
||||
> Everything the kernel needs must be gathered *before* we exit boot services,
|
||||
> since afterwards none of these calls are usable.
|
||||
|
||||
## What our loader actually does
|
||||
|
||||
The four milestones below are the spine of `boot()`. Along the way it also
|
||||
captures the **ACPI RSDP** from the UEFI configuration table (while boot
|
||||
services are still up), loads the system binaries into an in-RAM
|
||||
**initial ramdisk** (`loadSystemTree` — normally a single read of the pre-packed
|
||||
`boot\system.img` capsule, which already *is* the ramdisk wire format; it falls
|
||||
back to opening each manifest-listed path, and walks the `/system` and `/test`
|
||||
trees only as a last resort for hand-assembled sticks — the capsule's format, builder, and
|
||||
fallback chain are documented in [system-image.md](system-image.md). Best-effort either way — a kernel-only
|
||||
volume still boots), and builds the **bootstrap page tables** the kernel starts
|
||||
life on (`buildBootstrapTables`), all before the jump:
|
||||
|
||||
### 1. Query the framebuffer (`queryFramebuffer`)
|
||||
|
||||
We ask the firmware for the **Graphics Output Protocol (GOP)**, which describes
|
||||
the linear framebuffer — its address, resolution, pitch, and pixel format —
|
||||
and, when we can, switch the display to its native resolution first:
|
||||
|
||||
- Locate GOP via its *handle* (not `locateProtocol`), because the same handle
|
||||
also carries the display's **EDID**.
|
||||
- Read the EDID (trying the `EDID_ACTIVE` then `EDID_DISCOVERED` protocol on each
|
||||
GOP handle) and parse the first Detailed Timing Descriptor — by convention the
|
||||
panel's preferred (native) resolution. This is best-effort: firmware installs
|
||||
these protocols inconsistently, and OVMF with QEMU's stdvga doesn't expose them
|
||||
at all.
|
||||
- If we got a native resolution, enumerate the GOP modes with `queryMode` and
|
||||
`setMode` to the one that matches it (must be a linear 32bpp layout we can
|
||||
paint into). If EDID gave us nothing, we **keep the firmware's current default
|
||||
mode** rather than guess — with a valid EDID the firmware normally defaults to
|
||||
the native mode itself, so its choice beats second-guessing it with, say, the
|
||||
largest advertised mode.
|
||||
- Copy the resulting address/resolution/pitch/format into our own `Framebuffer`.
|
||||
|
||||
All of this *must* happen now, because after exit there's no GOP to ask. (See
|
||||
[framebuffer.md](framebuffer.md) for what those fields mean.)
|
||||
|
||||
### 2. Load the kernel (`loadKernel` + `loadElf`)
|
||||
|
||||
- Use the **LoadedImage** protocol to discover which device we booted from, then
|
||||
**SimpleFileSystem** to open that volume.
|
||||
- Open the kernel ELF at its FHS path (`system\kernel`), seek to the end to learn its
|
||||
size, rewind, and read the whole ELF into a firmware-allocated pool buffer. (`read`
|
||||
may return short, so we loop.)
|
||||
- Parse the ELF: validate the `\x7fELF` magic and the `x86_64` machine type, then
|
||||
walk the program headers. For every `PT_LOAD` segment we:
|
||||
- reserve the exact physical pages it asks to be loaded at (`p_paddr`) via
|
||||
`allocatePages`,
|
||||
- `@memcpy` the file-backed bytes to that address,
|
||||
- `@memset` the `.bss` tail (the part where `p_memsz > p_filesz`) to zero.
|
||||
|
||||
The kernel is linked to *run* in the higher half (virtual base
|
||||
`0xFFFFFFFF80000000`; `exe.image_base` in `build.zig` is the *virtual* address
|
||||
`0xFFFFFFFF80100000`) but is *loaded* low: the linker script's `AT()` clauses
|
||||
give every segment a low physical load address (`p_paddr`, with `.text` at
|
||||
`0x100000`, 1 MiB), which is what the loader allocates and copies into. The
|
||||
bootstrap page tables built before the jump map the high link addresses onto
|
||||
those low physical pages. If a segment's `p_paddr` collided with
|
||||
firmware-reserved memory, `allocatePages` would fail and we'd need to move the
|
||||
load addresses.
|
||||
|
||||
`loadElf` returns `e_entry`, the kernel's entry-point address.
|
||||
|
||||
### 3. Exit boot services (`exitBootServices`)
|
||||
|
||||
This is the handoff's trickiest step. To exit, the firmware demands the current
|
||||
**memory map** and its *key* — proof that we've seen the latest state of memory.
|
||||
But allocating the buffer to hold the memory map can itself *change* the map,
|
||||
invalidating the key. So it's a retry loop:
|
||||
|
||||
```
|
||||
get map info -> allocate buffer (+ spare descriptors) -> get map
|
||||
-> try exit with map.key
|
||||
-> if it failed, the map moved: free, retry
|
||||
```
|
||||
|
||||
Once `exitBootServices` succeeds, **the firmware's services are gone for good** —
|
||||
we must never touch `bs`, `con_out`, or any protocol again. The machine is now
|
||||
entirely ours.
|
||||
|
||||
### 4. Jump to the kernel
|
||||
|
||||
```zig
|
||||
handoff(cr3, entry, &boot_information);
|
||||
```
|
||||
|
||||
`handoff` is a single inline-asm block — `cli`, load the bootstrap page tables'
|
||||
`cr3`, place the `boot_information` pointer in RDI, then `callq *entry` — so
|
||||
nothing runs between the CR3 load and the jump. This never returns.
|
||||
|
||||
## The ABI subtlety: RCX vs RDI
|
||||
|
||||
There's a deliberate detail worth calling out. A UEFI binary is compiled with the
|
||||
**Microsoft x64** calling convention (first argument in register **RCX**). Our
|
||||
kernel is freestanding and uses the **SysV AMD64** convention (first argument in
|
||||
**RDI**). If we let each side use its target's default, the loader would place
|
||||
`boot_information` in RCX while the kernel looked for it in RDI — and the kernel
|
||||
would read garbage.
|
||||
|
||||
So the convention is pinned explicitly to SysV via the shared
|
||||
`boot_handoff.kernel_abi` (defined in `system/boot-handoff.zig`). The kernel's
|
||||
`kmainEntry` declares `callconv(kernel_abi)`; the loader honours the same contract
|
||||
by loading RDI by hand in `handoff`'s inline asm rather than trusting its own
|
||||
Microsoft-x64 default. `kernel_abi` lives in the shared `boot-handoff` module
|
||||
because it's a contract both binaries must agree on. See
|
||||
[sysv.md](sysv.md) for what "SysV" means and where else it shows up.
|
||||
|
||||
## The handoff contract
|
||||
|
||||
The loader and kernel are two *separate* binaries built for two different targets,
|
||||
so everything they exchange must have an identically-defined memory layout. That's
|
||||
what `system/boot-handoff.zig` provides — imported by both as the `boot-handoff` module.
|
||||
It is *only* the handoff: the kernel↔user ABI (`system/abi.zig`) and the device types
|
||||
(`library/device/model/device-abi.zig`) are separate contracts the bootloader never sees.
|
||||
|
||||
- `BootInformation` — the top-level struct passed to the kernel: the
|
||||
framebuffer, the memory map, the kernel's own `PT_LOAD` segments
|
||||
(`kernel_segments` + `kernel_segment_count`, so the kernel can re-map itself
|
||||
with correct permissions), the ACPI RSDP address, and the initial-ramdisk
|
||||
base/length.
|
||||
- `Framebuffer`, `PixelFormat`, `MemoryMap`/`MemoryRegion`, `KernelSegment`,
|
||||
`kernel_abi` — the shared field layouts and the calling convention.
|
||||
|
||||
The structs are `extern struct`, giving them a stable, C-compatible layout so the
|
||||
bytes the loader writes are the bytes the kernel reads.
|
||||
|
||||
## The whole flow at a glance
|
||||
|
||||
```
|
||||
power on
|
||||
-> UEFI firmware initialises hardware
|
||||
-> finds EFI/BOOT/BOOTX64.efi on the FHS volume, runs it (our efi.zig main)
|
||||
-> grab boot services (+ the ACPI RSDP from the configuration table)
|
||||
-> queryFramebuffer (via GOP: EDID native res, setMode, describe fb)
|
||||
-> loadKernel (read system/kernel ELF, load PT_LOAD segments low, .text at 0x100000)
|
||||
-> loadSystemTree (read the boot\system.img capsule as the in-RAM initial ramdisk;
|
||||
fallbacks: manifest-listed paths, then a /system + /test tree walk)
|
||||
-> buildBootstrapTables (identity + physmap + higher-half kernel mappings)
|
||||
-> exitBootServices (retry until the memory-map key holds)
|
||||
-> handoff: load bootstrap CR3, jump to e_entry, boot_information pointer in RDI
|
||||
-> kernel _start (architecture/x86_64/isr.s: switch to a kernel-owned stack,
|
||||
call kmainEntry -> paging, heap, device discovery,
|
||||
scheduler, SMP, user space)
|
||||
```
|
||||
|
||||
Bottom line: **UEFI's job is to give us a CPU, memory, and a framebuffer, then
|
||||
disappear.** `boot/efi.zig` is the thin bridge that collects those gifts into a
|
||||
`BootInformation`, tears down the firmware, and jumps into the kernel — after
|
||||
which we're on our own.
|
||||
@@ -0,0 +1,134 @@
|
||||
# The physical frame allocator
|
||||
|
||||
Once the kernel knows what RAM exists ([memory-map.md](memory-map.md)), it needs a
|
||||
way to *hand out* that RAM: give me a free page of physical memory, and later,
|
||||
here's one back. That's the **physical frame allocator** (a "physical memory
|
||||
manager", hence `system/kernel/pmm.zig`). It deals only in fixed 4 KiB **frames** — the
|
||||
natural unit because that's the granularity the CPU's paging hardware maps — and
|
||||
it is the primitive everything above it stands on: page tables, the kernel heap,
|
||||
per-process memory all ultimately ask the frame allocator for pages.
|
||||
|
||||
It's **generic kernel code**: it operates on the neutral `boot_handoff.MemoryRegion`
|
||||
array, so there's no UEFI in it and nothing architecture-specific beyond the 4 KiB
|
||||
page. (Contrast [architecture.md](architecture.md), which is where CPU-specific code lives.)
|
||||
|
||||
## Why a bitmap
|
||||
|
||||
There are a few classic designs; danos starts with the simplest that still
|
||||
supports freeing:
|
||||
|
||||
- **Bitmap** (chosen): one bit per frame, `1 = used`, `0 = free`. Freeing is
|
||||
trivial (clear a bit), it's very compact, and you can later extend it to
|
||||
allocate *contiguous* runs by scanning for consecutive zero bits. Allocation is
|
||||
a linear scan, but that's cheap and easy to reason about.
|
||||
- **Intrusive free-list / stack**: store the "next free frame" pointer inside each
|
||||
free frame; O(1) alloc and free. Elegant, but it can't satisfy contiguous
|
||||
multi-frame requests and can't answer "is *this* frame free?".
|
||||
- **Buddy allocator**: great for contiguous power-of-two blocks, but more
|
||||
machinery than a first allocator needs.
|
||||
|
||||
Compactness matters less than clarity here, but it's a nice property: 128 MiB of
|
||||
RAM is 32768 frames — a **4 KiB bitmap, a single frame**. Even 64 GiB needs only
|
||||
2 MiB of bitmap.
|
||||
|
||||
## How it works
|
||||
|
||||
State lives in `system/kernel/pmm.zig`: the `bitmap` slice, `total_frames`, `used_frames`,
|
||||
and a `next_hint` marking where the next allocation scan should start.
|
||||
|
||||
### init(map) — building it from the memory map
|
||||
|
||||
1. **Size it.** Find the highest address across all RAM regions — every kind
|
||||
*except* `mmio` — so reserved and ACPI spans sit *inside* the bitmap, marked
|
||||
used but trackable (e.g. so the boot buffers can be freed later);
|
||||
`total_frames = highest / page_size`. Only MMIO (device address space —
|
||||
remember the ~12 GiB of it from [memory-map.md](memory-map.md)) sits outside
|
||||
the bitmap and is simply never allocatable.
|
||||
2. **Place it (the bootstrap).** The bitmap needs storage before an allocator
|
||||
exists — a chicken-and-egg. Solution: pick the first `usable` region big enough
|
||||
to hold the bitmap and put it there, addressing it through the **physmap**
|
||||
(`boot_handoff.physicalToVirtual`). The loader's bootstrap page tables already
|
||||
provide the physmap and the kernel's own tables keep it, so the pointer stays
|
||||
valid across the paging switch.
|
||||
3. **Mark, then free.** Set the whole bitmap to `used` (`0xff`), then walk the
|
||||
`usable` regions clearing their bits. Doing it in that direction means every
|
||||
gap, reserved span, and hole is unallocatable *by default* — we only ever hand
|
||||
back memory the firmware explicitly called usable.
|
||||
4. **Take back the essentials.** Re-reserve the frames the bitmap itself occupies
|
||||
(they're inside a usable region we just freed), plus **frame 0**, so an address
|
||||
of `0` can keep meaning "no frame".
|
||||
|
||||
### alloc() → ?u64
|
||||
|
||||
Scan the bitmap from `next_hint` (wrapping once) for the first free bit, mark it
|
||||
used, advance the hint, and return `frame * page_size`. Returns `null` when no
|
||||
frame is free — genuine out-of-memory. The hint avoids rescanning the low,
|
||||
long-since-allocated frames on every call.
|
||||
|
||||
### free(addr)
|
||||
|
||||
Clear the frame's bit and, if it's below `next_hint`, pull the hint back so the
|
||||
reclaimed frame gets reused soon. Bogus or double frees (a frame already marked
|
||||
free, or one out of range) are ignored rather than corrupting the used count.
|
||||
|
||||
## Correctness points worth remembering
|
||||
|
||||
- **Generic walk.** Because `MemoryRegion` is danos's own type, the map is a plain
|
||||
slice — none of the variable descriptor-stride from the raw UEFI map.
|
||||
- **Physmap addressing.** The bitmap (and the region array) is reached through
|
||||
the physmap via `boot_handoff.physicalToVirtual`, which both the loader's
|
||||
bootstrap tables and the kernel's own tables provide — no remapping is needed
|
||||
when danos switches to its own paging. The one invariant: the bitmap must sit
|
||||
under the bootstrap physmap's reach (4 GiB), which holds because the placement
|
||||
scan (step 2 above) runs from the lowest usable region up and takes the first
|
||||
one big enough — on the supported configurations that lands well under 4 GiB.
|
||||
- **Frame 0 is reserved** so `0` stays a safe "none" sentinel — and the bitmap is
|
||||
never placed there. (An early bug did exactly that: a `usable` region at physical
|
||||
address 0 collided with a `0`-means-not-found sentinel and tripped a panic. The
|
||||
fix was an optional plus starting the bitmap at least one page in.)
|
||||
- **Everything non-usable is unallocatable by construction** — the "mark all used,
|
||||
then free usable" order gives that for free, so the kernel image, the loader's
|
||||
buffers, MMIO and firmware memory can never be handed out.
|
||||
|
||||
## Verifying it
|
||||
|
||||
`kmain` brings the allocator up and self-tests it. Booted in QEMU with 128 MiB:
|
||||
|
||||
```
|
||||
/system/kernel: frame allocator online
|
||||
free frames: 30520 (119 MiB) <- matches the map's 119 MiB usable
|
||||
alloc x3 : 0x3000 0x4000 0x5000 <- frame 0 reserved, bitmap at 0x1000, AP trampoline at 0x2000
|
||||
after free : 30520 frames free <- three freed, count restored
|
||||
```
|
||||
|
||||
(The AP-trampoline page is claimed with `allocBelow` right after `init`, before
|
||||
the demo allocations — hence they start at 0x3000.)
|
||||
|
||||
The `free frames` MiB agreeing with the memory map's `usable RAM`, the three
|
||||
distinct consecutive addresses, and the count returning to its start after freeing
|
||||
are the three signals that init, alloc and free are all correct.
|
||||
|
||||
## Boot-services memory comes pre-reclaimed
|
||||
|
||||
The UEFI boot-services memory (~44 MiB) is defunct and free once
|
||||
`ExitBootServices` runs, taking usable RAM from ~76 MiB up to ~119 MiB. The frame
|
||||
allocator does **nothing special** to get it: the loader already classified it as
|
||||
`usable` (see [memory-map.md](memory-map.md)), so it's just part of the `usable`
|
||||
regions `init` frees. Keeping that boot-protocol knowledge on the loader side is
|
||||
deliberate — the kernel has no notion of "reclaimable" or of UEFI at all.
|
||||
|
||||
The one live piece in that memory is the boot stack the kernel starts on; the loader
|
||||
leaves the single region containing it `reserved`, so `init` won't hand it out. A
|
||||
later step will move task 0 onto a kernel-owned stack, freeing that last ~1 MiB
|
||||
region too (and giving user mode the clean stack it wants).
|
||||
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Contiguous allocation** — done: `allocContiguous` scans for a run of clear
|
||||
bits, with an optional physical ceiling for DMA (`dma_alloc` is its user), and
|
||||
`allocBelow` serves the SMP trampoline.
|
||||
- **A kernel stack for task 0** — still open: the boot processor's idle task runs
|
||||
on the boot stack to this day, so that region can't be freed.
|
||||
- **Freeing the `reserved` `loader_data`** (the boot-time map buffers) — still
|
||||
open: the bitmap deliberately tracks those frames so they *can* be freed, but
|
||||
nothing frees them yet.
|
||||
@@ -0,0 +1,108 @@
|
||||
# The Framebuffer
|
||||
|
||||
## What a framebuffer is
|
||||
|
||||
A **framebuffer** is just a big region of memory where each element is one
|
||||
pixel's color. The display hardware continuously scans this memory and turns
|
||||
each value into light on the screen. There's no drawing API involved — you
|
||||
write a 32-bit value to the right address, and a pixel changes color. That's
|
||||
exactly what `Console.pixel` does:
|
||||
|
||||
```zig
|
||||
self.rowPtr(y)[x] = color; // system/kernel/console.zig
|
||||
```
|
||||
|
||||
Our `Framebuffer` struct (`system/boot-handoff.zig`) is the four facts you need to
|
||||
address it:
|
||||
|
||||
| Field | Meaning |
|
||||
|----------|---------|
|
||||
| `base` | the memory address where pixel data starts |
|
||||
| `width` | visible pixels per row (e.g. 1920) |
|
||||
| `height` | visible rows (e.g. 1080) |
|
||||
| `pitch` | **bytes** from the start of one row to the start of the next |
|
||||
|
||||
The bootloader (UEFI GOP, in our case) sets all this up and hands it over. The
|
||||
kernel just writes into it: no firmware, no driver — just pixels.
|
||||
|
||||
## The mental model: it's 1D memory pretending to be 2D
|
||||
|
||||
The screen is a grid, but memory is a flat line of bytes. So the pixels are
|
||||
stored row after row, laid end to end:
|
||||
|
||||
```
|
||||
row 0: [px0][px1][px2]...[width-1] <padding?>
|
||||
row 1: [px0][px1][px2]...[width-1] <padding?>
|
||||
row 2: ...
|
||||
```
|
||||
|
||||
To find pixel `(x, y)` you compute:
|
||||
|
||||
```
|
||||
address = base + y * (bytes per row) + x * (bytes per pixel)
|
||||
```
|
||||
|
||||
## So what is pitch?
|
||||
|
||||
**Pitch is "bytes per row"** — sometimes called *stride*. The obvious guess
|
||||
would be `pitch = width * 4` (4 bytes = 32 bits per pixel). And often it is.
|
||||
**But not always** — and that's the whole reason the field exists.
|
||||
|
||||
Hardware frequently wants each row to start at a nicely aligned address (a
|
||||
multiple of 32, 64, or a page). If `width` doesn't land on that boundary, the
|
||||
firmware pads the end of every row with a few extra unused bytes. That padding
|
||||
is invisible — it's never shown — but it's physically there in memory between
|
||||
the last pixel of one row and the first pixel of the next.
|
||||
|
||||
Example: a 1366-pixel-wide display at 32bpp:
|
||||
|
||||
- `width * 4` = 1366 × 4 = **5464 bytes** of actual pixels
|
||||
- but `pitch` might be **5504 bytes** (padded up to a multiple of 64)
|
||||
- those extra 40 bytes per row are dead space
|
||||
|
||||
This is exactly why `rowPtr` uses `pitch`, not `width`, to step between rows:
|
||||
|
||||
```zig
|
||||
inline fn rowPtr(self: *Console, y: u32) [*]volatile u32 {
|
||||
const base: [*]volatile u8 = @ptrFromInt(self.fb.base);
|
||||
return @ptrCast(@alignCast(base + y * self.fb.pitch)); // <- pitch, not width*4
|
||||
}
|
||||
```
|
||||
|
||||
Note the deliberate detail: `base` is cast to a **byte** pointer
|
||||
(`[*]volatile u8`) *before* adding `y * pitch`, because pitch is measured in
|
||||
bytes. Then it's cast to a `u32` pointer so that `[x]` indexes whole pixels. If
|
||||
you'd done the arithmetic on a `u32` pointer, `+ pitch` would step `pitch`
|
||||
*pixels* (4× too far).
|
||||
|
||||
### Why you must use pitch, not `width * 4`
|
||||
|
||||
If you assumed rows were `width * 4` apart on a display where
|
||||
`pitch > width * 4`, every row would start a little too early. The error
|
||||
accumulates: row 0 is fine, row 1 is off by (pitch − width×4) bytes, row 2 by
|
||||
twice that, and so on. The image ends up **skewed diagonally** — a slanted,
|
||||
sheared picture — because each row creeps sideways relative to where the
|
||||
hardware actually reads it.
|
||||
|
||||
Using `pitch` is what keeps each row landing exactly where the scanout expects
|
||||
it.
|
||||
|
||||
## Two subtleties worth noting
|
||||
|
||||
1. **`width` vs `pitch` in the loops.** In `fillRow` we iterate `x` up to
|
||||
`self.fb.width` — the *visible* count — but jump between rows with
|
||||
`pitch`. That's the correct pairing: touch only real pixels, but skip the
|
||||
full stride (including padding) to reach the next row. We never write into
|
||||
the padding, which is right. (A `copyRow` used to sit alongside it; it's
|
||||
gone — `scroll` was since rewritten as a writes-only screen clear, because
|
||||
reading VRAM back is uncached-slow on real hardware.)
|
||||
|
||||
2. **`volatile`.** The pointer is `volatile` because this memory is special —
|
||||
it's watched by the display hardware. `volatile` tells the compiler *"don't
|
||||
optimize these writes away or reorder/coalesce them"*; every store must
|
||||
actually hit memory, because something outside the CPU's knowledge (the
|
||||
scanout engine) is reading it.
|
||||
|
||||
Bottom line: **width is how wide the picture is; pitch is how wide the memory
|
||||
rows are.** They're usually equal (×4) but not guaranteed to be, so always
|
||||
advance rows by pitch.
|
||||
@@ -0,0 +1,71 @@
|
||||
# GOP
|
||||
|
||||
Graphics Output Protocol (GOP) is a UEFI driver interface that replaces legacy VGA BIOS functions to provide graphics console output in the pre-OS phase.
|
||||
|
||||
It allows the firmware to display boot screens and setup menus by providing direct access to the hardware frame buffer, enabling multiple GPUs to function equally without proprietary INT15 handshaking.
|
||||
GOP has no standardized mode numbers. Instead, the firmware (really the GPU's GOP driver) builds a list of modes at boot, and each is just an opaque index 0 .. MaxMode-1. Mode 0 might be 1920×1080 on your laptop and 800×600 on someone else's — the index carries no fixed meaning. To learn what a mode actually is, you have to ask:
|
||||
|
||||
- gop.mode.max_mode — how many modes exist
|
||||
- gop.query_mode(n) — returns the info (resolution, pixel format, pixels_per_scan_line) for mode n
|
||||
- gop.set_mode(n) — switch to it
|
||||
- gop.mode.info — the info for the currently active mode
|
||||
|
||||
queryFramebuffer used to read gop.mode.info directly and never call set_mode — taking whatever mode the firmware selected as its default. It now works out the monitor's native resolution from EDID and, if a matching mode exists, calls set_mode to switch to it. When EDID is unavailable it keeps the firmware's default mode rather than guessing — with a valid EDID present the firmware normally defaults to the native mode itself (this is exactly how the QEMU run boots at 1280x720: OVMF's stdvga driver reads the emulated EDID and defaults to the preferred mode, even though it never exposes the EDID protocols to us). See the "picks the native resolution" note at the bottom.
|
||||
|
||||
Does EFI detect native resolution?
|
||||
|
||||
Sometimes, but it's not guaranteed by the spec. Here's the actual chain:
|
||||
|
||||
1. The GOP driver reads the monitor's EDID over the DDC/I²C wire — a data blob the display publishes describing its supported resolutions, including its preferred (native) timing.
|
||||
2. The firmware then picks a default GOP mode. Many modern firmwares (and OVMF in a VM, driven by the emulated display) do default to the native/preferred resolution. But plenty of firmware defaults to a safe fallback like 1024×768 or 800×600 regardless of what the panel can do.
|
||||
|
||||
So gop.mode.info giving you native res is a common outcome, not a promise. Two more caveats:
|
||||
|
||||
- The mode list itself may not even contain the true native resolution — a GOP driver can expose only a handful of modes.
|
||||
- On a headless VM (your QEMU/OVMF setup), there's no real EDID; the "native" resolution is whatever the emulated GPU advertises. OVMF's default is typically 800×600 or 1024×768 unless you configure it (e.g. QEMU's -device virtio-vga with a set resolution, or the OVMF Platform config).
|
||||
|
||||
One note worth flagging: this picks the native resolution but doesn't force a particular pixel format — a native mode is only chosen if it's a paintable linear 32bpp layout (RGBX/BGRX); a bit_mask/blt_only native mode is skipped and we keep the firmware's default instead. That's the right trade-off for now since your console assumes linear 32bpp.
|
||||
|
||||
## Pixel formats
|
||||
|
||||
Every mode also carries a pixel format, and queryFramebuffer only accepts two of the four GOP formats. The full enum:
|
||||
|
||||
- red_green_blue_reserved_8_bit_per_color (RGBX) — accepted
|
||||
- blue_green_red_reserved_8_bit_per_color (BGRX) — accepted
|
||||
- bit_mask — rejected
|
||||
- blt_only — rejected
|
||||
|
||||
RGBX/BGRX tell you each pixel is a 32-bit value with a fixed byte order. That's what the console needs: a known layout it can write directly with rowPtr(y)[x] = color. The other two break one of those assumptions.
|
||||
|
||||
### bit_mask (spec: PixelBitMask)
|
||||
|
||||
The framebuffer is still linear memory you can write to directly — but the bits aren't in a standard RGBX/BGRX arrangement. Instead the firmware hands you a PixelBitmask struct describing where each channel lives:
|
||||
|
||||
```zig
|
||||
pub const PixelBitmask = extern struct {
|
||||
red_mask: u32,
|
||||
green_mask: u32,
|
||||
blue_mask: u32,
|
||||
reserved_mask: u32,
|
||||
};
|
||||
```
|
||||
|
||||
Each mask marks which bits of the pixel word belong to that channel. This is how the format expresses non-standard layouts — for example 16-bit RGB565 (red = 0xF800, green = 0x07E0, blue = 0x001F: 5/6/5 bits, only 16 bits per pixel), or an odd 32-bit order. To draw a color you'd have to read the masks, work out each channel's bit position and width, shift/scale your 8-bit R/G/B into place, and OR them together — per pixel. It's fully drawable, just not with a hardcoded 32-bit write, so we reject it rather than carry that machinery.
|
||||
|
||||
(RGBX/BGRX are really just two hardcoded special cases of a bitmask. The spec names them separately precisely so simple loaders can skip mask-decoding in the common case.) In practice bit_mask is rare on modern PC firmware — you'll almost always get RGBX or BGRX — so rejecting it costs basically nothing.
|
||||
|
||||
### blt_only (spec: PixelBltOnly)
|
||||
|
||||
This one is more fundamental: there is no linear framebuffer you can address at all. gop.mode.frame_buffer_base is meaningless — you have no pointer to pixel memory. The only way to put pixels on screen is through GOP's Blt ("block transfer") service, the third function pointer in the protocol: you build pixels in your own buffer and ask the firmware to copy ("blit") a rectangle onto the display. The firmware owns the actual scanout memory, wherever it lives (across a bus, behind a GPU command interface, in a layout the CPU can't map directly).
|
||||
|
||||
The catch that matters for us: Blt is a boot service. It stops working the instant you call ExitBootServices — which is exactly when the kernel runs. So a blt_only display gives a post-exit kernel no way to draw pixels at all, and there's genuinely nothing our framebuffer console could do with it. Rejecting it is the only correct response.
|
||||
|
||||
### Summary
|
||||
|
||||
| Format | Linear memory? | Layout | Console |
|
||||
| --- | --- | --- | --- |
|
||||
| rgbx / bgrx | yes | fixed 32bpp byte order | works — direct writes |
|
||||
| bit_mask | yes | arbitrary, described by masks | rejected — would need per-pixel mask decoding |
|
||||
| blt_only | no | no CPU-visible framebuffer; Blt service only | rejected — and Blt is gone after ExitBootServices anyway |
|
||||
|
||||
So the two-case accept list is the right line to draw: RGBX/BGRX are the only formats that give a post-ExitBootServices kernel a flat block of pixel memory it can write to without help from firmware that no longer exists.
|
||||
@@ -0,0 +1,144 @@
|
||||
# Halting
|
||||
|
||||
## Why a kernel needs to halt
|
||||
|
||||
An ordinary program ends by *returning* — `main` finishes, the C runtime calls
|
||||
`exit`, and the OS reclaims the process. A kernel has none of that. There is no
|
||||
OS underneath it, no runtime to return to, and no caller waiting. `_start` is the
|
||||
end of the line. So when the kernel has nothing left to do — whether it finished
|
||||
its work or hit a fatal error — it can't "quit". It has to explicitly park the
|
||||
CPU forever, because if execution ever ran off the end it would just keep fetching
|
||||
whatever bytes follow in memory and execute garbage.
|
||||
|
||||
That's what halting is: deliberately stopping the processor so it does nothing,
|
||||
safely, until the machine is reset or powered off.
|
||||
|
||||
## The core of it: `hlt`
|
||||
|
||||
Everything comes down to one x86 instruction. It's CPU-specific, so it lives in
|
||||
the arch module, `system/kernel/architecture/x86_64/cpu.zig` (see [architecture.md](architecture.md)), and the
|
||||
generic kernel calls it as `architecture.halt()`:
|
||||
|
||||
```zig
|
||||
/// Park the core forever. `hlt` drops it into a low-power idle until the next
|
||||
/// interrupt; the loop re-halts on every wake so the stop is permanent.
|
||||
pub fn halt() noreturn {
|
||||
while (true) asm volatile ("hlt");
|
||||
}
|
||||
```
|
||||
|
||||
**`hlt`** ("halt") tells the CPU core to stop executing instructions and drop into
|
||||
a low-power idle state. It's not a busy-wait — the core genuinely stops, drawing
|
||||
almost no power and generating almost no heat, until something wakes it.
|
||||
|
||||
This is much better than the naïve alternative, a spin loop:
|
||||
|
||||
```zig
|
||||
while (true) {} // "busy-wait" — DON'T do this to idle
|
||||
```
|
||||
|
||||
A bare `while (true) {}` keeps the core running flat out, executing the jump
|
||||
back to the top of the loop billions of times a second — 100% CPU, hot, and (on a
|
||||
laptop) draining the battery, all to accomplish nothing. `hlt` achieves the same
|
||||
"do nothing" outcome while letting the core sleep.
|
||||
|
||||
## Why the loop around it?
|
||||
|
||||
Here's the subtlety the comment points at: **`hlt` is not permanent.** It halts
|
||||
the core only until the *next interrupt* arrives. An interrupt is a signal — from
|
||||
a timer, a keypress, a device — that wakes the CPU so it can respond. When one
|
||||
fires, the core comes out of `hlt` and executes the next instruction.
|
||||
|
||||
If we wrote just a single `hlt`, the very first stray interrupt would wake the
|
||||
core and execution would continue past it — falling off the end of the function
|
||||
into whatever comes next in memory. Wrapping it in `while (true)` closes that
|
||||
door: every time an interrupt wakes the core, the loop immediately runs `hlt`
|
||||
again and it goes back to sleep. The net effect is a permanent halt that still
|
||||
sleeps between the interrupts it can't prevent.
|
||||
|
||||
(danos handles plenty of interrupts through its interrupt descriptor table —
|
||||
the timer tick waking a halted idle core is exactly how scheduling works — and
|
||||
non-maskable and system-management interrupts can wake a halted core regardless
|
||||
of what we handle. The loop makes the halt robust to all of them.)
|
||||
|
||||
## The `asm volatile` part
|
||||
|
||||
`hlt` has no equivalent in plain Zig, so we drop to inline assembly:
|
||||
|
||||
- **`asm`** emits the raw instruction directly into the function.
|
||||
- **`volatile`** tells the compiler *"this has side effects you can't see — do
|
||||
not optimize it away or reorder it."* Without it, an optimizing compiler might
|
||||
reason that the assembly produces no value anyone uses and delete it, or hoist
|
||||
it somewhere wrong. `volatile` pins it exactly where we wrote it.
|
||||
|
||||
## `noreturn`: telling the compiler it's the end
|
||||
|
||||
`halt()` is typed `noreturn` — a real Zig type meaning "this function never gives
|
||||
control back to its caller." That isn't decoration; it changes how the compiler
|
||||
treats the call:
|
||||
|
||||
- Code *after* a `noreturn` call is unreachable, so the compiler needn't emit a
|
||||
return sequence, and won't warn about "missing return value" in the callers.
|
||||
- It lets `kmain` and the exported entry shim `kmainEntry` themselves be
|
||||
`noreturn`, which is the honest signature for a kernel entry point — the
|
||||
bootloader jumps in and nothing ever jumps back out.
|
||||
|
||||
You can see the chain in the code: `_start` (an assembly stub in
|
||||
`system/kernel/architecture/x86_64/isr.s`) installs a kernel-owned stack and calls
|
||||
`kmainEntry` in `system/kernel/kernel.zig`, which is `noreturn`; it calls `kmain`,
|
||||
also `noreturn`, which ends by calling `architecture.halt()`, again `noreturn`.
|
||||
The "never returns" property is threaded all the way down.
|
||||
|
||||
## Where danos halts
|
||||
|
||||
There are three halt sites, and they're all the same idea:
|
||||
|
||||
1. **The BSP's idle loop** — `kmain` no longer runs out of work: it spawns
|
||||
`/system/services/init` as PID 1, drops itself to priority 0, and ends as the
|
||||
bootstrap core's idle task — still by calling `architecture.halt()`:
|
||||
|
||||
```zig
|
||||
scheduler.setPriority(0);
|
||||
status("\n/system/kernel: kernel idle; user space is running.\n");
|
||||
architecture.halt();
|
||||
```
|
||||
|
||||
The timer keeps preempting the idle context into init and whatever else is
|
||||
ready; between those interrupts, the halt loop is exactly the low-power park
|
||||
described above.
|
||||
|
||||
2. **Kernel panic** — the freestanding panic handler has no OS to report to, so
|
||||
it prints the message to the diagnostic log and the on-screen console —
|
||||
forcing the console back on even if a display service had it suppressed — and
|
||||
halts via the same `architecture.halt()`. A panic is unrecoverable here, so
|
||||
stopping the machine — rather than limping on with corrupted state — is the
|
||||
safe response.
|
||||
|
||||
3. **Bootloader failure** — in `boot/efi.zig`, if `boot()` fails *before* handing
|
||||
off to the kernel, `main` logs the error and parks the machine with the same
|
||||
loop so the message stays on screen:
|
||||
|
||||
```zig
|
||||
boot() catch |err| {
|
||||
log("\r\nEFI: boot failed: ");
|
||||
logBytes(@errorName(err));
|
||||
log("\r\n");
|
||||
while (true) asm volatile ("hlt");
|
||||
};
|
||||
```
|
||||
|
||||
(Here it's an inline loop rather than `architecture.halt()` because that lives in the
|
||||
kernel's arch module, and the loader is a separate binary from the kernel.)
|
||||
|
||||
## Summary
|
||||
|
||||
- A kernel can't "exit" — it must explicitly stop the CPU or it runs off into
|
||||
garbage.
|
||||
- **`hlt`** parks the core in a low-power idle until the next interrupt — far
|
||||
better than a 100%-CPU spin loop.
|
||||
- **`while (true) hlt`** makes that halt permanent, since any interrupt would
|
||||
otherwise wake the core and let execution continue.
|
||||
- **`asm volatile`** emits the instruction and forbids the compiler from removing
|
||||
it; **`noreturn`** encodes "control never comes back" into the type system.
|
||||
- danos halts in the kernel's idle loop, on a kernel panic, and on a bootloader
|
||||
error — the same "park the core safely" in all three.
|
||||
@@ -0,0 +1,78 @@
|
||||
# The kernel heap
|
||||
|
||||
The [frame allocator](frame-allocator.md) hands out fixed 4 KiB physical frames;
|
||||
the [VMM](paging.md) maps pages into virtual addresses. The **kernel heap** sits on
|
||||
top of both to provide what the rest of the kernel actually wants: `alloc(n)` /
|
||||
`free(p)` for arbitrary byte sizes. It's the first real consumer of `map()`, and
|
||||
the thing that unlocks dynamic data structures — lists, hash maps, driver state,
|
||||
eventually a process table.
|
||||
|
||||
It's generic kernel code (`system/kernel/heap.zig`): the allocator logic is
|
||||
architecture-neutral, using `architecture.mapPage` and the frame allocator underneath.
|
||||
|
||||
## A growable free-list allocator
|
||||
|
||||
The algorithm is a classic **first-fit free list**:
|
||||
|
||||
- The heap owns a virtual region. Free space is tracked as an **address-ordered
|
||||
singly linked list** of free blocks; each block begins with a 16-byte header
|
||||
(`size`, and a `next` link used while free).
|
||||
- **alloc(n)** walks the list for the first block big enough. If the block is much
|
||||
larger it's **split** — the front becomes the allocation, the remainder stays
|
||||
free. If nothing fits, the heap **grows** (below) and the search retries.
|
||||
- **free(p)** finds the block header just before `p` and inserts it back into the
|
||||
list, **coalescing** with the physically adjacent free blocks on either side so
|
||||
the space can be reused as one region rather than fragmenting away.
|
||||
|
||||
Allocations are 16-byte aligned; larger alignments aren't supported yet (the
|
||||
`std.mem.Allocator` `alloc` returns `null` for them).
|
||||
|
||||
## Growing on demand
|
||||
|
||||
The heap lives in the **higher half** of the address space (virtual base
|
||||
`0xFFFF_8000_0000_0000`) — unmapped, well clear of the low half, which belongs
|
||||
to user space (unmapped in the kernel's own tables; per-process user address
|
||||
spaces now map into it). (That base is x86_64-canonical; another architecture
|
||||
would pick its own.)
|
||||
|
||||
When the free list can't satisfy a request, `grow` extends the mapped region: it
|
||||
pulls fresh frames from the [frame allocator](frame-allocator.md) and `map`s each
|
||||
onto the end of the heap, then adds the new span as a free block (coalescing with
|
||||
the current tail). So the heap starts at one page and expands page-by-page as
|
||||
demand requires, up to a cap. This is exactly what the VMM's on-demand `map` was
|
||||
built for.
|
||||
|
||||
## A std.mem.Allocator
|
||||
|
||||
The heap is exposed as a **`std.mem.Allocator`** (`heap.allocator()`), Zig's
|
||||
standard allocator interface. That's a deliberate multiplier: it means the whole of
|
||||
Zig's standard library — `ArrayList`, `AutoHashMap`, `std.fmt.allocPrint`, and the
|
||||
rest — works directly on the kernel heap, no bespoke containers required.
|
||||
|
||||
## Verifying it
|
||||
|
||||
The `heap` test (see [testing.md](../testing.md)) exercises the allocator end to end:
|
||||
|
||||
```
|
||||
[PASS] alloc 4096 bytes
|
||||
[PASS] heap memory is writable and reads back
|
||||
[PASS] freed block is reused <- free list + coalescing works
|
||||
[PASS] many allocations (heap growth) stay valid <- grow() maps fresh frames
|
||||
[PASS] std.ArrayList on the kernel heap <- std containers work on it
|
||||
```
|
||||
|
||||
The "freed block is reused" check (free then re-alloc returns the same address) is
|
||||
the proof that free and the free list actually work, not just alloc; "heap growth"
|
||||
forces allocation past the initial page so `grow`/`map` runs; and the `ArrayList`
|
||||
check is the std-integration payoff.
|
||||
|
||||
## What's next (largely still true)
|
||||
|
||||
- **Thread/interrupt safety** — overtaken by the big kernel lock: SMP arrived
|
||||
with a single kernel lock taken at every kernel entry, which serializes all
|
||||
heap access. The heap still has no lock of its own, and needs none unless the
|
||||
big lock is ever split.
|
||||
- **Larger alignments** than 16 — still unsupported; page-aligned and DMA
|
||||
buffers come straight from the frame allocator instead.
|
||||
- **`resize`/`remap` in place** — still not done; growing an `ArrayList` copies.
|
||||
- **Reclaiming empty tail pages** — still not done; the heap only ever grows.
|
||||
@@ -0,0 +1,151 @@
|
||||
# Interrupts and exceptions
|
||||
|
||||
When something goes wrong on the CPU — a bad pointer, a divide by zero, a
|
||||
malformed page table — the processor raises an **exception**. If nothing is set
|
||||
up to catch it, the fault escalates: the CPU tries to invoke a handler, finds
|
||||
none, faults again trying to handle *that*, and on the third strike triple-faults,
|
||||
which on real hardware and in QEMU means a silent reset. Debugging by spontaneous
|
||||
reboot is miserable.
|
||||
|
||||
This is the machinery that catches those faults and prints what happened instead.
|
||||
It's all x86_64-specific, so it lives behind the [architecture](architecture.md) boundary in
|
||||
`system/kernel/architecture/x86_64/`. Only the 32 CPU-defined exception vectors were wired up at
|
||||
this stage; device interrupts (timer, keyboard, via the APIC) came later, on the
|
||||
same IDT — today it installs 48 gates (vectors 0-47) plus the ring-3 syscall gate
|
||||
at vector 128.
|
||||
|
||||
## First the GDT
|
||||
|
||||
In 64-bit long mode, segmentation is mostly switched off — but the CPU still
|
||||
requires valid **segment descriptors** for code and data, and, crucially, every
|
||||
IDT gate names a code-segment *selector* that must resolve in the current GDT. The
|
||||
firmware left a GDT in place, but we don't control it, so we install our own with
|
||||
known selectors: `0x08` kernel code, `0x10` kernel data.
|
||||
|
||||
`system/kernel/architecture/x86_64/gdt.zig` held three flat descriptors at this point — a required
|
||||
null entry, plus code and data — where the only bits that matter in long mode are
|
||||
the access byte and the code segment's long-mode (`L`) flag. (The table has since
|
||||
grown to seven entries: ring-3 user data and user code descriptors arrived with
|
||||
user mode, and the TSS descriptor — below — spans two slots.) Loading it (`gdt_flush` in
|
||||
`isr.s`) does two things: `lgdt`, then reload the segment registers. The data
|
||||
registers take a plain `mov`, but **CS can't** — so we reload it with a far
|
||||
return, pushing the new selector and a return address and letting `lretq` pop them
|
||||
into CS:RIP.
|
||||
|
||||
## Then the IDT
|
||||
|
||||
The **Interrupt Descriptor Table** maps each of 256 vectors to a handler. Each
|
||||
entry is a 16-byte *gate* holding the handler's address (split across three
|
||||
fields, a quirk of the format), the code selector (`0x08`), and flags: `0x8E`
|
||||
means present, ring 0, 64-bit interrupt gate. `system/kernel/architecture/x86_64/idt.zig` builds the
|
||||
table, points the first 32 vectors at their stubs (since grown to gates 0-47,
|
||||
plus the ring-3 syscall gate at vector 128), and loads it with `lidt`
|
||||
(`idt_flush`).
|
||||
|
||||
## The TSS and the double-fault stack
|
||||
|
||||
There's one more table, the **Task State Segment**. In long mode its main
|
||||
remaining job is the **Interrupt Stack Table (IST)**: an IDT gate can name an IST
|
||||
slot, and when that vector fires the CPU switches to the stack recorded there —
|
||||
*regardless* of what the interrupted stack looked like.
|
||||
|
||||
This matters most for the **double fault** (#DF, vector 8). A #DF means the CPU
|
||||
hit a fault *while trying to deliver another fault* — very often because the
|
||||
current stack pointer is bad, so pushing the exception frame itself faulted. If
|
||||
the #DF handler then tried to push onto that same bad stack, it would fault a
|
||||
third time and **triple-fault** — an instant reset. So the #DF gate is pointed at
|
||||
**IST1**, a small dedicated stack (`system/kernel/architecture/x86_64/tss.zig`) that's always valid.
|
||||
|
||||
Bringing it up: fill in the TSS's IST1 pointer, publish the TSS through a
|
||||
descriptor in the GDT (`gdt.setTssFor`), and load it into the task register with
|
||||
`ltr`. The TSS descriptor is a 16-byte system descriptor spanning two GDT slots,
|
||||
which is why the GDT grew from three entries to five (and later to seven, when
|
||||
user mode slotted the ring-3 user data/user code descriptors between the kernel
|
||||
data entry and the TSS pair).
|
||||
|
||||
## The stubs and the trap frame
|
||||
|
||||
On an exception the CPU pushes a small frame (SS, RSP, RFLAGS, CS, RIP) and, for
|
||||
*some* vectors, an **error code**. That inconsistency is a nuisance, so each stub
|
||||
in `system/kernel/architecture/x86_64/isr.s` normalises it: vectors that don't get a hardware error
|
||||
code push a dummy `0`, then every stub pushes its **vector number** and jumps to a
|
||||
shared tail, `isr_common`. The tail pushes all the general registers and calls the
|
||||
Zig handler with a pointer to the whole thing.
|
||||
|
||||
The result on the stack is a uniform **`CpuState`** — register block, then vector
|
||||
and error code, then the CPU's frame. Its field order in `idt.zig` is exactly the
|
||||
push order in `isr.s`; the two must stay in sync.
|
||||
|
||||
### Why a separate `.s` file
|
||||
|
||||
The stubs and table-loads are real assembly rather than Zig inline asm because
|
||||
they need things inline asm on this toolchain can't express: cross-symbol
|
||||
`jmp`/`call` (a stub jumping to `isr_common`, which calls the exported
|
||||
`interruptDispatch`), and the `lgdt`/`lidt` memory operands (which LLVM rejects
|
||||
inline). `build.zig` adds `isr.s` to the arch module.
|
||||
|
||||
## Reporting a fault
|
||||
|
||||
`isr_common` calls `interruptDispatch`, which forwards to a swappable `on_fault`
|
||||
hook. The generic kernel installs a reporter (`onException` in `kernel.zig`) that
|
||||
prints the exception name and vector, the error code, the faulting RIP
|
||||
and RSP, and — for a page fault (#PF, vector 14) — the faulting address from
|
||||
**CR2**. What happens next depends on where the fault came from:
|
||||
|
||||
- **User mode (CPL 3): kill the process, keep the machine.** The kernel is intact
|
||||
(the CPU trapped onto the task's kernel stack), so the faulting process is
|
||||
killed — address space, IRQ bindings, and IPC handles reclaimed; a client it
|
||||
owed a reply to is failed with `-EPEER` — and the core reschedules. A crashing
|
||||
driver takes itself down, never the OS. This is fault recovery step 2 of
|
||||
[resilience.md](resilience.md). NMI, double fault, and machine check are
|
||||
excluded: they report machine trouble regardless of what was running.
|
||||
- **Kernel mode: halt this core.** The trusted base itself is broken, so there is
|
||||
nothing safe to kill; the fault is still *contained* to the core (an
|
||||
application-processor fault leaves the rest of the system running), and the
|
||||
report makes it **visible** instead of a silent reset.
|
||||
|
||||
The hook is set before `arch.init()` in `kmain`, so a fault during setup is still
|
||||
caught.
|
||||
|
||||
## Verifying it
|
||||
|
||||
A temporary `ud2` (unconditional invalid-opcode instruction) in `kmain` produced,
|
||||
in red:
|
||||
|
||||
```
|
||||
CPU EXCEPTION: invalid opcode (vector 6)
|
||||
error code : 0x0
|
||||
RIP : 0x000000000010dadd <- the ud2, in the kernel image at 0x100000+
|
||||
RSP : 0x0000000007e8aed0
|
||||
```
|
||||
|
||||
Vector 6 with no error code, a RIP inside the loaded kernel, and a sane RSP
|
||||
together confirm the whole path: the GDT is active (we're still executing), the
|
||||
IDT vectored to the right stub, the stub built a correct `CpuState`, and the Zig
|
||||
handler read it and reported instead of triple-faulting.
|
||||
|
||||
Separately, pointing RSP at an unmapped address and faulting forced a **double
|
||||
fault** — reported cleanly (`double fault (vector 8)`) rather than triple-faulting
|
||||
into a reset, which only works because #DF ran on IST1. That's the proof the
|
||||
TSS/IST is wired up: the handler survived a completely broken stack.
|
||||
|
||||
## What's next (since done)
|
||||
|
||||
Both items originally deferred here have landed:
|
||||
|
||||
- **The IO-APIC**: [ioapic.zig](../../system/kernel/architecture/x86_64/ioapic.zig)
|
||||
routes external device lines onto vectors — discovered via ACPI's MADT, every
|
||||
input masked at init, lines unmasked one at a time as user-space drivers bind
|
||||
them (see [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). The keyboard followed
|
||||
exactly as predicted: the PS/2 bus driver (`system/drivers/ps2-bus/`) claims
|
||||
the 8042 controller and binds its IRQ 1 (and the aux mouse's IRQ 12) through
|
||||
this routing. USB HID keyboards arrive over xHCI instead, which interrupts via
|
||||
MSI, and the HPET's GSI routing exercises the same path.
|
||||
- **SSE state**: `isr_common` (and the syscall entry) now `fxsave`/`fxrstor` the
|
||||
full SSE/x87 register file around dispatch. This stopped being optional the
|
||||
moment kernel code touched XMM — a 16-byte struct copy is a `movdqu` — and its
|
||||
absence was the root cause of a long-lived corruption Heisenbug; the comments
|
||||
in `isr.s` tell the story.
|
||||
|
||||
With faults now debuggable, the paging work that follows — where a wrong
|
||||
page-table entry means an instant #PF — is far less painful.
|
||||
@@ -0,0 +1,100 @@
|
||||
# Logging
|
||||
|
||||
Output is a *diagnostic convenience, never a correctness dependency*: the kernel
|
||||
and every service must run correctly with zero output channels. On top of that
|
||||
rule, danos has **per-process logging** — every process's output is attributed
|
||||
by the kernel and lands in its own file on the flash volume, which is what makes
|
||||
a headless real machine (no serial port) debuggable. The display (the
|
||||
framebuffer surface) is a separate concern with one bootstrap exception: the
|
||||
kernel's framebuffer console (`system/kernel/console.zig`) joins the log sinks
|
||||
at boot, so the whole transcript shows on screen until the display service
|
||||
claims the framebuffer and silences it; after that only panic/fatal messages
|
||||
are mirrored to it explicitly (`system/kernel/kernel.zig`).
|
||||
|
||||
## The pipeline
|
||||
|
||||
```
|
||||
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /var/log/<boot-stamp>/<binary-path>.log
|
||||
kernel log.print ─┘ │
|
||||
└▶ serial / 0xE9 sinks (QEMU, -Dserial)
|
||||
```
|
||||
|
||||
1. **Emit.** A program calls `std.log.info("mounted {s}", .{path})` — the
|
||||
runtime's `logFn` (installed for every binary by the root shim,
|
||||
`library/kernel/logging.zig`) formats one line and issues one `debug_write`
|
||||
carrying the level. The payload does NOT contain the process's name.
|
||||
`logging.write` remains as the raw/bring-up path (panics, test
|
||||
fixtures); raw bytes ride the same ring, attributed all the same.
|
||||
|
||||
2. **Stamp.** The kernel wraps every payload LINE in a record stamped with the
|
||||
sender's pid, task name (its binary path, e.g. `/system/services/fat`),
|
||||
level, a per-boot sequence number, and a monotonic timestamp
|
||||
(`system/kernel/log.zig` + `log-ring.zig`). Attribution is structural — a
|
||||
payload cannot forge another sender's tag, and an embedded newline just ends
|
||||
the record, so the forged "prefix" lands inside the forger's own next line.
|
||||
|
||||
3. **Retain.** The 512 KiB ring overwrites oldest-first; sequence gaps make any
|
||||
loss countable. `klog_read` (#32) copies stream bytes from a free-running
|
||||
offset; `klog_status` (#45) returns the cursors plus the wall-clock time of
|
||||
boot. The framing (`abi.KlogRecordHeader`) is 32 bytes + name + payload,
|
||||
8-byte aligned.
|
||||
|
||||
4. **Render.** Registered sinks (serial under `-Dserial`, the 0xE9 debug
|
||||
console, and the framebuffer console until the display service claims the
|
||||
screen) get a live transcript: kernel/raw output verbatim, leveled records
|
||||
as `<binary path>: message` — one composed write per line, under the log's
|
||||
own spinlock (never the big kernel lock; panic paths try-acquire with a
|
||||
bound and fall back to sinks-only). Sinks are best-effort and self-guarding;
|
||||
a serial-less machine just goes quiet.
|
||||
|
||||
5. **Persist.** The **logger service** (`system/services/logger`) drains the
|
||||
ring every 250 ms and demultiplexes records into one file per source under
|
||||
`/var/log/<boot-stamp>/`, e.g.
|
||||
|
||||
```
|
||||
/var/log/2026-07-21T150434Z/kernel.log
|
||||
/var/log/2026-07-21T150434Z/system/services/fat.log
|
||||
/var/log/2026-07-21T150434Z/system/drivers/usb-storage.log
|
||||
```
|
||||
|
||||
The boot stamp is the RTC anchor from `klog_status` (FAT-safe: no colons; a
|
||||
dead RTC yields the 1970 directory rather than no logs). Each line carries
|
||||
the record's monotonic timestamp and level. Storage is best-effort and late:
|
||||
the ring buffers a whole boot many times over, and the first successful
|
||||
`makePath` of the per-boot directory (also the readiness probe) triggers a
|
||||
full backlog write. Files close — which is the fat server's SCSI cache
|
||||
flush — after a ~2 s quiet period, bounding data-at-risk without per-record
|
||||
flush thrash. At shutdown init stops the logger FIRST (it is last in the
|
||||
boot order), so its final drain runs over a live storage chain.
|
||||
|
||||
## Why a ring in the kernel, not a logging server
|
||||
|
||||
The storage stack must be able to log. If the fat server wrote its own log file
|
||||
through the VFS it would rendezvous-deadlock on itself; if processes sent
|
||||
records to a logging server over IPC, early boot would need a buffer that is —
|
||||
a ring, one hop later. The kernel ring is that buffer, placed where every
|
||||
process (and the kernel itself) can reach it with one syscall, before any
|
||||
service exists. The logger service is a *reader*, not a hop.
|
||||
|
||||
Two disciplines keep it honest:
|
||||
|
||||
- the logger announces itself **once** — a periodic status line would feed the
|
||||
very stream it drains;
|
||||
- lost records surface as an explicit `-- N records lost --` line, computed
|
||||
from sequence gaps, never silently.
|
||||
|
||||
## Last-resort channels
|
||||
|
||||
Unchanged, and independent of the sink list so they survive a total output
|
||||
failure: `checkpoint` (a one-byte POST code on port 0x80) and `recordPanic`
|
||||
(a fixed breadcrumb record, `log.panic_record`, findable in a RAM dump; magic
|
||||
written last so a reader only trusts a complete record).
|
||||
|
||||
## Accepted gaps
|
||||
|
||||
- A write-spamming process can evict other processes' unread records from the
|
||||
ring (a per-process quota is future work); the loss is at least visible via
|
||||
sequence gaps in every affected file.
|
||||
- `/var/log` files have no privacy until the VFS grows permissions.
|
||||
- Records emitted after the logger's final shutdown drain reach serial and the
|
||||
ring but not the files.
|
||||
@@ -0,0 +1,182 @@
|
||||
# The memory map
|
||||
|
||||
Before a kernel can manage memory, it has to *know what memory exists*: which
|
||||
physical address ranges are real RAM it may use, and which are firmware, hardware
|
||||
registers, or already occupied. That inventory is the **memory map**, and the
|
||||
firmware is the only thing that knows it. This page covers how danos gets that map
|
||||
from the firmware and hands it to the kernel — deliberately without dragging UEFI
|
||||
into the kernel.
|
||||
|
||||
## Why not just pass UEFI's map through?
|
||||
|
||||
UEFI hands the loader a perfectly good memory map. The tempting shortcut is to
|
||||
forward it to the kernel as-is. We don't, for two reasons:
|
||||
|
||||
1. **It would tie the kernel to UEFI.** The kernel would compare against UEFI's
|
||||
memory-type numbers and walk the array using UEFI's variable descriptor stride.
|
||||
That's UEFI vocabulary bleeding across the handoff — and danos wants to boot on
|
||||
systems that have no UEFI at all (a Raspberry Pi describes its memory with a
|
||||
*device tree* instead). See [architecture.md](architecture.md) for the same "keep the kernel
|
||||
platform-agnostic" principle applied to CPU code.
|
||||
2. **We already established the better pattern.** The loader doesn't hand the
|
||||
kernel a raw UEFI GOP either — [`queryFramebuffer`](gop.md) converts it to
|
||||
danos's own `Framebuffer`. The memory map follows the same discipline.
|
||||
|
||||
So the boundary is: **each boot path translates its native memory description into
|
||||
danos's own neutral format, and the kernel only ever sees that.**
|
||||
|
||||
## The neutral format
|
||||
|
||||
Defined in `system/boot-handoff.zig`, the shared loader↔kernel contract:
|
||||
|
||||
```zig
|
||||
pub const MemoryKind = enum(u32) {
|
||||
usable, // free RAM the kernel may allocate
|
||||
reserved, // firmware / kernel image / boot stack — real RAM, never hand out
|
||||
acpi_tables, // parse, then reclaim
|
||||
acpi_nvs, // preserve across sleep
|
||||
mmio, // device registers / reserved address space — not RAM at all
|
||||
};
|
||||
|
||||
pub const MemoryRegion = extern struct {
|
||||
base: u64, // physical start
|
||||
pages: u64, // length in page_size (4 KiB) units
|
||||
kind: MemoryKind,
|
||||
_pad: u32 = 0,
|
||||
};
|
||||
|
||||
pub const MemoryMap = extern struct {
|
||||
regions: usize, // pointer to a [len]MemoryRegion
|
||||
len: usize,
|
||||
};
|
||||
```
|
||||
|
||||
`MemoryKind` is danos's *own* vocabulary — not UEFI's ~15 types, just the
|
||||
distinctions the kernel actually acts on. And because danos defines `MemoryRegion`
|
||||
itself, `@sizeOf` is authoritative: the kernel walks a plain `[]MemoryRegion` with
|
||||
no variable-stride subtlety (that stride problem is a UEFI-ism, and it stays in the
|
||||
loader).
|
||||
|
||||
`BootInformation` carries it alongside the framebuffer (trimmed here to the
|
||||
fields this page is about — the full struct has since grown the kernel's
|
||||
PT_LOAD segments, the ACPI RSDP, and the initial-ramdisk span):
|
||||
|
||||
```zig
|
||||
pub const BootInformation = extern struct {
|
||||
framebuffer: Framebuffer,
|
||||
memory_map: MemoryMap,
|
||||
// ...kernel_segments, acpi_rsdp, initial_ramdisk_base/len
|
||||
};
|
||||
```
|
||||
|
||||
## The loader side (UEFI)
|
||||
|
||||
Two functions in `boot/efi.zig`: `exitBootServices` calls `convertMemoryMap`,
|
||||
which runs `classify` on each descriptor:
|
||||
|
||||
- **`classify`** maps each UEFI descriptor to a `MemoryKind`:
|
||||
`conventional_memory` **and** `boot_services_code`/`boot_services_data → usable`;
|
||||
`acpi_reclaim_memory → acpi_tables`; `acpi_memory_nvs → acpi_nvs`;
|
||||
`memory_mapped_io`/`memory_mapped_io_port_space → mmio`; **everything
|
||||
else → reserved** (the safe default). Our own `loader_data` — the kernel image and
|
||||
these buffers — falls into `reserved`.
|
||||
|
||||
Folding boot-services memory into `usable` is deliberate: we've already called
|
||||
ExitBootServices, so it's free RAM now, and doing the classification *here* (in
|
||||
the loader) means the kernel never learns about a UEFI-specific "reclaimable"
|
||||
state — it just sees usable RAM. The one catch is that our stack lives in
|
||||
boot-services memory and the kernel starts out running on it, so
|
||||
`convertMemoryMap` keeps the single region containing the current stack pointer
|
||||
`reserved`. All the boot-protocol knowledge stays on the loader side of the
|
||||
boundary; the kernel's frame allocator has no idea any of this happened.
|
||||
|
||||
One subtlety: **a region that isn't writeback-cacheable (the descriptor's `wb`
|
||||
attribute) is classified `mmio` regardless of type.** UEFI overloads
|
||||
`reserved_memory_type` for both reserved RAM *and* reserved address-space windows
|
||||
(PCIe config space, device BARs); the cache attribute is what actually tells them
|
||||
apart, since only real RAM is writeback-cacheable. Without this, a QEMU q35 guest
|
||||
reports ~12 GiB of "reserved" that is really a PCIe address hole near the 1 TB
|
||||
mark — not memory at all.
|
||||
- **`convertMemoryMap`** walks the UEFI descriptors (striding by
|
||||
`descriptor_size`, *not* `@sizeOf`), classifies each, and writes danos
|
||||
`MemoryRegion`s into an output buffer, coalescing adjacent same-kind regions.
|
||||
|
||||
### The ordering that makes it correct
|
||||
|
||||
This is the fiddly part, dictated by two UEFI rules: you can only allocate memory
|
||||
*before* `ExitBootServices`, and the memory map is only final *at* the moment you
|
||||
exit (its "key" proves you've seen the latest state). So `exitBootServices` does,
|
||||
per attempt:
|
||||
|
||||
1. `getMemoryMapInfo` to size things, then `allocatePool` **two** LoaderData
|
||||
buffers — one for the raw UEFI map, one for the converted regions. Allocating
|
||||
now, before exit, is mandatory.
|
||||
2. `getMemoryMap` then `exitBootServices(key)`. If either fails (allocating can
|
||||
perturb the map and invalidate the key), free both buffers and retry.
|
||||
3. **After** the exit succeeds, convert. Conversion is pure computation on memory
|
||||
we already hold — no boot-services calls — so it's safe once services are gone.
|
||||
|
||||
Both buffers are `LoaderData`, which survives `ExitBootServices`, so the converted
|
||||
array the kernel is pointed at stays valid. (The raw UEFI buffer is just scratch
|
||||
for the conversion.)
|
||||
|
||||
## The kernel side
|
||||
|
||||
The kernel receives a plain array and reads it with zero UEFI knowledge:
|
||||
|
||||
```zig
|
||||
const mm = boot_information.memory_map;
|
||||
const regions = @as(
|
||||
[*]const boot_handoff.MemoryRegion,
|
||||
@ptrFromInt(boot_handoff.physicalToVirtual(mm.regions)),
|
||||
)[0..mm.len];
|
||||
for (regions) |r| {
|
||||
if (r.kind == .usable) usable_pages += r.pages;
|
||||
}
|
||||
```
|
||||
|
||||
(`mm.regions` is a physical address, so it's dereferenced through the physmap —
|
||||
`physicalToVirtual` — since the kernel no longer runs under the loader's
|
||||
identity map.)
|
||||
|
||||
`kmain` summarises the map to prove the handoff works. Booted in QEMU with
|
||||
128 MiB, it reports:
|
||||
|
||||
```
|
||||
/system/kernel: physical memory
|
||||
total RAM : 0.12 GiB (127 MiB) - RAM the firmware reported
|
||||
usable : 121 MiB - free RAM (incl. reclaimed boot-services memory)
|
||||
reserved : 6 MiB - kernel image, boot stack, ACPI, runtime services
|
||||
regions : 28 - entries in the firmware memory map
|
||||
```
|
||||
|
||||
`usable` is ~121 of ~127 MiB because the loader already folded the boot-services
|
||||
memory into it — so the frame allocator gets it all with no special step. The ~6 MiB
|
||||
`reserved` is the kernel image, the boot stack's region, ACPI, and runtime services.
|
||||
`total` counts only writeback-cacheable RAM, so the ~12 GiB PCIe address hole is
|
||||
excluded (it's `mmio`), and the RAM categories summing back to the firmware's total
|
||||
is the sanity check that nothing was dropped.
|
||||
|
||||
## How Raspberry Pi will fit
|
||||
|
||||
No UEFI there, but the boundary is unchanged. The Pi's firmware jumps into the
|
||||
kernel with a **device-tree blob**; the AArch64 entry code will parse its
|
||||
`/memory` and `/reserved-memory` nodes and produce the *same* `MemoryRegion`
|
||||
array. The kernel's memory code — the frame allocator and everything above it —
|
||||
never knows the difference.
|
||||
|
||||
## What's next
|
||||
|
||||
This page is plumbing plus classification only. The map's first consumer, the
|
||||
**physical frame allocator**, is built directly on the `usable` regions here — which
|
||||
already include the reclaimed boot-services memory the loader folded in (see
|
||||
[frame-allocator.md](frame-allocator.md)). Of the two items once listed here, one is done:
|
||||
|
||||
- Freeing the `reserved` `loader_data` (these boot-time buffers) once the kernel
|
||||
is done reading the map — still open: the frame allocator's bitmap tracks those
|
||||
frames so they can be freed, but nothing frees them yet.
|
||||
- Capturing the ACPI RSDP from the UEFI configuration table before exit (the same
|
||||
"grab it before ExitBootServices" pattern) — done: the loader stows it in the
|
||||
boot handoff, and ACPI parsing consumes it from there ([acpi.md](acpi.md)).
|
||||
|
||||
See the roadmap in [efi.md](efi.md) for where this sits in the boot flow.
|
||||
@@ -0,0 +1,145 @@
|
||||
# Paging: the kernel's page tables and VMM
|
||||
|
||||
Every memory access the CPU makes goes through the **page tables**: hardware walks
|
||||
them to translate a virtual address into a physical one, and faults if there's no
|
||||
valid mapping. danos builds its own tables (rather than staying on the firmware's,
|
||||
which live in memory we'd like to reclaim and don't control), switches CR3 onto
|
||||
them, and — crucially — maps with **real permissions**.
|
||||
|
||||
It's x86_64-specific (the 4-level table format is an Intel/AMD thing), so it lives
|
||||
behind the [architecture](architecture.md) boundary in `system/kernel/architecture/x86_64/paging.zig`.
|
||||
|
||||
## The format
|
||||
|
||||
x86_64 uses **4 levels**: PML4 → PDPT → PD → PT, each a 512-entry table, with 9
|
||||
bits of the virtual address indexing each level and the low 12 bits the offset into
|
||||
the final 4 KiB page. Each entry holds a physical address plus flag bits —
|
||||
present, writable, and (bit 63) **no-execute**. danos maps nearly everything with
|
||||
4 KiB pages: precise, and the extra table memory is negligible against available
|
||||
RAM. (The physmap has since become the one exception: 2 MiB-aligned RAM there is
|
||||
mapped with **2 MiB huge pages** — a PS-bit leaf at the PD level — with 4 KiB
|
||||
pages filling the unaligned edges, so the table footprint scales sanely with big
|
||||
RAM. Kernel segments, heap, user space, and on-demand MMIO stay 4 KiB.)
|
||||
|
||||
## Higher half: the address-space layout
|
||||
|
||||
danos is a **higher-half kernel**. The kernel is linked to run at
|
||||
`0xFFFF_FFFF_8000_0000` but loaded low (the linker script's `AT()` gives each
|
||||
segment a physical load address at 1 MiB up; the bootloader maps the high link
|
||||
address to the low load address in its bootstrap tables and jumps in). The entire
|
||||
**low canonical half is reserved for user space**; the kernel lives in the top half
|
||||
alongside a **physmap** — a straight window onto all of physical memory at
|
||||
`physmap_base + phys`. Wherever the kernel needs to touch a physical address (a
|
||||
page-table frame, an ACPI table, a device register), it adds that constant:
|
||||
`boot_handoff.physicalToVirtual(phys)`. The layout constants live in `system/boot-handoff.zig`:
|
||||
|
||||
| region | virtual base | PML4 slot |
|
||||
|--------|--------------|-----------|
|
||||
| user image + stack | `0x0000_7000_0000_0000` | 224 (low half) |
|
||||
| kernel heap | `0xFFFF_8000_0000_0000` | 256 |
|
||||
| physmap (all RAM + MMIO windows) | `0xFFFF_8800_0000_0000` + phys | 272 |
|
||||
| kernel image | `0xFFFF_FFFF_8000_0000` | 511 |
|
||||
|
||||
The bootloader builds temporary **bootstrap tables** (identity + a 4 GiB physmap +
|
||||
the high kernel) so it can switch CR3 and jump to the high entry; the kernel then
|
||||
builds its own precise tables below and abandons them. Because both use the same
|
||||
`physmap_base`, any physmap pointer minted before the switch stays valid after it.
|
||||
|
||||
## What gets mapped, and with what permissions
|
||||
|
||||
The address space is built in four passes (`init`):
|
||||
|
||||
1. **All RAM in the physmap, RW + NX.** Every non-MMIO region from the
|
||||
[memory map](memory-map.md) is mapped at `physicalToVirtual(phys)`, read-write and
|
||||
*non-executable*. There is **no low/identity mapping** — the low half is user
|
||||
space. (Frames the kernel touches while still building these tables are reached
|
||||
through the loader's bootstrap physmap, which covers the low 4 GiB; both the
|
||||
frame allocator and the table builder scan low-address-up, so those frames stay
|
||||
under that limit.)
|
||||
2. **The framebuffer and the Local APIC**, the device memory the kernel touches
|
||||
directly, as physmap windows (RW + NX). Other MMIO is mapped on demand by
|
||||
`mapMmio`, also into the physmap; everything else is left unmapped, so a stray
|
||||
access faults instead of silently succeeding.
|
||||
3. **The kernel's own segments, overlaid with their true ELF permissions**, at
|
||||
their high link addresses mapped to their low physical load addresses. This is
|
||||
the interesting part.
|
||||
4. **Every higher-half PML4 entry pre-created** (an empty PDPT where none exists
|
||||
yet). The kernel half is then a fixed set of top-level slots, so a per-process
|
||||
address space can share it by copying `PML4[256..512)` once — growth beneath
|
||||
those slots (heap, on-demand MMIO) propagates to every address space because
|
||||
they share the PDPTs. `init` asserts no new higher-half PML4 entry appears
|
||||
afterward.
|
||||
|
||||
### W^X from the ELF program headers
|
||||
|
||||
Blanket RW+NX is fine for data but wrong for the kernel's own code, which must be
|
||||
executable — and its code must *not* be writable (W^X: no page is both). We get the
|
||||
right permissions per region straight from the kernel ELF: the **loader already
|
||||
parses the program headers**, so `efi.zig` records each `PT_LOAD` segment's
|
||||
address, size and R/W/X flags into `BootInformation`. Pass 3 re-maps those ranges with
|
||||
flags derived from the ELF flags:
|
||||
|
||||
| segment | ELF flags | mapped as |
|
||||
|---------|-----------|-----------|
|
||||
| `.text` | R + X | present, **not** writable, **not** NX |
|
||||
| `.rodata` | R | present, not writable, NX |
|
||||
| `.data`/`.bss` | R + W | present, writable, NX |
|
||||
|
||||
So code can execute but not be written, and data can be written but not executed.
|
||||
(Intermediate table entries are left writable and executable so the *leaf's* bits
|
||||
govern — a page is writable only if every level is, and non-executable if any level
|
||||
is.) NX itself has to be switched on first via `EFER.NXE`, or the NX bit would be a
|
||||
reserved bit and fault.
|
||||
|
||||
### The null guard
|
||||
|
||||
The whole low half is unmapped except for explicit user mappings, so page 0 (and
|
||||
every near-null address) is unmapped by construction. A null (or near-null) pointer
|
||||
dereference in the kernel takes a page fault instead of quietly reading or writing
|
||||
real memory — turning a whole class of silent bugs into an immediate, located crash.
|
||||
|
||||
## Switching on, and the on-demand API
|
||||
|
||||
Loading the PML4's physical address into **CR3** switches address spaces and
|
||||
flushes the TLB in one step. This works because *while we build* the kernel is
|
||||
still running on the loader's bootstrap tables, whose 4 GiB physmap makes freshly
|
||||
allocated table frames reachable at the same `physmap_base + phys` addresses;
|
||||
afterwards they're covered by pass 1.
|
||||
|
||||
`init` keeps the PML4 and the frame allocator around and exposes `map(virt, phys,
|
||||
writable)` / `unmap(virt)` (with `invlpg` TLB invalidation) — the primitive the
|
||||
kernel heap will build on to map pages on demand.
|
||||
|
||||
## Verifying it
|
||||
|
||||
Four tests (see [testing.md](../testing.md)) pin down the guarantees:
|
||||
|
||||
- **`vmm`** — map a fresh frame at an unused virtual address, write and read it
|
||||
back. Proves `map` works end to end.
|
||||
- **`fault-pf`** — an access far above all mapped RAM faults, with the address in
|
||||
CR2. Proves we're on our own (deliberately sparse) map.
|
||||
- **`fault-nx`** — calling into a data page (NX) faults on the instruction fetch.
|
||||
Proves NX is enforced.
|
||||
- **`fault-null`** — writing to address 0 faults. Proves the null guard.
|
||||
|
||||
> Toolchain notes, both hit while writing the tests: a volatile access to a
|
||||
> compile-time-*constant* address either trips the self-hosted backend's
|
||||
> `mov moffs` gap or (for address 0) Zig's null-pointer safety check — so the
|
||||
> null-guard test launders the address through empty asm and uses an `allowzero`
|
||||
> pointer to force a real hardware access. And `invlpg`, like `lgdt`, needs its
|
||||
> operand staged through a register in inline asm.
|
||||
|
||||
## What's next (mostly done since)
|
||||
|
||||
- **A kernel heap** — done, built on `map` exactly as anticipated
|
||||
([heap.md](heap.md)).
|
||||
- **A higher-half kernel** — done: the kernel is linked at
|
||||
`0xFFFFFFFF80000000` (`linker.ld`), loaded low and running high, and user
|
||||
processes own the low half.
|
||||
- **Per-address-space tables** — done: each user process gets its own root with
|
||||
the kernel half shared, and refcounted shared-memory mappings exist
|
||||
([ipc.md](../device-driver-development-guide/ipc.md)). Copy-on-write remains unbuilt — nothing has needed it yet.
|
||||
- **Uncacheable MMIO** — half done: user-space device and DMA mappings are
|
||||
strong-uncacheable and the framebuffer is write-combining via the PAT, but the
|
||||
kernel's own `mapMmio` path is still writeback — the LAPIC included (see
|
||||
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md)).
|
||||
@@ -0,0 +1,130 @@
|
||||
# The power service: events and shutdown
|
||||
|
||||
A laptop lid closes, a battery drains, someone presses the power button — and
|
||||
several parts of the system might care: a session manager dims the screen, a
|
||||
logger notes it, and ultimately *something* has to turn the machine off. None of
|
||||
them owns the hardware that reported the event, and the reporter should not know
|
||||
who is listening. So system power is a **service**: an event source **publishes**
|
||||
button/lid/battery/AC events, interested processes **subscribe**, and one
|
||||
privileged caller — init — can ask it to power the machine off. It is the same
|
||||
publish/subscribe shape as the [input service](../device-driver-development-guide/input.md), applied to power.
|
||||
|
||||
## Why a service, and why it is named for the domain, not the firmware
|
||||
|
||||
Where the events come from is firmware-specific — on x86 they ride the ACPI SCI
|
||||
([acpi.md](acpi.md)); on a Raspberry Pi they would come from PSCI or a mailbox.
|
||||
What subscribers want is not: *the lid closed* means the same thing regardless of
|
||||
who noticed. So the surface is **domain-named**. There is a `power-protocol`
|
||||
module and a well-known `ServiceId.power = 5`; on x86 the **acpi service**
|
||||
registers it, and on ARM a PSCI/mailbox service will register the *same* id.
|
||||
Subscribers call `ipc.lookup(.power)` and never learn which firmware they
|
||||
are on — the neutrality the whole [discovery](discovery.md) migration exists to
|
||||
preserve, carried one layer up into a running-system surface.
|
||||
|
||||
This is why the protocol is `power`, not "ACPI events": naming a cross-firmware
|
||||
surface after one firmware would leak x86 into code the ARM port must reuse
|
||||
unchanged.
|
||||
|
||||
## The protocol
|
||||
|
||||
The `power-protocol` module ([library/protocol/power/power-protocol.zig](../../library/protocol/power/power-protocol.zig))
|
||||
follows the vfs-protocol pattern — extern-struct messages, a version, reserved
|
||||
fields. Three operations:
|
||||
|
||||
| Direction | Operation | Purpose |
|
||||
|---|---|---|
|
||||
| subscriber → service | `subscribe` | receive published events; the subscriber's endpoint rides as the call's **capability** (the input/device-manager pattern) |
|
||||
| init → service | `shutdown` | orderly shutdown's last step: enter S5 (soft off) |
|
||||
| service → subscriber | `event` | a published `EventMessage`, delivered as a buffered message (never sent *to* the service) |
|
||||
|
||||
Events are published, not polled: like the input service, the service holds
|
||||
subscriber endpoints as capabilities and `ipc_send`s each event as a buffered
|
||||
message, so a slow or dead subscriber can never wedge the source. The event
|
||||
vocabulary is hardware-neutral:
|
||||
|
||||
- `power_button` — the button was pressed (a fixed ACPI event on x86).
|
||||
- `lid`, `ac`, `battery` — the named GPE-driven events.
|
||||
- `notify` — a device notification that maps to none of the above; its `code`
|
||||
(the ACPI `Notify` argument) and the notifying device's `hid` say which device
|
||||
and what happened.
|
||||
|
||||
An `EventMessage` carries the `event` tag plus `code` and an 8-byte `hid`, so a
|
||||
generic `notify` is fully described without a second round trip.
|
||||
|
||||
**`shutdown` is authority, not information.** It is the only operation that
|
||||
*does* something irreversible, so it is gated: the contract is that only init
|
||||
(PID 1) may request it, because init is the process that has already run the stop
|
||||
sequence over everything else. The acpi service implements this as a **soft
|
||||
gate** — it honors `shutdown` only from a process that is a *subscriber*, and
|
||||
init is the one subscriber. That stands in for "only the system supervisor may
|
||||
power off" without hard-coding a pid, so it still holds under tests where PID 1
|
||||
is not init.
|
||||
|
||||
## Orderly shutdown
|
||||
|
||||
Powering off cleanly is where the power service, the [process
|
||||
lifecycle](process-lifecycle.md), and [ACPI events](acpi.md) compose. init
|
||||
already supervises the services it starts; for shutdown it runs **one event loop
|
||||
over one endpoint** that carries three things at once: its children's exit
|
||||
notifications, the lifecycle **signals** it can receive (`terminate`), and the
|
||||
**power events** it subscribes to — plus a re-arming heartbeat timer proving PID
|
||||
1 is alive. (init subscribes with retries, because the power service registers
|
||||
`.power` well after init starts; a missing power service is not fatal — a
|
||||
`terminate` signal drives the same path.)
|
||||
|
||||
On a `power_button` event or a `terminate` signal, init:
|
||||
|
||||
1. logs that it is shutting down,
|
||||
2. runs the standard stop sequence — `process.stop(child, deadline,
|
||||
endpoint)` — over its children **in reverse spawn order**, so the VFS stops
|
||||
last (other services may flush through it), each child getting the
|
||||
*terminate → deadline → kill* escalation from
|
||||
[process-lifecycle.md](process-lifecycle.md), and
|
||||
3. requests `.power` `shutdown`.
|
||||
|
||||
The service then enters **S5** (soft off) by writing `SLP_TYP | SLP_EN` to the
|
||||
PM1 control register(s) from ring 3, with the `SLP_TYP` values taken from its own
|
||||
AML parse of the `_S5` object — the kernel has no S5 path of its own. If the
|
||||
write returns instead of powering the machine off, it logs loudly so a test
|
||||
fails rather than hangs.
|
||||
|
||||
**No new system call was needed for S5.** The broad io_port grant on the
|
||||
`acpi-tables` node ([discovery.md](discovery.md)) already put the PM1 control
|
||||
ports in the acpi service's hands, so writing S5 from ring 3 is something it
|
||||
could physically already do; formalizing it as a protocol operation added a
|
||||
contract, not authority. The kernel keeps only **reboot** (`acpi.reboot` in
|
||||
`system/kernel/acpi.zig` — the FADT reset register plus the legacy fallbacks, which
|
||||
need no AML); it has no poweroff path at all — S5 is not a kernel operation.
|
||||
|
||||
## Verifying it
|
||||
|
||||
Two QEMU scenarios exercise the path, both injecting a real ACPI power-button
|
||||
press via QMP `system_powerdown` (there is no other deterministic power event on
|
||||
this config):
|
||||
|
||||
- `power-button` proves the source: the acpi service's SCI handler logs the
|
||||
press and publishes `power_button` (the ACPI half is in [acpi.md](acpi.md)).
|
||||
- `orderly-shutdown` proves the whole composition: button → init logs shutting
|
||||
down → children stopped → the service enters S5 → QEMU exits. The ordered
|
||||
regex is the proof, and QEMU's self-exit through S5 is the pass.
|
||||
|
||||
## Scope
|
||||
|
||||
Interface-complete but validated on real hardware (the author's laptop) later,
|
||||
because QEMU does not emulate them: battery `_BST`/`_BIF` evaluation beyond the
|
||||
interface stubs, lid and AC events, and the embedded controller's `_Qxx`
|
||||
queries. Deliberately out of scope for now: reboot over the power protocol, S3
|
||||
sleep, per-device D-states (a future lifecycle-vocabulary extension, since
|
||||
"suspend" has the shape of a signal every driver must answer and has no consumer
|
||||
until laptop sleep), and thermal zones.
|
||||
|
||||
## See also
|
||||
|
||||
- [acpi.md](acpi.md) — where the events come from on x86: the SCI, the power
|
||||
button fixed event, and GPE/Notify dispatch in the acpi service.
|
||||
- [discovery.md](discovery.md) — why the surface is domain-named, and the
|
||||
firmware neutrality that makes a PSCI backend drop-in on ARM.
|
||||
- [process-lifecycle.md](process-lifecycle.md) — the stop sequence
|
||||
(`terminate → deadline → kill`) and signals init composes into shutdown.
|
||||
- [device-manager.md](../device-driver-development-guide/device-manager.md) — the supervision model init mirrors for
|
||||
its own children.
|
||||
@@ -0,0 +1,342 @@
|
||||
# Process lifecycle: signals over IPC
|
||||
|
||||
**Status: increments 1–4 built** (2026-07-12): claim release on death, exit
|
||||
reasons, published exit events, and signals + one-shot timers + the service
|
||||
harness are all in — the interface below is as-built. The primitives underneath
|
||||
predate this design ([process-management.md](process-management.md):
|
||||
spawn, the supervision link, kill, child-exit notifications); this document designs
|
||||
the layer above them — the standard vocabulary a danos process speaks about its own
|
||||
life, and the stable `process` interface that carries it. Nothing here is
|
||||
device- or driver-specific: a driver, the VFS, and a user application all stop,
|
||||
reload, and die the same way. The device manager is simply this design's first
|
||||
serious customer ([device-manager.md](../device-driver-development-guide/device-manager.md)).
|
||||
|
||||
**"POSIX" in this document means the concepts, never the letter of the standard.**
|
||||
danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE
|
||||
was a mistake) without inheriting the mechanism, the API, or the names. The naming
|
||||
rule is danos's own and it is strict: plain words that communicate intent
|
||||
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
|
||||
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
|
||||
a concept that already has one. Literal POSIX arrives later and lives elsewhere: the
|
||||
`std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C
|
||||
layer** on the same native surface (see [zig-self-hosting.md](../zig-self-hosting.md)) —
|
||||
musl's syscall surface retargeted at danos system calls and IPC protocols (files onto
|
||||
the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever
|
||||
networking becomes). Ported programs see POSIX; the system underneath never does.
|
||||
|
||||
## Why a standard vocabulary
|
||||
|
||||
A supervisor can only manage processes it has never heard of if "please exit" means
|
||||
the same thing to all of them. That is the one thing POSIX signals got deeply right:
|
||||
`SIGTERM` means the same thing to nginx and to a five-line script, which is why
|
||||
process supervision on Unix (init systems, container runtimes) is possible at all.
|
||||
danos wants that property from day one, because supervision-and-restart is the
|
||||
system's core motivation ([resilience.md](resilience.md)).
|
||||
|
||||
What POSIX got wrong — for a system like this — is the **delivery mechanism**:
|
||||
asynchronous control-flow hijack. A Unix handler runs on a stolen stack at an
|
||||
arbitrary instruction boundary, which is why the async-signal-safe function list
|
||||
exists, why `errno` must be saved, and why the canonical signal bug is a SIGTERM
|
||||
handler innocently calling `printf` mid-`malloc`. That entire bug class comes from
|
||||
the mechanism, not the vocabulary, and none of it is worth importing.
|
||||
|
||||
A microkernel already has the right channel: **a signal is a message.** QNX delivers
|
||||
POSIX signals over its message passing; seL4 has notification objects; Erlang turned
|
||||
"death is a message to whoever linked" into a reliability philosophy. danos has
|
||||
already done it once without naming it: a child's death arrives as a notification
|
||||
badge on the supervisor's endpoint — the microkernel's SIGCHLD, the IRQ-as-IPC
|
||||
pattern reused. Signals are the same pattern reused a third time.
|
||||
|
||||
## The mechanism
|
||||
|
||||
- **`signal_bind(endpoint)`** — a process nominates the endpoint its signals arrive
|
||||
on, exactly as `irq_bind` nominates where a device's interrupts land. The runtime
|
||||
does this at startup for any program that opts in.
|
||||
- **`process_signal(id, signal)`** — posts the signal as an asynchronous
|
||||
notification to the target's bound endpoint: badge = `notify_badge_bit |
|
||||
notify_signal_bit | pending signals`. Non-blocking for the sender, always.
|
||||
Signals address the *process*: `id` may name any member of a threaded process
|
||||
and resolves to its leader — whose endpoint the harness binds — with authority
|
||||
mirroring `process_kill` ([shared-fate-plan.md](shared-fate-plan.md)).
|
||||
- **Pending signals coalesce** in a per-process bitmask while the target has no
|
||||
signal endpoint bound, and the whole mask arrives as one notification at bind —
|
||||
POSIX's own semantics for non-realtime signals (two pending SIGTERMs are one
|
||||
SIGTERM). Once bound, each `process_signal` flushes the mask straight into the
|
||||
endpoint's notification ring, so a busy receiver drains separate posts as
|
||||
separate notifications — harmless, because the badge is a set of bits, never a
|
||||
count. The bitmask *is* the design: signals carry no payload. Anything with a
|
||||
payload is a protocol message.
|
||||
- **Authority**: the supervisor may signal its children — the same link that is
|
||||
already the kill authority. A process may signal itself. Anything broader waits
|
||||
for transferable process handles.
|
||||
- **No binding, no problem**: a process that never calls `signal_bind` is not
|
||||
broken — its signals pend unread and only `process_kill` works on it. Simple
|
||||
programs stay simple; the vocabulary is opt-in, the kill authority is not.
|
||||
|
||||
Because delivery is a message into the process's own event loop, there is no
|
||||
async-signal-safe list in danos: a handler is ordinary code running at a point the
|
||||
process chose. The bug class is gone by construction, not by discipline.
|
||||
|
||||
## The vocabulary: POSIX.1-1990, sorted honestly
|
||||
|
||||
The full 1990 set, and what each becomes. Two intrinsically problematic cases get a
|
||||
defense below the table.
|
||||
|
||||
| POSIX.1-1990 | danos disposition | Notes |
|
||||
|---|---|---|
|
||||
| SIGTERM | signal `terminate` | finish up and exit; the supervisor's polite half |
|
||||
| SIGHUP | signal `reload` | re-read configuration / re-scan |
|
||||
| SIGINT | signal `interrupt` | interactive interrupt; meaningful once a console can send it, in the vocabulary now so numbering is stable |
|
||||
| SIGQUIT | signal `quit` | as SIGINT, without the core-dump baggage |
|
||||
| SIGALRM | signal `alarm` | timer expiry as a message; the Unix SIGALRM+`longjmp` timeout hacks are impossible here. In the vocabulary, unbuilt: no consumer yet, and when one appears it is runtime sugar over the existing timer — zero kernel work |
|
||||
| SIGUSR1, SIGUSR2 | signals `user_1`, `user_2` | service-defined |
|
||||
| SIGCHLD | **already exists** — the exit notification | the badge carries the child id, dodging the classic coalescing bug (Unix code must loop `waitpid`) |
|
||||
| SIGKILL | `process_kill` — kernel mechanism | its definition is "cannot be handled"; it was never really a signal |
|
||||
| SIGABRT | exit reason `aborted` | synchronous self-termination is an exit, not an event; recorded for any nonzero exit code |
|
||||
| SIGSEGV, SIGILL, SIGFPE | exit reasons, **never delivered** | see below |
|
||||
| SIGPIPE | **an error return**, not a signal | see below |
|
||||
| SIGSTOP, SIGTSTP, SIGTTIN, SIGTTOU, SIGCONT | deferred | job control needs terminals, sessions, and process groups; stop/continue is scheduler territory |
|
||||
|
||||
**The fault signals (SIGSEGV, SIGILL, SIGFPE) are intrinsically wrong for messages.**
|
||||
They are *synchronous* — raised at a specific faulting instruction, not "sometime
|
||||
soon". A message cannot be delivered to a process whose next instruction re-faults;
|
||||
it never reaches its event loop to read it. POSIX only makes fault handlers "work"
|
||||
via the async hijack (run the handler *instead of* the instruction), and even there,
|
||||
returning from a SIGSEGV handler without curing the cause is undefined behavior.
|
||||
danos's architecture already has the better answer: fault → the kernel kills the
|
||||
process ([resilience.md](resilience.md) step 2, built) → the supervisor reads the
|
||||
reason → restart. Recovery is restart, not a handler. This is also truer to the 1990
|
||||
standard than handling is: the standard's default action for all three was
|
||||
"terminate the process".
|
||||
|
||||
**SIGPIPE deserves special contempt.** Its default kills a process that writes to a
|
||||
closed pipe — which is why "the whole server died because one client disconnected"
|
||||
is roughly every network daemon's first production bug, and why every mature codebase
|
||||
contains the same fix: ignore SIGPIPE, handle the `EPIPE` error return. danos made
|
||||
the right choice natively already — a reply owed to a dead peer fails with `-EPEER`.
|
||||
Errors from operations are error returns from those operations. The posix layer can
|
||||
synthesize SIGPIPE for ported code that expects it.
|
||||
|
||||
### Statements, not questions
|
||||
|
||||
A signal and a protocol message both travel over IPC — the difference is the
|
||||
**contract**, not the transport. danos IPC has two primitives, both already in
|
||||
daily use: the **asynchronous notification** (a badge — bits that coalesce into a
|
||||
pending mask; the sender never blocks; no payload, *no reply path*; how IRQs and
|
||||
exit events arrive) and the **synchronous call** (a rendezvous — payload both
|
||||
ways, the caller waits for the reply; how VFS requests work). A signal is the
|
||||
first kind: a *statement*. `terminate` wants no reply — the exit notification is
|
||||
its acknowledgement.
|
||||
|
||||
A health probe is the second kind: a *question*, worthless without its answer —
|
||||
and the answer's absence within a deadline is the very thing being measured.
|
||||
Asked as a signal it has no reply channel (a coalescing bit can't carry an answer,
|
||||
and the authority rule forbids a child signalling its supervisor back); asked as a
|
||||
call, the timeout-is-the-diagnosis semantics come free. So there is no `health`
|
||||
signal. Liveness is the common **`ping`**: a reserved request every harness-run
|
||||
service answers automatically on its main endpoint — still free for the service
|
||||
author, still one obvious way — and a supervisor's probe is a `ping` call with a
|
||||
deadline.
|
||||
|
||||
## The two iron rules
|
||||
|
||||
1. **Cleanup is the kernel's job.** A process can die with no warning — fault,
|
||||
kill, power. Correctness must never depend on a `terminate` handler running. On
|
||||
any death the kernel releases the address space, IPC handles, IRQ bindings,
|
||||
owed replies, and **device, I/O-port, and interrupt claims and MSI vectors** —
|
||||
the last of these was once the known gap in
|
||||
[process-management.md](process-management.md), closed by increment 1
|
||||
(`releaseTaskResourcesLocked`, on every death path). A signal handler is
|
||||
for *graceful* work — flushing, deregistering, saving — never for *necessary*
|
||||
work.
|
||||
2. **Kill is not a signal, and exit reasons are load-bearing.** The standard stop
|
||||
sequence is *terminate → deadline → `process_kill`*; the unhandleable kill stays
|
||||
a kernel mechanism. And a supervisor deciding whether to restart must know *how*
|
||||
the child died: clean exit (meant to — don't restart), fault (restart with
|
||||
backoff), killed (the supervisor did it). The exit notification carries only the
|
||||
id; the reason is recorded before the notification posts and read with the
|
||||
supervisor-gated `process_exit_reason` query. Restart policy cannot be
|
||||
written without it.
|
||||
|
||||
## Who learns of a death
|
||||
|
||||
A death has three audiences, and conflating them is how systems end up with either
|
||||
zombie state or privileged snooping:
|
||||
|
||||
1. **The supervisor** — gets the exit notification on the endpoint it gave at spawn
|
||||
(built), then reads the `ExitReason` with the `process_exit_reason` query
|
||||
(increment 2). The supervisor is the only
|
||||
audience that needs the *reason*, because it is the only one deciding whether to
|
||||
restart.
|
||||
2. **The peer owed a reply** — already built: a client that dies mid-request fails
|
||||
the server's reply with `-EPEER`; a server that dies fails its waiting clients
|
||||
the same way. This covers the *synchronous* case only.
|
||||
3. **The subscribers** — the new piece, and it is the input service's
|
||||
publish/subscribe shape ([input.md](../device-driver-development-guide/input.md)) applied to exits. A stateful
|
||||
service accumulates per-client state across many requests: a filesystem server
|
||||
(FAT today) holds a dead client's open file handles, the input service holds
|
||||
its subscriptions, a future network stack holds its sockets. None of these
|
||||
are the client's supervisor, and none learn anything from a failed reply if
|
||||
the client simply never calls again.
|
||||
So the kernel **publishes every exit** to whoever subscribed:
|
||||
`process_subscribe(endpoint)` adds a subscriber, and each death posts a
|
||||
notification to every subscriber (badge = `notify_exit_bit | process id` — the
|
||||
same encoding supervisors already decode, the IRQ-as-IPC pattern once more). The
|
||||
subscriber filters for ids it holds state for and releases what the dead client
|
||||
held. Correlating is free of bookkeeping: an IPC sender's badge already *is* its
|
||||
task id (`ipc.Received`), so the id a service has been keying client
|
||||
state by all along is the id the exit event carries.
|
||||
|
||||
Subscription, not broadcast-to-everyone: only processes that asked receive
|
||||
events, the kernel keeps a bounded subscriber table, and delivery is the same
|
||||
non-blocking coalescing notification as everything else — a dying process never
|
||||
waits on its mourners. Subscribing is ungated, like `process_enumerate`: what is
|
||||
running (and dying) is not a secret between cooperating processes. Subscribers
|
||||
do not receive the exit reason — the filesystem server does not care *why*
|
||||
the client died.
|
||||
|
||||
This is the service-side mirror of iron rule 1: **a service must never depend on
|
||||
its clients cleaning up after themselves.** Handle release on client death is the
|
||||
service's job, triggered by the published exit event — never by a courtesy
|
||||
"closing now" message that a crashed client will never send.
|
||||
|
||||
## The stable interface: `process`
|
||||
|
||||
`process` already owns what a process receives at birth (`Init`, the
|
||||
argv contract). It grows to own the other end of life.
|
||||
|
||||
**The runtime is the stable interface; the numbers are not.** danos applications do
|
||||
not make system calls — they call the runtime library, and the system-call numbers,
|
||||
notification bits, and signal bit positions beneath it are a **private kernel ↔
|
||||
runtime contract** that may change at any time (settled 2026-07-12). This is why
|
||||
the runtime exists. Today kernel and runtime ship from one tree in one image, so
|
||||
"stability" is simply building them together. When driver binaries start shipping
|
||||
as separately-versioned applications — the whole point of the restart design — the
|
||||
binary's embedded runtime version becomes compatibility metadata (the same idea as
|
||||
the protocol version in the device manager's `hello`), and the kernel refuses what
|
||||
it cannot serve. Signals therefore need no reserved numbering scheme: the enum
|
||||
below is vocabulary, not ABI.
|
||||
|
||||
```zig
|
||||
/// The signal vocabulary. The value is the bit position in the pending mask — a
|
||||
/// private kernel/runtime detail, free to change while they ship together.
|
||||
pub const Signal = enum(u5) {
|
||||
terminate = 0, // SIGTERM: finish up and exit
|
||||
reload = 1, // SIGHUP: re-read configuration
|
||||
interrupt = 2, // SIGINT
|
||||
quit = 3, // SIGQUIT
|
||||
alarm = 4, // SIGALRM
|
||||
user_1 = 5, // SIGUSR1
|
||||
user_2 = 6, // SIGUSR2
|
||||
};
|
||||
|
||||
/// A decoded pending mask: the coalesced set of signals a notification delivered.
|
||||
pub const SignalSet = struct {
|
||||
pending: u32,
|
||||
pub fn has(set: SignalSet, signal: Signal) bool { ... }
|
||||
};
|
||||
|
||||
/// Nominate `endpoint` as this process's signal endpoint (signal_bind). The
|
||||
/// runtime's service harness calls this; a bare program may call it directly and
|
||||
/// fold signals into its own replyWait loop.
|
||||
pub fn bindSignals(endpoint: usize) bool { ... }
|
||||
|
||||
/// Decode a received badge into signals, or null if the badge is not a signal
|
||||
/// notification (mirrors ipc.Received.isChildExit).
|
||||
pub fn signalsFrom(badge: u64) ?SignalSet { ... }
|
||||
|
||||
/// Send `signal` to process `id`. Supervisor-gated, like kill; non-blocking.
|
||||
pub fn sendSignal(id: u32, signal: Signal) bool { ... }
|
||||
|
||||
/// The standard stop sequence: terminate, wait up to `deadline_ms` for the exit
|
||||
/// notification on `exit_endpoint` (the endpoint the child was spawned with),
|
||||
/// then process_kill. The one call a supervisor needs.
|
||||
pub fn stop(id: u32, deadline_ms: u64, exit_endpoint: usize) void { ... }
|
||||
|
||||
/// Subscribe `endpoint` to published exit events (process_subscribe). Every
|
||||
/// process death posts an asynchronous notification: badge = notify_exit_bit |
|
||||
/// process id — the same encoding a supervisor's exit notification uses, decoded
|
||||
/// by the same ipc.Received helpers. For stateful services: release what the dead
|
||||
/// client held (file handles, subscriptions, sockets). Ungated, like
|
||||
/// process_enumerate.
|
||||
pub fn subscribeExits(endpoint: usize) bool { ... }
|
||||
|
||||
/// How a process ended — queried after the exit notification (the kernel records
|
||||
/// it first, so the two never race). What restart policy reads. (Built in M17.2.)
|
||||
pub const ExitReason = enum(u8) {
|
||||
exited, // returned from main / clean exit
|
||||
aborted, // deliberate failure exit — any nonzero exit code (SIGABRT's ghost)
|
||||
segmentation_fault, // SIGSEGV's ghost
|
||||
illegal_instruction, // SIGILL's ghost
|
||||
arithmetic_fault, // SIGFPE's ghost
|
||||
protection_fault, // general protection fault
|
||||
fault, // any other CPU exception
|
||||
killed, // process_kill
|
||||
};
|
||||
```
|
||||
|
||||
Two deliberate absences. There is no `mask`/`block` API — a process that is not
|
||||
ready for a signal simply has not waited on its endpoint yet; the pending mask *is*
|
||||
the blocked set. And there is no per-signal handler registration at this layer —
|
||||
dispatch is the process's own `switch` over `SignalSet`, or the service harness's
|
||||
callbacks (`on_terminate`, `on_reload`) for programs that want defaults.
|
||||
|
||||
### The service harness
|
||||
|
||||
`service` owns the `replyWait` loop and folds every event source — signals,
|
||||
child exits, protocol messages — into callbacks, with the vocabulary's defaults:
|
||||
`terminate` returns from the loop (clean exit), the common `ping` is answered automatically,
|
||||
`reload` is ignored unless overridden. One loop, no locking, nothing reentrant. A
|
||||
service author writes domain logic; the lifecycle contract is satisfied by the
|
||||
harness. A process that bypasses the harness and ignores its signals meets the
|
||||
deadline-then-kill escalation — you cannot force a process to implement an
|
||||
interface, but you can make compliance free and non-compliance fatal.
|
||||
|
||||
### The musl layer later
|
||||
|
||||
The POSIX C layer is a **musl port**: musl's arch/syscall layer retargeted so that
|
||||
what musl believes are kernel syscalls become danos runtime calls and IPC — `open`
|
||||
and `read` onto the VFS protocol, `kill`/`sigaction`/`waitpid` onto this document's
|
||||
vocabulary, `exit` onto the runtime's exit path. `sigaction` handlers registered
|
||||
through it are invoked by the runtime's loop when the signal message arrives —
|
||||
synchronous underneath, async-looking to ported code, delivered at wait boundaries
|
||||
the way most Unix programs already experience signals (at syscalls). No stack hijack
|
||||
ever happens, `SA_RESTART` semantics come free because nothing was interrupted, and
|
||||
SIGPIPE can be synthesized from `-EPEER` for the programs that expect it. C programs
|
||||
get POSIX; danos-native programs never pay for it.
|
||||
|
||||
## Increments
|
||||
|
||||
1. **Kernel: release device/port/IRQ claims and MSI vectors on death** — the
|
||||
cleanup half of iron rule 1, and the prerequisite for any restart story. Test:
|
||||
kill a claiming driver, spawn it again, the claim succeeds.
|
||||
2. **Exit reason in the death notification** (`ExitReason` above).
|
||||
3. **Exit events**: `process_subscribe` in the kernel (bounded subscriber table,
|
||||
publishes on every death), `process.subscribeExits`; the userspace VFS
|
||||
router was the first subscriber — releasing a dead client's handles was its
|
||||
proof test — and the FAT server inherited the role when the router moved into
|
||||
the kernel (clients now hold the filesystem server's node ids directly).
|
||||
4. **Signals**: `signal_bind` + `process_signal` + the pending mask in the kernel;
|
||||
`process` grows the interface above; the service harness handles
|
||||
`terminate` and answers the common `ping`; `stop()` for supervisors.
|
||||
|
||||
[device-manager.md](../device-driver-development-guide/device-manager.md) builds directly on all four.
|
||||
|
||||
## Settled questions (2026-07-12)
|
||||
|
||||
- **Signal numbering is not ABI**: the runtime is the stable interface; the numbers
|
||||
beneath it are a private kernel ↔ runtime contract (see "The stable interface").
|
||||
- **Liveness is a `ping` call, not a signal**: signals are statements, questions
|
||||
are synchronous calls (see "Statements, not questions"). A service wanting *deep*
|
||||
health ("can I reach my hardware?") defines its own protocol message on top.
|
||||
- **Process handles: deferred.** Pids + the supervisor gate cover everything
|
||||
planned; transferable handles (Fuchsia-style, delegating signalling without
|
||||
delegating kill) wait for the capability table to grow types beyond endpoints.
|
||||
- **`alarm`: in the vocabulary, unbuilt.** No consumer yet; when one appears it is
|
||||
runtime sugar over the existing timer (arm a timer that posts your own signal) —
|
||||
zero kernel work, so deferring costs nothing.
|
||||
- **Subscription granularity: all exits**, subscriber-side filtering — one
|
||||
subscription per service, a bounded kernel table. Per-id subscriptions only if
|
||||
event volume ever matters (hundreds of processes, not before).
|
||||
- **Client identity across the exit boundary: no convention needed** — an IPC
|
||||
sender's badge already is its task id (see "Who learns of a death").
|
||||
@@ -0,0 +1,130 @@
|
||||
# Process Management
|
||||
|
||||
How danos lists, supervises, and kills processes — the microkernel answer to
|
||||
`ps`, `kill`, and `SIGCHLD`/`wait`.
|
||||
|
||||
## Why system calls, not `/proc`
|
||||
|
||||
Unix systems sit on a spectrum. Classic BSD/macOS list processes through
|
||||
syscalls (`sysctl(KERN_PROC)`) and kill through `kill(2)`; Linux renders the
|
||||
process table as `/proc` for *reading* but still kills through a syscall; Plan 9
|
||||
made the file tree the whole interface (`echo kill > /proc/n/ctl`). Microkernels
|
||||
mostly abandon ambient PIDs: Minix and QNX route everything through a user-space
|
||||
process-manager server, and Fuchsia/seL4 control processes only through handles.
|
||||
|
||||
danos rules out `/proc` **as the primitive**: the path router lives in the
|
||||
kernel (`fs_resolve`), but what is mounted under a path is served by a
|
||||
user-process filesystem server (the way FAT serves `/mnt/usb`) — a `/proc`
|
||||
would be one more such server, which would put a user process in the path of
|
||||
process control. If that server (or anything under it) hangs, nothing could be
|
||||
listed or killed, *including the hung server*. The control plane for processes
|
||||
must not depend on a process. So the primitives are kernel system calls; a
|
||||
read-only `/proc` rendering can be layered on later, and a POSIX-style
|
||||
process-manager server can be built *from* these primitives when one is needed.
|
||||
|
||||
## The three primitives
|
||||
|
||||
### `process_enumerate(buffer, maximum) -> total`
|
||||
|
||||
A snapshot of the task table into a caller buffer of `abi.ProcessDescriptor`
|
||||
(id, supervisor, state, priority, name) — the exact shape of
|
||||
`device_enumerate`, so `ps` is a user program over a snapshot, not a kernel
|
||||
service. The total may exceed what fit; call again with a larger buffer. Kernel
|
||||
tasks are included with an empty name — an honest listing shows the idle tasks
|
||||
too. Ungated and read-only: what is running is not a secret between cooperating
|
||||
bring-up processes.
|
||||
|
||||
### `system_spawn(..., exit_endpoint) -> child id`, and the supervision link
|
||||
|
||||
`system_spawn` records the caller as the child's **supervisor** and returns the
|
||||
child's process id (ids are monotonic, never reused — a stale id can only miss).
|
||||
That link is the kill authority: it answers "who may kill process 7?" without
|
||||
inventing users or permissions, the same way a device *claim* is the capability
|
||||
for `mmio_map`. It composes with the supervision hierarchy the device manager
|
||||
already forms: init supervises the services it starts, the device manager
|
||||
supervises the drivers it matches. (A transferable process *handle* — Fuchsia
|
||||
style — can replace the id once the handle table grows types beyond endpoints.)
|
||||
|
||||
`exit_endpoint` (a handle, or `abi.no_cap`) is the supervisor's death-watch: when
|
||||
the child ends — clean exit, CPU fault, or `process_kill` — the kernel posts an
|
||||
asynchronous notification to that endpoint, exactly like a bound IRQ. The badge
|
||||
carries `abi.notify_badge_bit | abi.notify_exit_bit | child_id`, so one endpoint
|
||||
supervises many children and can even share with IRQ notifications. This is the
|
||||
microkernel's SIGCHLD: no new mechanism, just the IRQ-as-IPC pattern reused, and
|
||||
a supervisor's event loop (`ipc.replyWait`) already knows how to receive it. The
|
||||
child holds a reference to the endpoint from birth, so the notification cannot
|
||||
dangle even if the supervisor dies first.
|
||||
|
||||
### `process_kill(id) -> 0 / -ESRCH / -EPERM`
|
||||
|
||||
Only the supervisor may kill; kernel tasks are not killable processes. The kill
|
||||
is a **whole-process** kill ([shared-fate-plan.md](shared-fate-plan.md)): `id`
|
||||
may name any member of a threaded process — it resolves to the group's leader,
|
||||
authorization is checked against the *leader's* supervisor, and every thread
|
||||
dies. Like a signal, delivery is prompt but asynchronous — 0 means the kill is
|
||||
accepted and irrevocable; the exit notification (badged with the leader, posted
|
||||
once the last member is gone) confirms completion.
|
||||
|
||||
## How a kill lands (the kernel mechanics)
|
||||
|
||||
Everything below runs under the big kernel lock, where task states cannot move.
|
||||
|
||||
- **Target ready or blocked** (not on any core): reaped on the killer's own
|
||||
call. The reap releases what death always releases (IRQ bindings first, then
|
||||
a client the target still owed a reply to is failed with `-EPEER`, IPC handles
|
||||
closed, the exit notification posted last) — plus the unlinking only a
|
||||
*remote* death needs: out of the ready queue, out of an endpoint's sender FIFO
|
||||
(`Task.ipc_wait_endpoint`), out of a receive wait queue (`Task.wait_queue`),
|
||||
and out of any server's owed-reply slot, so nothing ever dequeues a dangling
|
||||
pointer. Destroying the address space is safe because no core can have it
|
||||
loaded: every switch away from a task loads the next task's tables.
|
||||
- **Target running on another core**: it cannot be torn down mid-instruction,
|
||||
so it is condemned (`Task.kill_pending`) and dies at whichever comes first:
|
||||
- its next **system_call entry** — checked before dispatch, so a condemned
|
||||
process cannot spawn, claim, or message anything on its way out;
|
||||
- its core's next **timer tick** — but only when the task is not inside one
|
||||
of its own system calls (`Task.in_system_call`): the tick may have
|
||||
interrupted kernel code mid-operation, where teardown would leak whatever
|
||||
the operation held. User-mode execution is always a safe kill point. The
|
||||
tick-time terminate abandons the interrupt frame exactly like the fault
|
||||
path (the LAPIC is acknowledged before the tick hook runs);
|
||||
- any core's tick finding it **blocked or ready** (it entered a syscall and
|
||||
parked after being condemned) — reaped by the same remote-reap path.
|
||||
|
||||
A pure user-mode spin loop that never makes a system call therefore dies
|
||||
within one tick; nothing a process does can outrun the kill.
|
||||
|
||||
The scheduler stays below the process layer: finishing a kill (IRQ bindings,
|
||||
handles, the notification) is called *up* through two hooks process.zig
|
||||
registers at boot (`terminate_current_hook`, `reap_task_hook`), mirroring how
|
||||
the architecture layer calls up into `tick`.
|
||||
|
||||
## Known gaps (bring-up honesty)
|
||||
|
||||
- ~~Device claims are not released on death~~ Closed (M17.1): every path out of a
|
||||
process releases its device claims alongside its IRQ and MSI bindings
|
||||
(`releaseTaskResourcesLocked`), so a restarted driver can claim its hardware
|
||||
again — the cleanup half of [process-lifecycle.md](process-lifecycle.md)'s iron
|
||||
rule 1. The `claim-release` test proves the kill → release → re-claim cycle.
|
||||
- ~~Kernel stacks of dead tasks are leaked~~ Closed (threading-plan M8): a task
|
||||
exiting on its own core queues on the core's reap list in `.reaping` state, and
|
||||
the next switch away (or tick) frees its kernel stack; one killed while off-CPU
|
||||
has its stack freed synchronously by the reap itself. Both paths are accounted
|
||||
by `live_stack_bytes`, which returns to baseline when no extra tasks are live.
|
||||
- ~~There is no exit status in the notification~~ Closed (M17.2): the kernel
|
||||
records how every process ends — exited, a fault class, or killed — before it
|
||||
posts the exit notification, and the supervisor reads it with
|
||||
`process_exit_reason` (`process.exitReason`). This is the input to
|
||||
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
|
||||
for the clean case can still ride alongside later.
|
||||
- Enumerate writes through the caller's raw pointer under the bring-up trust
|
||||
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
|
||||
isolation break).
|
||||
|
||||
## Tests
|
||||
|
||||
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
|
||||
notifications), `supervision` (the whole user-side surface via the process-test
|
||||
service: spawn supervised → enumerate → kill blocked and spinning children →
|
||||
notifications → gone), `claim-release` (a killed claim-holder's device is
|
||||
claimable again). See test/qemu_test.py.
|
||||
@@ -0,0 +1,70 @@
|
||||
# The release ISO — flashable boot media
|
||||
|
||||
`zig build release-x86-64` produces **`zig-out/danos-x86-64.iso`**, the file you
|
||||
hand to someone who wants to try danos on a real machine: point
|
||||
[balenaEtcher](https://etcher.balena.io) (or Raspberry Pi Imager, or plain `dd`)
|
||||
at it, flash a USB stick, and boot the stick. The same file also burns to
|
||||
optical media. `zig build check-iso-image` validates it without booting.
|
||||
|
||||
```
|
||||
zig build release-x86-64
|
||||
# Etcher: select danos-x86-64.iso → select the stick → Flash
|
||||
# or: sudo dd if=zig-out/danos-x86-64.iso of=/dev/rdiskN bs=4m (macOS; triple-check N)
|
||||
```
|
||||
|
||||
## Why an ISO when danos-usb.img already boots
|
||||
|
||||
`danos-usb.img` is a raw FAT32 **superfloppy** — a filesystem starting at
|
||||
sector 0, no partition table. UEFI firmware accepts that from a USB stick (it
|
||||
probes whole-disk FAT before giving up), which is why `dd`-ing the .img works
|
||||
and why QEMU and the test harness boot it directly. But it is a
|
||||
developer-shaped artifact: flashing apps expect an ISO, and a superfloppy
|
||||
can't be burned to a CD/DVD or carry a partition table for pickier firmware.
|
||||
|
||||
The ISO wraps that same FAT image — bit-identical, built by the same
|
||||
`tools/make-fat-image.py` — in a container that boots everywhere release media
|
||||
gets consumed. One payload, two images: the .img stays the raw volume the QEMU
|
||||
harness mounts and boots, the .iso is what leaves the building.
|
||||
|
||||
## How a hybrid ISO boots twice
|
||||
|
||||
The trick (the same one Linux distribution ISOs use, usually via `xorriso
|
||||
-isohybrid…`) is that ISO9660 reserves its first 32 KiB as a **system area** it
|
||||
never touches — exactly where an MBR lives on a disk. So one file can carry two
|
||||
tables of contents, both pointing at the same embedded FAT image:
|
||||
|
||||
* **Flashed to USB (Etcher, dd):** firmware sees a disk whose sector 0 is an
|
||||
MBR with one partition of type `0xEF` (EFI System Partition) covering the
|
||||
embedded FAT image. It mounts that ESP and runs `\EFI\BOOT\BOOTX64.efi` —
|
||||
the standard removable-media path ([efi.md](efi.md)).
|
||||
* **Burned to optical media:** firmware reads the ISO9660 volume descriptors
|
||||
at sector 16 and finds an **El Torito** boot record. Its catalog has one
|
||||
entry, platform ID `0xEF` (EFI), whose start LBA is — again — the embedded
|
||||
FAT image. The firmware exposes that image as a virtual disk and runs the
|
||||
same `BOOTX64.efi` off it.
|
||||
|
||||
Neither path involves the legacy BIOS boot-sector machinery: danos is
|
||||
UEFI-only ([system-requirements.md](../system-requirements.md)), so the MBR holds
|
||||
no boot code, just the partition entry, and the El Torito entry is EFI-class,
|
||||
not floppy emulation.
|
||||
|
||||
One El Torito wrinkle: the catalog's sector-count field is 16-bit (units of
|
||||
512 bytes), so it can name at most 32 MiB — less than the 64 MiB FAT image.
|
||||
That is fine in practice: firmware sizes the FAT filesystem from its own BPB,
|
||||
and the boot files sit in the first few MiB of the image (clusters are
|
||||
allocated from the front) either way. The USB path has no such cap.
|
||||
|
||||
## The builder
|
||||
|
||||
`tools/make-iso-image.py` follows the house rule of
|
||||
[make-fat-image.py](../../tools/make-fat-image.py): pure Python 3 standard
|
||||
library, no external tools (no xorriso, mkisofs, or isohybrid), with a
|
||||
`--verify` mode the `check-iso-image` step runs — it checks that the MBR
|
||||
partition and the El Torito catalog agree on where the FAT image lives and
|
||||
that a FAT32 boot sector is actually there. Every timestamp field in the ISO
|
||||
is zeroed, so the build is reproducible byte-for-byte.
|
||||
|
||||
The ISO9660 filesystem around the boot machinery is minimal but real: a root
|
||||
directory listing `BOOT.CAT` (the catalog) and `EFI.IMG` (the FAT image), so
|
||||
`file`, mount tools, and archive browsers can open the ISO and see what's in
|
||||
it.
|
||||
@@ -0,0 +1,165 @@
|
||||
# Resilience: fault isolation and live restart
|
||||
|
||||
Steps 1–4 of the ordering below are **built** (M17–M18, 2026-07-13): user-mode
|
||||
isolation; fault → kill the process → keep the core (`onException`; the
|
||||
`fault-recovery` test); the supervisor notification **with exit reasons**
|
||||
([process-lifecycle.md](process-lifecycle.md) — clean exit, fault class, or
|
||||
killed, recorded before the notice posts); and the **restart policy itself**
|
||||
([device-manager.md](../device-driver-development-guide/device-manager.md)): the device manager supervises every
|
||||
driver, restarts crashes with backoff, caps crash loops, and re-claims work
|
||||
because the kernel releases a dead process's claims. The `driver-restart` and
|
||||
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
|
||||
end to end. What remains of this document's ladder is scope, not mechanism:
|
||||
more of the system moved into restartable processes (the discovery migration,
|
||||
[discovery.md](discovery.md), is the next rung). This is the property danos is really chasing:
|
||||
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
|
||||
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
|
||||
the reason the [microkernel](../vision.md) shape was chosen, and it's a *separate* goal
|
||||
from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) —
|
||||
one that's less pervasive to build (see [vision.md](../vision.md)).
|
||||
|
||||
## The idea: "let it crash" + supervision
|
||||
|
||||
The philosophy is older than microkernels and shows up across systems: don't try to
|
||||
make every component perfect — make failures **contained and recoverable**. Isolate
|
||||
each component, watch it, and when it dies, restart it from a known-good state. A
|
||||
small trusted core supervises a fleet of restartable, untrusted parts.
|
||||
|
||||
Prior art worth studying (see Further reading): **MINIX 3's reincarnation server** (a
|
||||
driver crashes, a supervisor restarts it live — the closest thing to your goal),
|
||||
**QNX** (restartable drivers on a message-passing microkernel), **Erlang/OTP
|
||||
supervision trees** ("let it crash", not a kernel but the canonical design), and the
|
||||
historical **Tandem NonStop** (fault-tolerant by process pairs).
|
||||
|
||||
## Why a microkernel makes this possible
|
||||
|
||||
The blast radius of a fault is the address space it happens in. In a monolith, a
|
||||
driver bug can corrupt anything — the kernel *is* the driver. In a microkernel,
|
||||
drivers and services are **isolated user-space processes**, so a fault is trapped by
|
||||
the kernel and confined to that one process. The kernel — the one thing that *can't*
|
||||
be restarted, because it's the trusted base — stays tiny, which is precisely why a
|
||||
small kernel is a *more recoverable* kernel: less code that can take the whole system
|
||||
down. **Keeping the kernel minimal is a resilience strategy, not just an aesthetic.**
|
||||
|
||||
## The building blocks
|
||||
|
||||
1. **Address-space isolation.** A fault in one component can't corrupt another or the
|
||||
kernel. This is the [user-mode milestone](../vision.md) (ring 3, per-process page
|
||||
tables) — the shared prerequisite for *any* of this, and it's needed regardless.
|
||||
2. **Fault detection** — how the system notices a component is dead or sick:
|
||||
- **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps
|
||||
to the kernel, which kills the process and notifies the supervisor. The clean,
|
||||
easy case — and danos already reports CPU faults (see
|
||||
[interrupts.md](interrupts.md)); user mode turns "halt on fault" into "kill the
|
||||
process and tell the supervisor."
|
||||
- **Hang**: a livelocked or infinite-looping component needs a **watchdog /
|
||||
heartbeat** and the ability to **preempt and kill** it. The preemptive scheduler
|
||||
already built ([scheduling.md](scheduling.md)) is what makes a runaway component
|
||||
killable — a nice case of a scheduling mechanism serving resilience without any
|
||||
real-time *guarantee*.
|
||||
- **Misbehaviour**: IPC timeouts, failed health checks.
|
||||
3. **A supervisor / reincarnation server.** A user-space server holding the *policy*:
|
||||
what components exist, their dependencies, and each one's restart strategy. When a
|
||||
component dies, it decides whether/how to restart it. (MINIX 3 calls this the
|
||||
reincarnation server; Erlang calls it a supervisor.)
|
||||
4. **A resource model that supports clean teardown.** When a component dies, its
|
||||
resources — memory, MMIO grants, IPC channels, IRQ routes — must be **reclaimed**,
|
||||
and a restarted replacement must be able to **re-acquire** them. This is where a
|
||||
**capability** model shines (seL4's reference design): a component holds
|
||||
capabilities to its resources; killing it **revokes** them, which frees everything
|
||||
in one clean sweep, and the supervisor hands the replacement fresh caps. A simpler
|
||||
grant/ownership table can work too — capabilities are the principled version.
|
||||
5. **Re-initialisable drivers.** A driver must start from a known state and
|
||||
re-establish its hardware. Some hardware is easy to reset; some holds state that's
|
||||
hard to recover — a real limit on what "just restart it" can fix.
|
||||
|
||||
## The hard part: restarting *correctly*
|
||||
|
||||
Detecting and killing is the easy half. The genuinely tricky questions are about the
|
||||
*rest of the system* when a component dies:
|
||||
|
||||
- **In-flight IPC**: messages sent to the dead component, or replies its clients are
|
||||
blocked waiting for. The channel has to break cleanly and unblock the waiters with
|
||||
an error rather than hang them forever (a design constraint that reaches back into
|
||||
[ipc.md](../device-driver-development-guide/ipc.md) — channels need a "peer died" outcome).
|
||||
- **Clients**: how does a client discover the service it was talking to is gone and
|
||||
has been replaced? Options: capability revocation makes stale handles fail; or a
|
||||
**name server** re-binds clients to the new instance; or clients retry through a
|
||||
stable endpoint.
|
||||
- **State**: the cheapest model is **stateless restart** — the replacement starts
|
||||
fresh and clients re-establish whatever they need. Richer options (checkpointed
|
||||
state, state handed to a standby) are more work and more failure modes. Start
|
||||
stateless.
|
||||
|
||||
These are the constraints most worth *bumping into and researching* — they're where
|
||||
resilience gets genuinely interesting.
|
||||
|
||||
## Kernel mechanism vs user-space policy
|
||||
|
||||
The microkernel split applies to fault management itself:
|
||||
|
||||
- **Kernel (mechanism):** isolation, trapping faults, enforcing capabilities/grants,
|
||||
IPC, creating/destroying address spaces, granting/revoking resources, preempting a
|
||||
runaway task.
|
||||
- **User space (policy):** the supervisor decides *what* to restart, *when*, and
|
||||
*how* — dependency order, retry limits, escalation. None of that belongs in the
|
||||
kernel.
|
||||
|
||||
So the kernel gains a few primitives (kill an address space, reclaim its resources,
|
||||
deliver a "child died" notification); everything smart lives in a user-space server.
|
||||
|
||||
## What's *not* recoverable this way
|
||||
|
||||
Honest boundaries:
|
||||
|
||||
- **The kernel itself.** It's the trusted base; if it faults, this mechanism can't
|
||||
save it. The mitigation is to keep it tiny — the microkernel bet.
|
||||
- **Corrupted hardware state.** Isolation limits the blast radius to one process, but
|
||||
if a driver wedged the device itself, a restart may not un-wedge it.
|
||||
- **Shared-resource corruption** that happened *before* the fault was detected. Clean
|
||||
capability revocation limits this, but it's why fault *detection latency* matters.
|
||||
|
||||
## Suggested ordering
|
||||
|
||||
1. **User mode + address-space isolation** — the shared prerequisite (also on the
|
||||
path for everything else). **Done.**
|
||||
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
|
||||
"confine to the process and report it." **Done** (the kill and reclaim; the
|
||||
supervisor notification waits for step 3's supervisor). A killed server's
|
||||
pending client is unblocked with `-EPEER` rather than hung.
|
||||
3. **A minimal supervisor server** that can (re)start a process.
|
||||
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
|
||||
table.
|
||||
5. **First restartable driver** — the keyboard — as the end-to-end proof: crash it on
|
||||
purpose, watch it come back.
|
||||
|
||||
## Relationship to real-time
|
||||
|
||||
Resilience needs **structural** features (isolation + supervision + a resource
|
||||
model); real-time needs a **pervasive** timing invariant. They're separable, and
|
||||
resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](../vision.md)).
|
||||
Note the overlap, though: **preemptive scheduling** and **priorities** — already
|
||||
built — serve resilience too (you can preempt and kill a misbehaving component, and
|
||||
run the supervisor at high priority). So danos keeps the useful *mechanisms* of the
|
||||
real-time work without owing anyone a timing *guarantee*.
|
||||
|
||||
## Further reading
|
||||
|
||||
- Herder, Bos, Gras, Homburg, Tanenbaum — the **MINIX 3** papers, esp. *"Construction
|
||||
of a Highly Dependable Operating System"* and *"Fault Isolation for Device
|
||||
Drivers"* — the reincarnation server, the closest match to danos's goal.
|
||||
- **QNX** architecture — a shipping microkernel with restartable drivers.
|
||||
- **Erlang/OTP** supervision trees and the *"let it crash"* philosophy — the design
|
||||
pattern, distilled.
|
||||
- **seL4** capability model — the principled basis for clean resource teardown.
|
||||
- **Tandem NonStop** (historical) — fault tolerance via process pairs.
|
||||
|
||||
## Related
|
||||
|
||||
- [vision.md](../vision.md) — the goals this serves (learning by doing; resilience over
|
||||
hard real-time).
|
||||
- [scheduling.md](scheduling.md) — preemption, which makes runaway components killable.
|
||||
- [ipc.md](../device-driver-development-guide/ipc.md) — channels that need a "peer died" outcome for clean restart.
|
||||
- [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and
|
||||
restart" instead of "halt".
|
||||
- [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context.
|
||||
@@ -0,0 +1,149 @@
|
||||
# Scheduling
|
||||
|
||||
The scheduler turns danos from a linear "boot then halt" kernel into a **running
|
||||
multitasking system**. It's **fixed-priority preemptive**: the highest-priority
|
||||
ready task always runs, and tasks at the same priority take turns. That model is
|
||||
chosen for [real-time](../vision.md) — it's predictable (you can reason about which
|
||||
task runs when) and its decisions are O(1), unlike a fair-share scheduler.
|
||||
|
||||
The scheduler proper (`system/kernel/scheduler.zig`) is generic; the context switch and new-task
|
||||
stack setup are architecture-specific (`system/kernel/architecture/x86_64/`, see [architecture](architecture.md)).
|
||||
|
||||
## Tasks
|
||||
|
||||
A **task** is a kernel thread: ring-0 code with its own 16 KiB stack (allocated
|
||||
from the [heap](heap.md)). A task struct holds its saved stack pointer, priority,
|
||||
state, and a ready-queue link. The currently-running kernel context (kmain)
|
||||
registers itself as task 0, so there's always something to switch *from*.
|
||||
|
||||
## The context switch
|
||||
|
||||
Switching tasks means swapping stacks. `switch_context(old, new)` (in `isr.s`)
|
||||
saves the **callee-saved** registers on the current stack, stores the stack pointer
|
||||
into the old task, loads the new task's stack pointer, restores *its* callee-saved
|
||||
registers, and `ret`s — landing wherever the new task was last suspended. Only
|
||||
callee-saved registers are handled explicitly: to the compiler this looks like a
|
||||
normal function call, so it already preserves the caller-saved ones itself (this is
|
||||
the [SysV](sysv.md) convention doing the work).
|
||||
|
||||
A **freshly spawned** task has never run, so there's nothing to restore. Its stack
|
||||
is faked to look as if it had just called `switch_context`: `init_task_stack` lays
|
||||
down a return address pointing at `task_trampoline` and zeroed callee-saved slots
|
||||
(smuggling the entry function in via the `r15` slot). When first switched to, the
|
||||
`ret` lands in the trampoline, which enables interrupts and calls the entry.
|
||||
|
||||
## Two ways to switch, one flag discipline
|
||||
|
||||
`schedule()` — pick the best task and switch — runs from two places:
|
||||
|
||||
- **`yield()`** — a task voluntarily gives up the CPU.
|
||||
- **`tick()`** — the 1000 Hz [timer](../device-driver-development-guide/device-interrupts.md) preempts the running
|
||||
task. This is what lets a task that never yields still share the CPU.
|
||||
|
||||
The subtlety in mixing them is the **interrupt flag (IF)**. The rule: `switch_context`
|
||||
is always entered with interrupts *disabled* — naturally so inside the timer ISR,
|
||||
and explicitly (`cli`) in `yield`. Then every task ends up with interrupts enabled
|
||||
again through whichever path resumes it:
|
||||
|
||||
- a task suspended in `yield` re-enables them (`sti`) right after `schedule` returns;
|
||||
- a task suspended mid-ISR resumes through the interrupt return (`iretq`), which
|
||||
restores the `RFLAGS` it had when it was preempted (IF set);
|
||||
- a brand-new task enables them in the trampoline.
|
||||
|
||||
One related detail: the timer interrupt is **acknowledged (EOI) before** its handler
|
||||
runs, so a handler that switches tasks and doesn't return promptly can't stall the
|
||||
LAPIC from delivering the next tick.
|
||||
|
||||
## Priority selection, in O(1)
|
||||
|
||||
Ready tasks live in a **FIFO queue per priority level** (8 levels), plus a
|
||||
**bitmap** with one bit per non-empty level. Picking the next task is: find the
|
||||
highest set bit (one instruction), take the front of that level's queue. No list
|
||||
walking, no scanning — the decision cost is constant regardless of how many tasks
|
||||
exist, which is what a real-time scheduler needs.
|
||||
|
||||
- **Highest priority wins.** A ready high-priority task always runs before a
|
||||
lower-priority one.
|
||||
- **Round-robin within a level.** When a task is descheduled it goes to the *back*
|
||||
of its level's queue, so equal-priority tasks share the CPU fairly.
|
||||
|
||||
## Affinity: pinning a task to a core
|
||||
|
||||
By default a task runs on **any** core — the ready queue above is global, and any
|
||||
idle core pulls the highest-priority task from it (work-conserving; see
|
||||
[smp.md](smp.md)). A task can instead be **pinned** to one core with
|
||||
`spawnOn(entry, priority, cpu)`, giving it an *affinity*: it will only ever run
|
||||
there, never migrating.
|
||||
|
||||
Mechanically, each core has its **own** pinned queue (same 8-level FIFO + bitmap)
|
||||
alongside the global one. A pinned task is enqueued only into its core's pinned
|
||||
queue; selection compares the top of the global queue and the running core's pinned
|
||||
queue and takes the higher priority (still O(1) — two bit-scans and a compare), with
|
||||
a pinned task winning an equal-priority tie so it can't be starved by global work.
|
||||
Because every queue is mutated under the [big kernel lock](smp.md), one core enqueuing
|
||||
into another core's pinned queue is safe.
|
||||
|
||||
This is the *explicit-affinity* model (no surprise migration mid-deadline), which is
|
||||
the more real-time-predictable direction. `spawnOn` refuses to pin to an offline or
|
||||
out-of-range core — it creates the task unpinned instead, so it still runs somewhere
|
||||
rather than stranding in a queue no core services, and returns whether the pin took.
|
||||
|
||||
## Sleeping and the idle task
|
||||
|
||||
A task can **block** — give up the CPU until an event, rather than busy-wait
|
||||
(busy-waiting is the enemy of a real-time system: it wastes cycles a
|
||||
higher-priority task should get). The first form is time-based: **`sleep(ms)`**
|
||||
marks the task blocked with a wake deadline and switches away. On every tick the
|
||||
timer wakes any task whose deadline has passed (a bounded scan, so it stays
|
||||
deterministic), which makes it ready again; the scheduler then runs it when its
|
||||
priority comes up. `sleep` measures its deadline on the [calibrated
|
||||
clock](../device-driver-development-guide/device-interrupts.md), so it's real time.
|
||||
|
||||
When *every* task is blocked, something still has to run — so there's an **idle
|
||||
task** at the lowest priority that just `hlt`s until the next interrupt (see
|
||||
[halting.md](halting.md)). Because it's always runnable, the scheduler always has a
|
||||
task to pick, and the "nothing to run" case never arises.
|
||||
|
||||
## Event-based blocking
|
||||
|
||||
The other form of blocking is waiting for an **event** rather than a duration. A
|
||||
**wait queue** is a set of tasks parked until something happens: `wait(wq)` blocks
|
||||
the caller on it, `wake(wq)` moves the highest-priority waiter back to ready
|
||||
(preempting if it now outranks the running task). A task links into a wait queue
|
||||
through the same field the ready queues use — it's in exactly one queue at a time.
|
||||
These are the primitives locks, semaphores and [IPC](../device-driver-development-guide/ipc.md) are built on.
|
||||
|
||||
Blocking safely needs **composable critical sections**. A blanket `cli`/`sti` pair
|
||||
doesn't nest: an IPC channel that `cli`s and then calls `wait` would have `wait`'s
|
||||
`sti` re-enable interrupts too early, mid-operation. So the blocking primitives use
|
||||
`saveInterrupts` / `restoreInterrupts` — capture the interrupt flag, disable, and
|
||||
later restore *only if it was set* — which nests correctly. The invariant that
|
||||
makes it all work: `schedule()` is always entered with interrupts disabled, so a
|
||||
task always resumes from a switch with interrupts disabled and can restore its
|
||||
caller's state.
|
||||
|
||||
## Verifying it
|
||||
|
||||
Three tests (see [testing.md](../testing.md)) prove the guarantees:
|
||||
|
||||
- **`sched`** spawns three tasks that busy-loop *without ever yielding*. They all
|
||||
make progress — which can only happen if the timer is **preempting** between them
|
||||
and the context switch is correct (nothing yields voluntarily).
|
||||
- **`priority`** (with preemption off, for determinism) spawns tasks at three
|
||||
priorities; they run and exit **highest-priority first** — `[6, 4, 2]`.
|
||||
- **`sleep`** blocks a task for 50 ms and checks the elapsed time on the clock — a
|
||||
real block (the idle task runs meanwhile), not a busy-wait.
|
||||
- **`event`** blocks a task on a wait queue; waking it (from another task) resumes
|
||||
it, and since it's higher priority it preempts immediately.
|
||||
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Priority inheritance** — still open. Tasks now do block on shared resources
|
||||
(IPC rendezvous, the big kernel lock), and nothing yet bounds priority
|
||||
inversion — a [real-time](../vision.md) requirement.
|
||||
- **Task exit / a reaper** — done. A dying task goes on its core's reap list in a
|
||||
`.reaping` state; the timer tick drains the list, frees the stack back to the
|
||||
heap, and recycles the task-table slot.
|
||||
- **Per-address-space tasks** — done. User processes each own an address space,
|
||||
and the context switch reloads CR3 when the target's tables differ (see
|
||||
[paging.md](paging.md)).
|
||||
@@ -0,0 +1,310 @@
|
||||
# Shared fate: whole-process death (plan)
|
||||
|
||||
**Status: implemented 2026-07-22 (branch shared-fate), M1–M4 all landed; leader
|
||||
`thread_exit` → `-EPERM` as decided. One scope addition forced by M4: the
|
||||
per-task DMA/shared-memory arena cursors moved to the per-space object (the
|
||||
`shm-mapping-ref` test could not distinguish corruption-by-remap from
|
||||
corruption-by-free while sibling threads overlapped the arena) — the same move
|
||||
the mmap/MMIO cursors made in threading M7.**
|
||||
|
||||
[threading.md](threading.md) promises that a process dies *whole* — a fault in any
|
||||
thread, or a kill, takes down every thread. The kernel doesn't do that yet: every
|
||||
death path (`exit`, `thread_exit`, a ring-3 fault, `process_kill`) tears down
|
||||
exactly one `Task`, and the address-space refcount keeps the space alive for the
|
||||
siblings — so a faulting worker orphans its threads, which keep running in the
|
||||
possibly-corrupted address space (`process.zig` `killCurrentProcess`;
|
||||
`scheduler.zig` `exitUserLocked`). This plan closes that gap the way Linux, Windows,
|
||||
and Fuchsia all did: **the process is the unit of fate; only a voluntary
|
||||
`thread_exit` is per-thread.** The supervisor contract — one exit notification,
|
||||
then `process_exit_reason` — is deliberately unchanged.
|
||||
|
||||
## The contract
|
||||
|
||||
| Event | Who dies | Reason the supervisor reads (leader's record) |
|
||||
|---|---|---|
|
||||
| CPU fault in **any** thread (recoverable vector) | the whole group | the fault class (`segmentation_fault`, …) |
|
||||
| `process_kill` on **any member id** | the whole group | `killed` |
|
||||
| `exit(code)` from **any** thread | the whole group | `exited` (0) / `aborted` (≠0) |
|
||||
| `thread_exit` from a worker | that worker only | — (per-task record: `exited`) |
|
||||
| `thread_exit` from the **leader** | nobody — refused, `-EPERM` (decided, below) | — |
|
||||
| NMI, double fault, machine check | the core halts (unchanged) | — |
|
||||
|
||||
`exit` gaining group semantics is the `exit_group` lesson from Linux: the runtime's
|
||||
main-return path calls `exit`, and a process whose main returned must not leave
|
||||
workers running. `thread_exit` (what the worker trampoline calls) keeps today's
|
||||
per-thread behavior, refcount and all.
|
||||
|
||||
**The leader-`thread_exit` rule (decided: refuse).** The syscall is reachable from
|
||||
the leader even though the runtime never does it. Options weighed: (i) **refuse
|
||||
with `-EPERM`** — cheapest and honest; the group ends only through
|
||||
`exit`/fault/kill; (ii) escalate to `exit(0)` — Linux-flavored, but silently turns
|
||||
a (buggy) library call into process death; (iii) a Linux-style zombie leader whose
|
||||
slot survives until the group ends — the most faithful, and by far the most
|
||||
machinery. **(i) chosen at sign-off**; rows and tests below follow it.
|
||||
|
||||
## Group identity: a leader id
|
||||
|
||||
Nothing on `Task` names a process today — threads are tied to their process only by
|
||||
an equal `address_space`, their name is `"thread"`, and their `supervisor` is
|
||||
whichever *task* spawned them (possibly another worker), so supervision links form a
|
||||
chain, not a group. Scanning by `address_space` is also fragile during teardown,
|
||||
because both death paths zero it.
|
||||
|
||||
So: **add `leader: u32` to `Task`** — the Linux tgid, in danos clothes.
|
||||
`spawnProcessSupervised` sets `leader = own id`; `spawnThreadSupervised` copies the
|
||||
*caller's* leader; kernel tasks keep `leader = 0`, which is never followed. The
|
||||
leader id is exactly the id `system_spawn` returned to the supervisor, so the
|
||||
outside world already speaks it. Group membership = equal `leader`, where a *live
|
||||
member* means `state` ∈ {`.ready`, `.blocked`, `.running`} — the same filter
|
||||
`taskByIdLocked` applies; `.reaping` corpses are excluded. While the field is being
|
||||
introduced, add `leader` to `ProcessDescriptor` too (the ABI is private, so this is
|
||||
cheap now and lets `process_enumerate` consumers group threads).
|
||||
|
||||
`process_kill` re-derives its authority through the leader — with the existing
|
||||
guard order preserved: a kernel task (`address_space == 0`) is `-ESRCH` *before*
|
||||
any leader resolution (the kernel test asserts exactly that). Then: resolve the
|
||||
target, follow `target.leader`, require `leader.supervisor == caller`. The kill
|
||||
capability becomes per-*process*, aimed at any member id, and the odd
|
||||
today-behavior where a worker can be individually killed by its spawning thread
|
||||
disappears. (Audited: nothing in-tree kills a worker tid or relies on
|
||||
thread-supervisor kill semantics.)
|
||||
|
||||
## The group-dying latch
|
||||
|
||||
`AddressSpaceRef` — one per space, refcounted by its member tasks, recycled with a
|
||||
full struct re-init — gains the group-death state:
|
||||
|
||||
```zig
|
||||
dying: bool = false, // set by the first trigger; never cleared
|
||||
group_reason: abi.ExitReason, // what the leader's record will say
|
||||
exit_endpoint: ?*ipc.Endpoint, // the leader's counted ref, moved here
|
||||
leader: u32,
|
||||
```
|
||||
|
||||
The latch answers three attacks the red team confirmed against a latch-free
|
||||
design:
|
||||
|
||||
- **The spawn gate.** A member already *inside* `thread_spawn` on another core
|
||||
when the fan-out runs (it passed the syscall-entry `kill_pending` check, then
|
||||
spun on the BKL) would otherwise complete the spawn after the fan-out's lock
|
||||
hold ends — a fresh, uncondemned member that escapes the kill and, worse, holds
|
||||
a space reference that keeps the group-death hook from ever firing. Fix:
|
||||
`retainAddressSpace` (equivalently `spawnUserLocked`) **refuses a dying
|
||||
space**; the in-flight `thread_spawn` fails with `-ESRCH` under the same lock
|
||||
that would have created the member.
|
||||
- **Concurrent triggers.** A second member faulting (or exiting) on another core
|
||||
while the first fan-out runs must not re-run the fan-out, double-bump
|
||||
`fault_kill_count`, or re-stamp reasons. Every kill path checks the latch first:
|
||||
already dying → skip straight to `terminateCurrentLocked`, no stamp, no count.
|
||||
First trigger wins, deterministically. `exit_reason` and `fault_kill_count`
|
||||
writes move under the BKL as part of this.
|
||||
- **Notification ownership.** The leader's `exit_endpoint` is a counted
|
||||
birth-to-death reference dropped at notify time. The stamp pass **moves** that
|
||||
reference onto the `AddressSpaceRef` and nulls `Task.exit_endpoint` in the same
|
||||
hold, so the leader's own `releaseTaskResourcesLocked` sees null (no early
|
||||
notify, no double drop); the group-death hook notifies and drops exactly once.
|
||||
|
||||
## The fan-out: `killGroupLocked`
|
||||
|
||||
One new function in `process.zig`, running under a **single BKL hold** (built from
|
||||
the `*Locked` primitives — the lock is non-recursive, and `terminateCurrentLocked`
|
||||
never returns, which forces the shape):
|
||||
|
||||
```
|
||||
killGroupLocked(leader: u32, reason: ExitReason, trigger: ?*Task)
|
||||
0. Latch: AddressSpaceRef.dying = true, stash {reason, leader,
|
||||
leader's exit_endpoint (moved)}.
|
||||
1. Stamp pass: the LEADER's exit_reason = reason — the leader's
|
||||
record is the one the supervisor can read, so it carries the
|
||||
group reason even when the trigger is a worker. The trigger
|
||||
also keeps `reason` (its own record tells the truth); every
|
||||
other live member gets .killed. All members get kill_pending.
|
||||
Stamping precedes any teardown, because recordExitLocked
|
||||
snapshots the reason first thing.
|
||||
2. Reap pass, to fixpoint: reap every member in .ready or .blocked
|
||||
via reapTaskLocked, re-reading Task.state each iteration — a
|
||||
member's teardown can wake another member (-EPEER wakes, joiner
|
||||
wakes), flipping it .blocked → .ready behind the scan cursor.
|
||||
Terminates in ≤ one pass per member: the scrub calls in
|
||||
releaseTaskResourcesLocked (abandonSenderLocked,
|
||||
removeFromWaitQueueLocked, forgetIpcClientLocked,
|
||||
killOwnedEndpointsLocked) run before destroy, so no wake path
|
||||
holds a pointer to a reaped member.
|
||||
3. Members .running on other cores stay condemned (kill_pending);
|
||||
a condemned member dies at its next syscall entry, at its own
|
||||
core's next tick while in user mode, or — once it blocks or is
|
||||
preempted — at any core's next tick reap. There is no kill IPI.
|
||||
(The entry check reads kill_pending unlocked; benign on
|
||||
x86-TSO — a missed read is caught by the next delivery point —
|
||||
but make the field atomic when touching it.)
|
||||
4. If the current task is a member (fault, exit, in-group kill):
|
||||
terminateCurrentLocked, last, because it switches away and the
|
||||
reap paths free the kernel stack being stood on.
|
||||
If the caller is outside the group (supervisor kill): return.
|
||||
```
|
||||
|
||||
The invariants this preserves, each load-bearing today:
|
||||
|
||||
- **Only `.ready`/`.blocked` tasks are reaped synchronously.** A member running on
|
||||
another core can only be condemned — it tears itself down after switching CR3
|
||||
off the dying page tables (the stack it stands on is freed later by the reap
|
||||
list), and its address-space reference protects the page tables its CR3 still
|
||||
points at. Force-destroying the space under a running sibling is the one
|
||||
unrecoverable mistake available here.
|
||||
- **The refcount decides when the space dies.** Reaping N members drops N
|
||||
references; the last drop — possibly on a condemned sibling's core, a tick
|
||||
later — destroys the space. No path forces it.
|
||||
- **`fault_kill_count` bumps once per group**, not per member (`fault-recovery`
|
||||
asserts `== 1` exactly); the latch is what enforces this under racing faults.
|
||||
- **Per-tid resource sweeps stay per-tid.** Each member's
|
||||
`releaseTaskResourcesLocked` releases what *that tid* owns — claims, GSI/MSI
|
||||
bindings, registered endpoints, handles. That keying is correct under shared
|
||||
fate (and is today's hazard: a lone worker death already yanks its claims out
|
||||
from under live siblings). A worker that *does* carry an `exit_endpoint` (the
|
||||
ABI allows it; the runtime passes `no_cap`) keeps today's per-task posting at
|
||||
its own teardown — only the leader's notification moves.
|
||||
|
||||
## When is the group dead? The notification
|
||||
|
||||
Today each task posts its own exit notification as the *last* step of its release,
|
||||
so a supervisor observes a fully-released child. For a group that guarantee must
|
||||
hold for the **whole group**: if the leader's notification fires while a condemned
|
||||
sibling still runs on another core, the device manager can respawn the driver into
|
||||
a claim conflict with a not-yet-dead sibling.
|
||||
|
||||
The clean fix falls out of the refcount: **the group is dead exactly when the
|
||||
address space is destroyed.** `releaseAddressSpace`'s last-drop path calls a new
|
||||
`group_exit_hook` (the scheduler already calls up through hooks —
|
||||
`terminate_current_hook` — precisely to keep this layering), which:
|
||||
|
||||
1. **re-stamps the leader's exit record** with the stashed `group_reason` — the
|
||||
record is written (again) at group-death time, so "the reason is recorded
|
||||
before the notification posts" stays true and a supervisor can never be
|
||||
notified and then read `-ESRCH` because the burst evicted an old record;
|
||||
2. posts the leader's exit notification (and subscriber broadcast) from the
|
||||
stashed endpoint, and drops that reference — exactly once.
|
||||
|
||||
Both `releaseAddressSpace` call sites (`exitUserLocked`, `destroyTaskLocked`) run
|
||||
under the BKL, so the hook does too; its wakes are safe at both (verified). For a
|
||||
single-threaded process the behavior is *externally indistinguishable* from
|
||||
today — the order of notify vs. destroy inverts, but both sit inside one lock
|
||||
hold, so no other core can observe the space destroyed but the notification
|
||||
unposted, or vice versa. That sentence is the correctness argument; it is also the
|
||||
first invariant to re-examine if the BKL is ever split, along with
|
||||
`killGroupLocked`'s single-hold atomicity. (Hand-built spaces that were never
|
||||
retained take the immediate-destroy path and are out of the hook's scope.)
|
||||
|
||||
Workers' `exit_subscribers` broadcasts still fire per task — the FAT server's
|
||||
dead-client sweep is keyed by tid and needs those.
|
||||
|
||||
**Signals.** `signal_bind` is per-task and the service harness binds on the main
|
||||
thread, so signals address the leader in practice; that stays. During a group
|
||||
death, `process_signal` may return `0` (accepted by a condemned member — never
|
||||
delivered, every delivery point kills first) or `-ESRCH` (member already reaped);
|
||||
init's stop sequence already tolerates both, and its timer escalation to
|
||||
`process_kill` covers the gap. `process_signal` follows `process_kill`'s
|
||||
leader re-key for consistency.
|
||||
|
||||
## The shared-memory frame hazard
|
||||
|
||||
`dropSharedMemoryReference` frees a region's physical frames when the last *handle*
|
||||
reference drops, but mappings die only with the address space. If the last handle
|
||||
lived in a torn-down member while any task still has the region mapped, that task
|
||||
holds a live mapping onto freed frames — and the red team showed this is **not**
|
||||
group-specific: a plain `thread_exit` of the handle-holding thread, or a last-ref
|
||||
drop by a task *outside* the dying group during the condemned window, hits the same
|
||||
use-after-free.
|
||||
|
||||
So the fix is a property of the **object**, not the dropper: give
|
||||
`SharedMemoryObject` a per-*mapping* reference — `shared_memory_map` (and create's
|
||||
self-map) retains; each space's destruction releases. "Last reference" then means
|
||||
*no handles and no mappings*, both hazard paths collapse into the existing
|
||||
refcount, and no group-kill special case is needed at all.
|
||||
|
||||
## Deliberately unchanged
|
||||
|
||||
- Worker `thread_exit`: per-thread, full per-tid resource sweep, refcount drop.
|
||||
- The condemned-but-running window: a member on another core can finish its
|
||||
in-flight syscall and run user code for up to a tick before dying — identical to
|
||||
today's single-task `process_kill` semantics ("prompt but asynchronous, like a
|
||||
Unix signal"). A kill IPI would shrink it; it is not part of this plan.
|
||||
- `thread_join` returns 0 for a killed thread; joiners inside a dying group are
|
||||
woken and then reaped like any member.
|
||||
- The `.reaping` state, reap lists, and stack reaper.
|
||||
|
||||
## Accepted limits (documented, not fixed here)
|
||||
|
||||
- **Notify-ring overflow**: a group death posts one subscriber badge per member
|
||||
into 8-slot rings; a >8-member group can drop badges. Group size is bounded by
|
||||
the 48-task table; today's largest production group is 2 (display) and
|
||||
thread-test already reaches 5.
|
||||
- **Exit-record ring pressure**: one 64-entry ring, one record per member — made
|
||||
harmless for the supervisor by the hook's group-death re-stamp.
|
||||
- **Enumerate shows a partial group** mid-death: reaped members vanish at once,
|
||||
condemned members linger up to a tick (audited: no in-tree consumer
|
||||
misbehaves; the `leader` field in `ProcessDescriptor` lets future consumers
|
||||
group correctly).
|
||||
- **Pre-existing reap race, not widened**: a task preempted *mid-syscall* is
|
||||
`.ready` with `in_system_call = true`, and the tick's reap loop will reap it —
|
||||
an existing hazard the fan-out inherits but must not add new instances of.
|
||||
Filed to investigate separately.
|
||||
- **Per-task DMA/shm cursors** — *fixed during M4 after all*: the
|
||||
`shm-mapping-ref` test tripped the overlap (the sibling's churn regions mapped
|
||||
over the worker's region), so both cursors moved to the `AddressSpaceRef`
|
||||
like the mmap/MMIO cursors before them. The post-implementation review then
|
||||
found the other half: the shm/DMA page-table walks and their pmm/heap calls
|
||||
ran *outside* the big kernel lock — pre-existing, but fatal once siblings
|
||||
were invited to race them (and a plausible root for the long-standing
|
||||
intermittent AP ring-3 fault at the shm base). All three paths now follow
|
||||
the mmap discipline: metadata and allocation under one hold, the map itself
|
||||
per-page under brief holds.
|
||||
- **Mapping-record slots are never recycled**: 16 per space, one per
|
||||
`shared_memory_create`/`map`, freed only at space destruction (there is no
|
||||
shm unmap). A long-lived compositor that churns surfaces will hit the cap;
|
||||
the failure is a clean refused create, and slot recycling can ride whatever
|
||||
adds `shared_memory_unmap`.
|
||||
- **Two properties lack direct tests**: the spawn gate (an in-flight
|
||||
`thread_spawn` racing the fan-out — inherently nondeterministic to arrange;
|
||||
covered by code inspection and the `-ESRCH` path) and the `process_signal`
|
||||
leader re-key (exercised only implicitly by the signals case).
|
||||
|
||||
## Milestones
|
||||
|
||||
- **M1 — the leader id.** `Task.leader` (kernel tasks: 0, never followed), set on
|
||||
both spawn paths; `leader` added to `ProcessDescriptor`; `process_kill` and
|
||||
`process_signal` re-keyed (kernel-task `-ESRCH` guard *before* leader
|
||||
resolution). No fan-out yet. Existing tests must pass untouched.
|
||||
- **M2 — the latch + fan-out.** `AddressSpaceRef.dying` + stash;
|
||||
`retainAddressSpace` refuses dying spaces; `killGroupLocked`; wire the fault
|
||||
path, `exit`, and `process_kill` into it; leader `thread_exit` → `-EPERM`;
|
||||
`exit_reason`/`fault_kill_count` writes under the BKL; `kill_pending` made
|
||||
atomic. Group notification via the `group_exit_hook` re-stamp + post.
|
||||
- **M3 — shared-memory mapping refs.** `SharedMemoryObject` counts mappings;
|
||||
space destruction releases them; frames free only at zero handles *and* zero
|
||||
mappings.
|
||||
- **M4 — tests + docs.** New `-Dtest-case`s (all `smp: 4` where cross-core
|
||||
matters), driving `thread-test` with new argv modes:
|
||||
- `thread-fault-group`: a worker faults; assert both tasks gone from
|
||||
`enumerate`, `fault_kill_count == 1`, `process_exit_reason(leader) ==
|
||||
segmentation_fault`, address-space and stack-bytes counters return to base.
|
||||
- `kill-threaded-group`: `process_kill(leader)` with a worker spinning on
|
||||
another core; assert the worker dies by the deferred path, exactly one exit
|
||||
badge, delivered only after both members are dead, and
|
||||
`process_exit_reason(leader) == .killed`.
|
||||
- `kill-via-worker-tid`: `process_kill(worker)` kills the whole group;
|
||||
`-EPERM` for a non-supervisor aiming at the worker.
|
||||
- `racing-triggers`: two members fault/exit simultaneously on different cores;
|
||||
assert a deterministic leader reason and `fault_kill_count == 1`.
|
||||
- `exit-group`: a *worker* calls `exit(3)`; group dies, leader reason
|
||||
`.aborted`.
|
||||
- `leader-thread-exit`: asserts the chosen rule (`-EPERM`, workers unaffected).
|
||||
- `thread-exit-solo`: regression — worker `thread_exit` still leaves siblings
|
||||
running.
|
||||
- `group-claim-release`: a member claims a device; assert the claim is free and
|
||||
the leader notification arrives only after every member is dead.
|
||||
- `shm-mapping-ref`: last handle dropped by a dying thread; sibling's mapping
|
||||
stays valid until space death (M3 regression).
|
||||
Then update [threading.md](threading.md) (the shared-fate gap note),
|
||||
[process-lifecycle.md](process-lifecycle.md),
|
||||
[process-management.md](process-management.md), and
|
||||
[ipc.md](../device-driver-development-guide/ipc.md)/[drivers.md](../device-driver-development-guide/drivers.md) mentions.
|
||||
@@ -0,0 +1,275 @@
|
||||
# SMP: multiple cores, the microkernel way
|
||||
|
||||
A design/research note that predates the build — danos now runs on **multiple
|
||||
cores** by default (see [Implementation status](#implementation-status) and
|
||||
[scheduling.md](scheduling.md)). This maps how microkernels — especially the L4
|
||||
family and seL4 — handle **symmetric multiprocessing (SMP)**, the plan and
|
||||
reading list the port followed. It also flags where those choices depend on
|
||||
whether danos is chasing **real-time** or **resilience** (see the note at the end).
|
||||
|
||||
## First, the vocabulary
|
||||
|
||||
Three independent things people conflate (see [scheduling.md](scheduling.md) for the
|
||||
danos specifics):
|
||||
|
||||
- **Task capacity** (`max_tasks`) — how many tasks can *exist*. A table size.
|
||||
- **Cores** — how many tasks *run at the same instant*. One running task per core.
|
||||
- **Real-time** — whether timing is *predictable*. Comes from bounded operations
|
||||
(our O(1) scheduler), not from core count.
|
||||
|
||||
danos now runs on multiple cores. The firmware starts only the **bootstrap processor
|
||||
(BSP)**; the kernel wakes the other cores (**application processors**, APs) with
|
||||
INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor tables,
|
||||
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
|
||||
the `smp` self-test confirms worker tasks executing on all four cores at once under
|
||||
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
|
||||
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
|
||||
and thread-to-core affinity (see [Implementation status](#implementation-status)).
|
||||
|
||||
## The common microkernel instinct: don't share kernel state
|
||||
|
||||
Monolithic kernels (Linux) share large amounts of state across cores behind many
|
||||
fine-grained locks. Microkernels lean the other way — the kernel does *little* (IPC,
|
||||
scheduling, capabilities), so the pressure is to make kernel state **per-core** and
|
||||
coordinate cores with **inter-processor interrupts (IPIs) or messages** rather than
|
||||
shared, locked data structures. Two hallmarks follow:
|
||||
|
||||
- **Thread-to-core affinity.** Threads are usually *bound* to a core; migration is an
|
||||
explicit operation, not automatic load-balancing.
|
||||
- **Policy in user space.** *Which* core a thread runs on tends to be a user-level
|
||||
decision (a scheduler/manager server); the kernel just provides the mechanism to
|
||||
run it there and to signal across cores. This matches the microkernel creed:
|
||||
mechanism in the kernel, policy outside.
|
||||
|
||||
Within that instinct, the L4 family split on **how much to lock**.
|
||||
|
||||
## seL4: the "big kernel lock" — and why it's not a hack
|
||||
|
||||
seL4's choice is striking: a **single big kernel lock (BKL)**. Only one core runs
|
||||
*kernel* code at a time; **user code runs fully in parallel** on all cores. A core
|
||||
that traps into the kernel takes the global lock, does its (short) work, releases it.
|
||||
|
||||
Why so coarse? **Formal verification.** seL4's whole value is a machine-checked
|
||||
correctness proof, built for a *uniprocessor* kernel — reasoning about one thread of
|
||||
kernel execution. Fine-grained SMP locking explodes the interleavings you'd have to
|
||||
reason about. The big lock **serialises kernel execution so the single-core reasoning
|
||||
still holds**. It trades kernel scalability for verifiability.
|
||||
|
||||
And it works better than it sounds, *because the kernel does so little*: the lock is
|
||||
held for short, bounded intervals, while the real work (drivers, services) runs in
|
||||
user space in parallel, outside the lock. Scheduling is otherwise **per-core** (each
|
||||
core its own ready queues), threads carry an **affinity**, and cross-core IPC costs an
|
||||
IPI.
|
||||
|
||||
> Nuance: the fully *verified* seL4 configuration is the uniprocessor one. The
|
||||
> SMP/big-lock version isn't covered by the same end-to-end proof — extending
|
||||
> verification to multicore has been ongoing research. So the big lock is partly
|
||||
> "stay close to the thing we proved."
|
||||
|
||||
seL4 also layers **MCS** (mixed-criticality scheduling) on top: **scheduling
|
||||
contexts** carrying a time *budget* and *period*, so a thread can't overrun its share
|
||||
— temporal isolation, reasoned about per core. This is the seriously real-time part.
|
||||
|
||||
## Fiasco.OC / NOVA: per-CPU, finer-grained
|
||||
|
||||
Not all L4s took the big lock. **Fiasco.OC** (TU Dresden L4, part of L4Re) is
|
||||
**per-CPU**: per-CPU run queues, CPU-local kernel objects, IPIs for the rare
|
||||
cross-CPU operations. Threads bind to a CPU; moving one is explicit. Scales better
|
||||
than a big lock, more complex, and without seL4's verification constraint forcing the
|
||||
issue. **NOVA** (a microhypervisor) is similarly per-CPU. The shared pattern: make
|
||||
everything CPU-local you can, and when cores must interact, **send a message/IPI**
|
||||
instead of touching shared data.
|
||||
|
||||
## The extreme: the "multikernel"
|
||||
|
||||
Taken to its logical end you get **Barrelfish** (ETH Zurich): treat a multicore
|
||||
machine as a *network of cores*, each running its **own kernel instance**, sharing
|
||||
**no** kernel memory, communicating **only by message passing** — the microkernel's
|
||||
IPC philosophy applied to the kernel's own structure. seL4's "clustered multikernel"
|
||||
explorations use the same idea: groups of cores, each cluster a big-lock domain,
|
||||
clusters talking by messages. The insight: if you're already committed to messages
|
||||
for user-space isolation, structure the kernel across cores the same way and sidestep
|
||||
shared-memory locking entirely.
|
||||
|
||||
## Does the right choice depend on real-time vs resilience?
|
||||
|
||||
Yes — and this is the branch that matters for danos right now.
|
||||
|
||||
- **If the goal is hard real-time:** favour **per-core scheduling with fixed
|
||||
affinity**. A thread never gets surprise-migrated mid-deadline, and each core's
|
||||
timeline can be reasoned about in isolation. Global load-balancing (Linux's default)
|
||||
is great for throughput and *bad* for determinism, which is why RT microkernels
|
||||
mostly pin threads. seL4's MCS scheduling contexts are the reference model.
|
||||
- **If the goal is resilience / restartability:** the SMP priority shifts to **fault
|
||||
isolation and recovery**, not timing. What matters is that a failed component (a
|
||||
driver, a service) on any core can be **killed and restarted** without taking the
|
||||
system down — which is a property of address-space isolation + a supervising
|
||||
restart server (below), *largely orthogonal to how cores are scheduled*. A big lock
|
||||
is perfectly fine here; you're optimising for "a crash is contained and
|
||||
recoverable," not "latency is bounded to N µs."
|
||||
- **If the goal is throughput:** you'd care about lock contention and per-core
|
||||
queues — the least microkernel-flavoured of the three.
|
||||
|
||||
These pull in different directions, so **picking the primary goal comes before
|
||||
picking the SMP design.** (danos's founding assumption was real-time; that's under
|
||||
active reconsideration in favour of resilience — see [vision.md](../vision.md).)
|
||||
|
||||
## What this would mean for danos
|
||||
|
||||
Whatever the top goal, the *sequence* is the same and seL4 validates starting simple:
|
||||
|
||||
1. **Enumerate cores** — needs [device discovery](discovery.md) (ACPI MADT on x86,
|
||||
device tree on ARM). SMP is a concrete consumer of that work. **Done on x86:** the
|
||||
MADT parse records every usable Local APIC — with the `apic_id` an AP wake targets —
|
||||
and `platform.cpus()` returns the list (see [discovery.md](discovery.md)). The boot
|
||||
log reports the count; the ARM (device-tree) path still needs it.
|
||||
2. **Wake the APs** — INIT–SIPI–SIPI on x86; PSCI/spin-tables on ARM. Each core brings
|
||||
up its own tables, timer, and idle task. **Done on x86** — cores climb to long mode,
|
||||
set up their own GDT/TSS, and enter the scheduler; tasks run in parallel across all
|
||||
cores ([status](#implementation-status)).
|
||||
3. **Start with a big kernel lock.** It's a legitimate first design, not a shortcut —
|
||||
philosophically aligned with a tiny kernel, and it lets the single-core correctness
|
||||
model you already have (the interrupt-flag discipline in
|
||||
[scheduling.md](scheduling.md)) stay largely intact: one lock around kernel entry
|
||||
instead of rethinking every critical section. **Done** — see
|
||||
`system/kernel/sync.zig`.
|
||||
4. **Later, if contention bites,** evolve toward **per-core run queues + explicit
|
||||
affinity** (the Fiasco.OC direction) — also the more real-time-predictable model.
|
||||
5. **Placement stays a user-space policy** — the kernel runs a thread on the core it's
|
||||
told to, a user-level manager decides which.
|
||||
|
||||
Big-lock-first → per-core-later. The affinity/MCS depth is only worth it if real-time
|
||||
turns out to be the actual goal.
|
||||
|
||||
## Implementation status
|
||||
|
||||
The "wake + schedule" build (real parallel task execution) is going in as a sequence
|
||||
of green checkpoints — each step keeps the single-core test suite passing before the
|
||||
next lands.
|
||||
|
||||
**Done:**
|
||||
|
||||
- **Core enumeration** — the MADT parse records every usable Local APIC (with its
|
||||
`apic_id`, which an AP wake targets); `platform.cpus()` returns the list. See
|
||||
[discovery.md](discovery.md).
|
||||
- **The big kernel lock** (`system/kernel/sync.zig`) — one coarse spinlock guarding the
|
||||
scheduler queues and IPC, always held with local interrupts disabled. It is held
|
||||
*across* a context switch and released by whichever task resumes (the hand-off
|
||||
rule); `task_trampoline` releases it for a freshly-spawned task. `scheduler.zig` and
|
||||
`ipc.zig` run every critical section under it. Uncontended on one core, so behaviour
|
||||
is identical to the old interrupt-flag model.
|
||||
- **Per-CPU state** — a `PerCpu` struct (running task, idle task, APIC id) per core,
|
||||
its pointer kept in the x86 **GS base** (`IA32_GS_BASE`). No `swapgs` was needed at
|
||||
the time — there was no user mode yet; with ring 3 in place, every ring transition
|
||||
now swaps it against the user's own GS base under the `swapgs` discipline (see
|
||||
`system/kernel/architecture/x86_64/per-cpu.zig`), and each AP enables the fast
|
||||
system-call path (`initSystemCall`) for itself at bring-up. The old global
|
||||
`current` is now `thisCpu().current`. The ready
|
||||
queues stay **global** under the lock — work-conserving, so any idle core will pull
|
||||
the highest-priority ready task; per-core queues are a later optimisation.
|
||||
- **AP wake to long mode** — `architecture.startSecondary` drives INIT–SIPI–SIPI (via the
|
||||
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
|
||||
real mode at a low page and runs the [trampoline](../../system/kernel/architecture/x86_64/trampoline.s)
|
||||
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
|
||||
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
|
||||
all four cores report `online`.
|
||||
|
||||
The trampoline earns its complexity from four hardware facts:
|
||||
- a STARTUP IPI vectors a core to physical `vector << 12` (a *byte* vector), so the
|
||||
trampoline must live **below 1 MiB** — the kernel reserves that page from the frame
|
||||
allocator at boot, before paging/heap draw down the scarce low frames;
|
||||
- the blanket RAM identity map is **NX** (W^X), but the AP fetches the trampoline
|
||||
from it under paging, so that one page is made executable for bring-up;
|
||||
- the blob is copied to a page whose address isn't known at link time, so it is
|
||||
**position-independent**: it derives its own base from `CS` and, crucially,
|
||||
addresses data *segment-relative in real mode* (where the segment base already
|
||||
supplies the page base) but *base-register-relative in protected/long mode* (flat
|
||||
segments, base 0). Getting that distinction wrong was the first bug found;
|
||||
- an AP starts with a bare `CR0`/`CR4`, but the kernel is built **with SSE** (the
|
||||
x86_64 baseline) and the compiler emits SSE for things as ordinary as a struct
|
||||
copy — so the trampoline must set `CR4.OSFXSR`/`OSXMMEXCPT` and fix `CR0.EM`/`MP`,
|
||||
or the first SSE instruction on the AP `#UD`s. The BSP inherited those bits from
|
||||
UEFI; the AP has to set them itself. This was the second bug — it masqueraded as a
|
||||
fault in `lgdt` (the first kernel code after entry that the compiler vectorised).
|
||||
|
||||
- **Per-core tables + scheduler entry** — each AP loads **its own GDT** (with its own
|
||||
TSS descriptor) and **its own TSS** (its own IST/`rsp0` stack), loads the shared
|
||||
IDT, enables its LAPIC and timer, then calls the generic `secondaryMain`: it turns
|
||||
its bring-up context into the core's idle task (as task 0 is for the BSP), marks the
|
||||
core online, and enters the run loop. With interrupts on, each core's own timer tick
|
||||
preempts its idle context into whatever the global ready queue offers — so all cores
|
||||
pull real work in parallel. The `smp` test spawns CPU-bound workers and confirms they
|
||||
execute on all four cores at once, and `fault-ap-df` pins a #DF to an AP and checks
|
||||
that core catches it on **its own** IST (a broken per-core TSS would triple-fault) —
|
||||
reported as "core N: …", so a fault is always attributed to the core it happened on,
|
||||
and is contained to that core (the rest of the system keeps running).
|
||||
- **Thread affinity** — `spawnOn(entry, priority, cpu)` pins a task to a core (its own
|
||||
per-core pinned queue, merged with the global queue at selection; see
|
||||
[scheduling.md](scheduling.md#affinity-pinning-a-task-to-a-core)). The `affinity`
|
||||
test confirms a pinned task never migrates. This is the mechanism the fault-on-AP
|
||||
test rides on, and the *explicit-affinity* real-time-predictable model.
|
||||
- **Right-sized footprint** — the per-CPU ceiling (`parameters.maximum_cpus`, one
|
||||
constant shared by discovery, the scheduler, and the per-core GDT/TSS) is generous
|
||||
(128), but
|
||||
the *large* per-core resources — the kernel and IST (double-fault) stacks — are
|
||||
**heap-allocated at bring-up**, only for cores that actually come online. Only the
|
||||
BSP's IST stack is static, because it must exist before the frame allocator does.
|
||||
This kept the kernel image small (a static `[128][16 KiB]` IST array would have been
|
||||
2 MiB of `.bss`); it's a few tens of KiB instead.
|
||||
- **`single_threaded` off** — the kernel was built `single_threaded = true`, which
|
||||
compiles `std.atomic` down to plain non-atomic ops. Harmless on one core, but it
|
||||
quietly breaks the big kernel lock across cores; it's now `false`.
|
||||
- **Re-armable wake + retry** — the trampoline frame is reserved for the system's
|
||||
life, but kept **inert between wakes**: zeroed and non-executable, armed (blob
|
||||
copied in, page made executable) only for the moment a core is actually climbing,
|
||||
then disarmed again. So there's never a dormant executable page, and a core can be
|
||||
(re)woken at any time — `architecture.startSecondary` is one self-contained attempt (arm →
|
||||
INIT–SIPI–SIPI → disarm), and its `INIT` resets a wedged core, so retrying just
|
||||
works. Boot retries a non-responding core up to three times; the same primitive is
|
||||
the groundwork a future **power manager** would drive to bring cores up (and,
|
||||
eventually, its counterpart to take them offline — which additionally needs the
|
||||
core's tasks migrated off first).
|
||||
|
||||
**Next (refinement, not first-light):**
|
||||
|
||||
- **IPIs** — cross-core wake/preempt. Not needed for correctness: an idle core wakes
|
||||
on its own timer tick and pulls ready work then; IPIs only cut that latency from
|
||||
≤1 ms to near-instant.
|
||||
- **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock
|
||||
contention ever bites. (Thread *affinity* already exists — see above; this is the
|
||||
further step of giving each core its own primary run queue for load distribution.)
|
||||
- **Fault recovery** — today a fault halts (only) the faulting core. Turning that into
|
||||
"kill the task, keep the core running" is the [resilience](resilience.md) track (it
|
||||
needs the task's lock/resource state handled), and for taking a core fully offline,
|
||||
its tasks migrated first.
|
||||
|
||||
## Further reading
|
||||
|
||||
**Microkernel SMP & scheduling**
|
||||
- Klein et al., *"seL4: Formal Verification of an OS Kernel"* (SOSP 2009) — the
|
||||
verification that shapes seL4's whole SMP stance.
|
||||
- Lyons et al., *"Scheduling-Context Capabilities: A Principled, Light-Weight OS
|
||||
Mechanism for Managing Time"* (EuroSys 2018) — seL4 MCS, the real-time model.
|
||||
- The **seL4 whitepaper** and "towards a verified multiprocessor seL4" material — the
|
||||
big-lock / clustered-multikernel reasoning.
|
||||
- **Fiasco.OC / L4Re** documentation (TU Dresden) — the per-CPU alternative.
|
||||
- Baumann et al., *"The Multikernel: A New OS Architecture for Scalable Multicore
|
||||
Systems"* (SOSP 2009) — Barrelfish, the share-nothing extreme.
|
||||
|
||||
**Resilience / self-healing (if that's the real goal)**
|
||||
- Herder et al., *"Fault Isolation for Device Drivers"* and the **MINIX 3**
|
||||
*reincarnation server* — a driver crashes, a supervisor restarts it live. The
|
||||
closest existing system to "re-initialise parts of the OS."
|
||||
- **QNX** — commercial microkernel RTOS built on message passing and restartable
|
||||
drivers; good study of the combination.
|
||||
- **Erlang/OTP** *supervision trees* and the *"let it crash"* philosophy — not a
|
||||
kernel, but the canonical design for "isolate failures and restart the failed
|
||||
part," directly relevant to danos's restartability motivation.
|
||||
|
||||
## Related
|
||||
|
||||
- [scheduling.md](scheduling.md) — the single-core scheduler SMP would extend.
|
||||
- [discovery.md](discovery.md) — enumerating cores is a device-discovery problem.
|
||||
- [ipc.md](../device-driver-development-guide/ipc.md) — the message passing cross-core coordination rides on.
|
||||
- [vision.md](../vision.md) — the goals question (real-time vs resilience) this note
|
||||
keeps bumping into.
|
||||
@@ -0,0 +1,95 @@
|
||||
# System Calls
|
||||
System calls (syscalls) are the bridge between your programs and the operating system's restricted core (kernel).
|
||||
|
||||
> **Status:** danos has real user processes. User programs enter the kernel
|
||||
> via the `syscall` instruction (STAR/LSTAR/SFMASK set per core; the entry stub in
|
||||
> `isr.s` does the `swapgs` + kernel-stack switch and reuses the interrupt
|
||||
> dispatcher); the `int 0x80` gate is kept alongside as a minimal test path.
|
||||
> The live table is `system/abi.zig` (private, renumberable — see
|
||||
> [vdso.md](vdso.md) for the public boundary): process lifecycle + threads,
|
||||
> memory (mmap/dma/shared-memory), synchronous + async IPC with capability passing,
|
||||
> device access, time, the tagged-log diagnostics (`debug_write` with a level,
|
||||
> `klog_read`/`klog_status`), and filesystem NAMING (`fs_resolve`/`fs_node`/
|
||||
> `fs_mount`/`fs_unmount` — the kernel VFS root routes paths and serves the
|
||||
> read-only /system initrd mount; file DATA stays with userspace filesystem
|
||||
> servers over the vfs-protocol, docs/vfs-protocol.md).
|
||||
|
||||
## The Mechanism of a Syscall
|
||||
|
||||
A system call follows a highly orchestrated, 6-step sequence to ensure hardware safety and process security:
|
||||
|
||||
1. Setting the Arguments: The program places a unique ID corresponding to the requested service (e.g., \(sys\_write\)) into a specific CPU register (like RAX in x86_64) along with its required parameters.
|
||||
2. Executing the Trap: The program triggers a special CPU instruction, such as syscall or int 0x80. This acts as a software interrupt.
|
||||
3. Mode Switch: The CPU atomically flips its execution privilege from unprivileged User Mode (Ring 3) to the highly privileged Kernel Mode (Ring 0).
|
||||
4. Lookup and Execution: The kernel looks up the syscall ID in a dispatch table and runs the designated internal routine to perform the actual work (like fetching data from the hard drive).
|
||||
5. Returning the Status: The kernel places the result of the operation—or an error code—back into the RAX register.
|
||||
6. Return to User Mode: The kernel executes a return instruction (such as sysret), the CPU shifts back to User Mode, and your program continues executing.
|
||||
|
||||
## Approaches
|
||||
There are two approaches to system calls, stable public ABI and unstable private ABI. A public ABI uses a standard defined set of numbers. It makes it easy to guess, easy to write compilers and tools for. private ABI tend to change the numbers to obscure the numbers, preventing attackers from bypassing libc using CPU instructions directly. Private ABIs force users to use a library like libSystem, which the kernel can inject the syscall numbers. An alternative solution is to just have 1 system number but make it generic by the address to a struct in memory.
|
||||
|
||||
Practical minimum primitives (Monolithic kernel)
|
||||
| Category | Primitive | Detail |
|
||||
|----------|--------------|---------------------------------------------------------------------------------------|
|
||||
| lifecyle | spawn / exec | Loads a binary from storage into memory and begins execution. |
|
||||
| lifecyle | exit | Terminates the current process and frees its memory back to the kernel. |
|
||||
| I/O | read | Requests data from a hardware device or file descriptor into user memory. |
|
||||
| I/O | write | Pushes data from user memory out to a device or file descriptor (like a screen). |
|
||||
| Memory | brk / mmap | Requests the kernel to allocate more physical or virtual memory pages to the process. |
|
||||
| Control | ioctl | A catch-all "escape hatch" call to send hardware-specific commands to device drivers. |
|
||||
|
||||
|
||||
The Microkernel Minimum Set
|
||||
|
||||
Everything else---including`read()`,`write()`,`malloc()`, and`fork()`---will run in user space as servers (e.g., a VFS server, a memory manager server) that threads communicate with using these three calls:
|
||||
|
||||
1. **`IPC_Call(endpoint, message_buffer)`(Synchronous Send + Receive)**
|
||||
- **What it does:**The calling thread sends a message block to a service endpoint and immediately blocks (sleeps) until that service processes the request and sends a reply back.
|
||||
- **Why it's minimal:**Combining*Send*and*Receive*into a single atomic atomic system call eliminates the need for separate tracking and prevents a massive amount of context-switching overhead. This is the foundation of high-performance microkernels like[seL4](https://sel4.systems/)and L4.[[1](https://en.wikipedia.org/wiki/L4_microkernel_family),[2](https://www.microchip.com/en-us/products/microprocessors/64-bit-mpus/pic64hx/ecosystem)]
|
||||
2. **`IPC_ReplyWait(endpoint, reply_buffer)`(Respond + Wait for Next)**
|
||||
- **What it does:**Used strictly by your background user-space servers (like your disk driver or filesystem). It sends a reply to the last client that called it, and immediately puts the server to sleep until the next request arrives.[[1](https://news.ycombinator.com/item?id=33078441)]
|
||||
3. **`Yield()`/`Thread_Ctrl()`**
|
||||
- **What it does:**Allows a thread to voluntarily give up its CPU time slice, or allows a root task to spawn/kill threads.
|
||||
4. **`ipc_send(endpoint, message_buffer)`(Asynchronous Send)**
|
||||
- **What it does:**Posts a small payload to an endpoint's bounded queue and returns *without* blocking — no rendezvous, no reply. The receiver picks it up through the same `IPC_ReplyWait`, as a buffered message. It is the async counterpart of `IPC_Call`, for one-to-many broadcasts where a synchronous rendezvous would let one dead or slow receiver hang the sender. The [input service](../device-driver-development-guide/input.md) — keyboard-event fan-out — is its first user. A full queue drops the oldest message (a buffered message is discrete data, unlike a coalescing interrupt notification).
|
||||
|
||||
* * * * *
|
||||
|
||||
Hardware Implementation: x86_64 vs. aarch64
|
||||
|
||||
Because you are targeting both platforms, you must design a clean**Architecture Abstraction Layer (AAL)**. Each architecture uses completely different assembly instructions, CPU registers, and privilege levels to jump from user space (Ring 3 / EL0) to kernel space (Ring 0 / EL1).[[1](https://android.googlesource.com/kernel/common/+/84d3e59750bbd/arch/arm64/Kconfig),[2](https://blog.codingconfessions.com/p/making-system-calls-in-x86-64-assembly),[3](https://alex.dzyoba.com/blog/os-segmentation/),[4](https://dev.to/ripan030/linux-kernel-interrupt-handling-part-2-fe1)]
|
||||
|
||||
Here is how you will map your bare minimum system calls on both architectures:
|
||||
|
||||
1\. x86_64 Implementation
|
||||
|
||||
On 64-bit Intel and AMD processors, you ignore the old`int 0x80`software interrupts. Instead, you use the high-performance`syscall`and`sysret`instructions.[[1](https://alex.dzyoba.com/blog/os-segmentation/)]
|
||||
|
||||
- **The Trap:**The user-space program executes the`syscall`instruction.[[1](https://namastedev.com/blog/kernel-vs-user-space-2/)]
|
||||
- **The Registers:**The hardware automatically moves the instruction pointer, but you must pass your arguments in specific registers. A common microkernel convention mimics the System V AMD64 ABI:
|
||||
- `rax`: System Call ID (e.g.,`0`for IPC_Call,`1`for IPC_ReplyWait)
|
||||
- `rdi`: Argument 1 (Endpoint ID / Destination)
|
||||
- `rsi`: Argument 2 (Pointer to the message payload buffer)
|
||||
- `rdx`: Argument 3 (Size of the message)[[1](https://dev.to/kaamkiya/hello-world-in-assembly-x86-64-2kb8)]
|
||||
- **Kernel Setup:**Your kernel must configure the Model Specific Registers (MSRs)---specifically`IA32_STAR`and`IA32_LSTAR`---during boot to point to your kernel's system call entry assembly code.[[1](https://nfil.dev/kernel/rust/coding/rust-kernel-to-userspace-and-back/),[2](https://johannst.github.io/notes/arch/x86_64.html)]
|
||||
|
||||
2\. aarch64 (ARM 64-bit) Implementation
|
||||
|
||||
On ARMv8-A and ARMv9-A architectures, privilege levels are called Exception Levels. User space runs at**EL0**, and your microkernel runs at**EL1**.[[1](https://community.nxp.com/pwmxy87654/attachments/pwmxy87654/imx-processors/183079/1/AN12212.pdf),[2](https://developer.arm.com/-/media/Arm%20Developer%20Community/PDF/Learn%20the%20Architecture/Exception%20model.pdf?revision=a62f2bf2-b08a-4a4f-8cbe-38c67ddf4434),[3](https://people.kernel.org/linusw/),[4](https://drewdevault.com/blog/Helios-aarch64/)]
|
||||
|
||||
- **The Trap:**The user-space program executes the`svc #0`(Supervisor Call) instruction.[[1](https://hackmd.io/@xlYUTygoRkyuQQlwXuWDWQ/SJWQuIsIZe)]
|
||||
- **The Registers:**ARM provides a clean, plentiful register set. You typically pass your arguments in the standard parameter registers:
|
||||
- `x0`: System Call ID
|
||||
- `x1`: Argument 1 (Endpoint ID / Destination)
|
||||
- `x2`: Argument 2 (Pointer to message payload buffer)
|
||||
- `x3`: Argument 3 (Size of the message)[[1](https://github.com/lelegard/arm-cpusysregs),[2](https://medium.com/@vincentcorbee/http-server-in-arm64-assembly-apple-silicon-m1-077a55bbe9ca)]
|
||||
- **Kernel Setup:**Your kernel must set up an Exception Vector Table and write its base address to the`VBAR_EL1`register. When`svc`is executed, the CPU jumps to the synchronous exception offset in that table.[[1](https://dev.to/ripan030/linux-kernel-interrupt-handling-part-2-fe1),[2](https://eastrivervillage.com/blog/archive/2018/06/),[3](https://www.wadixtech.com/blog/armv8-a-exception-levels-el0-to-el3)]
|
||||
|
||||
* * * * *
|
||||
|
||||
Managing the Payload Challenge
|
||||
|
||||
Because it is a microkernel, performance lives or dies by how fast your`IPC_Call`can move data from Client to Server. You have two minimal choices for handling the`message_buffer`pointer:[[1](https://anazimzada2020.medium.com/microkernel-architectural-pattern-5e4e9184170e)]
|
||||
|
||||
- **The Copy Method (Simplest to start):**Your kernel pauses the client, reads the data from the client's memory space, switches page tables to the server, and copies the data into the server's buffer.
|
||||
- **The Shared Memory Method (Fastest):**The kernel sets up a temporary, shared virtual memory page between the client and server. The client writes to it, calls`syscall`/`svc`, and the server reads it instantly without the kernel copying any bytes
|
||||
@@ -0,0 +1,131 @@
|
||||
# system.img — the boot capsule
|
||||
|
||||
## What it is
|
||||
|
||||
`boot/system.img` is the **boot capsule**: every bundled user binary — init, the
|
||||
services, the drivers, the test programs — packed into **one file** on the boot
|
||||
volume. It is not a filesystem image and it is not compressed; it is exactly the
|
||||
kernel's **initial-ramdisk wire format** (`system/initial-ramdisk.zig`, format
|
||||
v2), written to disk ahead of time. The EFI loader reads it in a single
|
||||
sequential pass and hands the bytes to the kernel unmodified.
|
||||
|
||||
The capsule is a *performance artifact*, not a source of truth. The boot
|
||||
volume's `/system` and `/test` file trees remain the canonical layout (see
|
||||
[danos-file-system-hierarchy-FSH.md](../file-system-development/danos-file-system-hierarchy-FSH.md));
|
||||
the capsule is a pre-baked snapshot of the same binaries, derived from the same
|
||||
build graph, so the running system is identical whether the loader read the
|
||||
capsule or walked the tree.
|
||||
|
||||
## Why it exists
|
||||
|
||||
Firmware file I/O has exactly one fast shape: **one open + one sequential
|
||||
read**. Everything else is a lottery. Loading the system per-file — dozens of
|
||||
opens, seeks, and short reads through the firmware's FAT driver — measured
|
||||
**minutes** on real hardware, against milliseconds in QEMU/OVMF. Packing the
|
||||
binaries into a single file turns the whole of user space into the shape
|
||||
firmware is good at.
|
||||
|
||||
Because the capsule already *is* the ramdisk wire format, the loader doesn't
|
||||
even repack it: `loadCapsule` (`boot/efi.zig`) validates the magic and passes
|
||||
the buffer straight through as `BootInformation.initial_ramdisk_base`/`len`.
|
||||
|
||||
## The format
|
||||
|
||||
The container is deliberately trivial — danos owns both producer and consumer,
|
||||
so it need be no fancier. Little-endian throughout:
|
||||
|
||||
```
|
||||
Header magic: u32 = "DNR2" (0x32524E44), count: u32
|
||||
Entry × count name: [64]u8 (NUL-padded FHS path), offset: u64, len: u64
|
||||
blobs... each entry's file bytes, at its offset within the image
|
||||
```
|
||||
|
||||
- **Names are full FHS paths** (`/system/services/init`), not basenames — that
|
||||
is what "v2" means. The 64-byte capacity matches `abi.maximum_process_name`,
|
||||
so a task named after its binary path is never truncated. Paths longer than
|
||||
63 bytes are a build error (`pack-system-image.py` rejects them).
|
||||
- **The v1 magic (`"DNRD"`, basename entries) is rejected**, not tolerated: a
|
||||
stale image should fail loudly at `Reader.init`, not misparse names.
|
||||
- `initial_ramdisk.Reader` is the one validated view over the bytes — magic
|
||||
check, table bounds, per-blob bounds — used by the kernel and shared with the
|
||||
loader. `Reader.find` resolves a binary by exact path first, then by unique
|
||||
basename, ASCII case-insensitively (the entries come from a FAT volume, whose
|
||||
name lookups are case-insensitive by definition).
|
||||
|
||||
## How it is built
|
||||
|
||||
`build.zig` maintains one `bundled` list — every user binary and its FHS home.
|
||||
Three artifacts are derived from that same list, in the same build graph, so
|
||||
they cannot drift apart:
|
||||
|
||||
1. **The tree**: each binary installed at its FHS path (`zig-out/system/...`
|
||||
and `zig-out/test/...`, mirrored onto the FAT boot volume by
|
||||
`tools/make-fat-image.py`).
|
||||
2. **The manifest** (`system/manifest`): the FHS path of every bundled binary,
|
||||
one per line — the loader's per-file fallback input.
|
||||
3. **The capsule**: `tools/pack-system-image.py` packs the same binaries into
|
||||
the v2 container, installed at `zig-out/boot/system.img` and placed on the
|
||||
boot volume at `boot/system.img`.
|
||||
|
||||
Note what the capsule does *not* contain: the kernel (`system/kernel` is loaded
|
||||
separately by `loadKernel`, as an ELF) and the EFI loader itself. It is user
|
||||
space only.
|
||||
|
||||
## How it is loaded
|
||||
|
||||
`loadSystemTree` (`boot/efi.zig`) tries three strategies, most portable first —
|
||||
the running system cannot tell which one ran, because all three produce the
|
||||
same in-RAM ramdisk image:
|
||||
|
||||
1. **The capsule** — open `boot\system.img`, read it whole, check the magic,
|
||||
hand it over as-is. The normal path on any build-produced volume.
|
||||
2. **The manifest** — read `system\manifest` and open each listed path *by
|
||||
name*. FAT name lookup is case-insensitive and firmware-portable, unlike
|
||||
directory enumeration. The loader assembles the v2 image in RAM itself.
|
||||
3. **The tree walk** — enumerate `/system` and `/test` recursively (`/test`
|
||||
is optional: a stick without fixtures still boots). Last resort for
|
||||
hand-assembled sticks with neither file: some firmware FAT drivers return
|
||||
bare 8.3 names uppercase from enumeration, which is why this is the
|
||||
fallback and not the primary path.
|
||||
|
||||
All three are best-effort: a **kernel-only volume still boots** — the kernel
|
||||
just has no user binaries to spawn and reports the absence.
|
||||
|
||||
One operational consequence of the ordering: the capsule *shadows* the tree.
|
||||
If you hand-edit binaries on a stick that also carries a `boot/system.img`,
|
||||
your edits are invisible — the loader boots the capsule's snapshot. Delete
|
||||
`boot/system.img` from the volume to force the manifest/tree path.
|
||||
|
||||
## What the kernel does with it
|
||||
|
||||
The loader records the image's physical base and length in `BootInformation`;
|
||||
the kernel (`kernel.zig`) then publishes the same bytes twice, to two
|
||||
consumers:
|
||||
|
||||
- **The process layer** (`process.zig`): `system_spawn` looks binaries up in
|
||||
the ramdisk via `Reader.find` — exact FHS path, or unique basename for
|
||||
pre-path callers — and loads them as fresh ring-3 processes. The stored path
|
||||
becomes the task's name.
|
||||
- **The VFS root** (`vfs.zig`, `setInitialRamdisk`): the image is mounted as
|
||||
kernel-backed, read-only mounts — one per top-level tree named by the entry
|
||||
paths, so `/system` and, when the fixtures are bundled, `/test`. Directory
|
||||
nodes are derived from the entry paths (the unique parents), so the trees
|
||||
are listable and their files readable over the normal VFS protocol — the
|
||||
FHS boot tree every process sees comes straight out of the capsule bytes.
|
||||
|
||||
The image is never copied after the handoff and never mutated: the initrd is
|
||||
immutable, which is what makes the VFS's node serving lock-free.
|
||||
|
||||
## What it is not
|
||||
|
||||
- **Not `danos-usb.img`.** That is the 64 MiB FAT32 *boot volume* built by
|
||||
`tools/make-fat-image.py` — the thing a machine actually boots, which
|
||||
*contains* `boot/system.img` alongside the loader, kernel, manifest, and
|
||||
tree. See [efi.md](efi.md) and [release-iso.md](release-iso.md).
|
||||
- **Not a mountable filesystem.** No FAT, no block device, no driver — just a
|
||||
header, a table, and concatenated blobs, parsed by ~90 lines of
|
||||
`initial-ramdisk.zig`.
|
||||
- **Not required.** It is the fast path, with two slower equivalents behind
|
||||
it.
|
||||
- **Not a place where state lives.** It is regenerated on every build from the
|
||||
bundled binaries; nothing writes to it, at build time or runtime.
|
||||
@@ -0,0 +1,122 @@
|
||||
# SysV: the kernel's calling convention
|
||||
|
||||
Several places in danos say "the kernel is SysV" — most visibly `system/boot-handoff.zig`:
|
||||
|
||||
```zig
|
||||
pub const kernel_abi: std.builtin.CallingConvention = .{ .x86_64_sysv = .{} };
|
||||
```
|
||||
|
||||
**SysV** is short for the **System V AMD64 ABI**, the calling convention that
|
||||
Unix-like systems (Linux, the BSDs, macOS) use on x86-64. This page explains what
|
||||
that means and why danos has to pin it explicitly.
|
||||
|
||||
## What a calling convention is
|
||||
|
||||
At the machine level there's no language keeping two functions honest when one
|
||||
calls the other — just registers and a stack. So there has to be a shared
|
||||
agreement on the mechanics:
|
||||
|
||||
- which registers carry the **arguments**, and in what order,
|
||||
- where the **return value** goes,
|
||||
- which registers the callee must **preserve** versus may freely clobber,
|
||||
- the **stack alignment** required at a call,
|
||||
- how larger things (structs, floats, varargs) are passed.
|
||||
|
||||
That agreement is the calling convention. Both sides of a call must be compiled to
|
||||
the *same* one, or they read arguments out of the wrong registers and get garbage.
|
||||
("ABI" — Application Binary Interface — is the broader term, also covering type
|
||||
sizes and object-file format; here we mean the calling-convention part.)
|
||||
|
||||
## What SysV specifies (the parts that matter here)
|
||||
|
||||
Integer and pointer arguments go in this register sequence:
|
||||
|
||||
| arg | 1 | 2 | 3 | 4 | 5 | 6 |
|
||||
|------|---------|-----|-----|-----|----|----|
|
||||
| reg | **RDI** | RSI | RDX | RCX | R8 | R9 |
|
||||
|
||||
The return value comes back in **RAX**. RBX, RBP and R12–R15 are **callee-saved**
|
||||
(a function must restore them before returning); the rest are caller-saved. The
|
||||
stack must be 16-byte aligned at a `call`. And there's a **red zone** — 128 bytes
|
||||
below RSP that a function may use as scratch without adjusting RSP.
|
||||
|
||||
The name is historical: it descends from AT&T's *System V* Unix, whose ABI
|
||||
documents this lineage comes from. The modern spec is the "System V Application
|
||||
Binary Interface, AMD64 Architecture Processor Supplement."
|
||||
|
||||
## Why danos pins it: RDI vs RCX
|
||||
|
||||
The reason this is called out explicitly is a clash with the *other* common x86-64
|
||||
convention, **Microsoft x64** — used by Windows **and UEFI** — where the first
|
||||
argument arrives in **RCX**, not RDI.
|
||||
|
||||
danos's two binaries default to different conventions:
|
||||
|
||||
- `boot/efi.zig` is built for the UEFI target, so its default C convention is
|
||||
Microsoft x64 (first argument → RCX).
|
||||
- The kernel is freestanding, so its convention is SysV (first argument → RDI).
|
||||
|
||||
When the loader jumps to the kernel passing the `BootInformation` pointer, both sides have
|
||||
to agree *which register that pointer lands in*. Left to their defaults, the loader
|
||||
would place it in RCX while the kernel looked in RDI — and the kernel would read
|
||||
garbage. So both sides reference the same `boot_handoff.kernel_abi` (SysV): the loader's
|
||||
function-pointer type and the kernel's `_start` both carry
|
||||
`callconv(boot_handoff.kernel_abi)`, and the pointer reliably arrives in RDI. That is the
|
||||
whole reason `kernel_abi` lives in the shared contract — see [efi.md](efi.md) for
|
||||
the handoff it governs.
|
||||
|
||||
## The process-entry stack (argc/argv)
|
||||
|
||||
The SysV ABI also fixes what a *fresh process* finds on its stack — and danos
|
||||
follows it, so its own runtime and any future C libc read arguments the same way.
|
||||
At the first user instruction, `rsp` is 16-byte aligned and points at (addresses
|
||||
growing upward):
|
||||
|
||||
```
|
||||
rsp → argc u64
|
||||
argv[0] … argv[argc-1] pointers into the strings area below
|
||||
NULL argv terminator
|
||||
NULL envp terminator (no environment yet)
|
||||
{AT_PAGESZ, page size} auxiliary vector
|
||||
{AT_NULL, 0} auxiliary-vector terminator
|
||||
argv string bytes NUL-terminated
|
||||
───────────────────────── stack top (stack_top_virtual)
|
||||
```
|
||||
|
||||
The kernel builds this block at the top of the process's stack — 8 pages (32 KiB,
|
||||
`parameters.user_stack_pages`) mapped RW+NX below a fixed top, with the page below
|
||||
them left unmapped as a **guard**, so a stack overflow faults (killing only that
|
||||
process) instead of silently corrupting the image
|
||||
(`buildEntryStack` in `system/kernel/process.zig`); `argv[0]` is always the path
|
||||
or initial-ramdisk name the process was spawned as, and `system_spawn`'s optional
|
||||
argument blob becomes `argv[1..]`. The runtime's `_start`
|
||||
(`library/kernel/start.zig`) hands the block to `rt_start`, which builds a
|
||||
`process.Init` from it and passes that to the program's `main`
|
||||
(`pub fn main(init: process.Init)`; a parameterless `main()` is also
|
||||
accepted). A C runtime's `crt0` would walk
|
||||
the identical layout unmodified — that's the compatibility being bought. The
|
||||
`args` test proves the round trip.
|
||||
|
||||
## Where else it surfaces
|
||||
|
||||
- **The red zone → `red_zone = false`.** `build.zig` disables the red zone for the
|
||||
kernel. Interrupts push their frame onto the current stack; if the interrupted
|
||||
code was using its 128-byte red zone, that push would stomp it. Turning the red
|
||||
zone off is the standard fix for kernel code — a direct consequence of SysV
|
||||
*having* a red zone. (See [interrupts.md](interrupts.md).)
|
||||
- **`callconv(.c)` == SysV here.** The exception/interrupt dispatcher
|
||||
(`interruptDispatch`) is declared `callconv(.c)`, which resolves to SysV on this
|
||||
target. That's why the assembly stub in `isr.s` moves the `CpuState` pointer into
|
||||
**RDI** before `call`ing it — the same first-argument rule.
|
||||
|
||||
So "the kernel is SysV" means: its functions pass arguments in RDI/RSI/RDX/…,
|
||||
return in RAX, preserve the SysV callee-saved registers, and assume a red zone —
|
||||
and every boundary that calls into the kernel (the loader, the interrupt stubs)
|
||||
has to speak that same convention at the point of the call.
|
||||
|
||||
## A note on other architectures
|
||||
|
||||
This is x86-64-specific. An AArch64 port ([architecture.md](architecture.md)) has its own calling
|
||||
convention (arguments in X0–X7, and so on) — a different ABI entirely. `kernel_abi`
|
||||
would be set per-architecture, but the *principle* is the same: the loader/entry
|
||||
boundary and the kernel must agree on how arguments are passed.
|
||||
@@ -0,0 +1,471 @@
|
||||
# Threading — build plan (`runtime.Thread` over a private thread ABI)
|
||||
|
||||
The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone
|
||||
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
|
||||
[display-v2-plan.md](../device-driver-development-guide/display-v2-plan.md). Read threading.md first for the *why*.
|
||||
|
||||
## Locked decisions (do not relitigate)
|
||||
|
||||
- **`runtime.Thread` mirrors `std.Thread`'s API; the implementation is danos-native.**
|
||||
Not literal `std.Thread` — that would break the [private ABI](syscall.md).
|
||||
- **Threads are a narrow, per-binary opt-in.** Default concurrency stays process + IPC
|
||||
([resilience.md](resilience.md)); only a service that asks is built
|
||||
`single_threaded = false`.
|
||||
- **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an
|
||||
idle core still halts ([halting.md](halting.md)).
|
||||
- **New syscalls are private**: extend [abi.zig](../../system/abi.zig) `SystemCall` after
|
||||
`shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`,
|
||||
`futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number.
|
||||
- **Restart granularity stays the process** — a faulting thread kills its process; the
|
||||
supervisor restarts the process, which respawns its threads.
|
||||
|
||||
## Conventions
|
||||
|
||||
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations,
|
||||
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
|
||||
`addUserBinary` (with the new `threaded` flag where a binary spawns threads) and get
|
||||
packed into the initial-ramdisk; new syscalls extend [abi.zig](../../system/abi.zig)
|
||||
`SystemCall` + a `library/runtime` wrapper; test services live beside the code they
|
||||
exercise and register a `ServiceId` if they must be looked up.
|
||||
|
||||
## How to verify along the way
|
||||
|
||||
**Every gate is serial-checkable — no screenshots** (this plan runs unattended). A
|
||||
thread proves it ran by writing to **shared memory** the parent reads back, and proves
|
||||
parallelism by stamping the **core index** it ran on (like the `smp`/`affinity` cases).
|
||||
|
||||
- `zig build test` — host unit tests (closure packing, mutex state machine, futex
|
||||
wrapper encodings).
|
||||
- `python3 test/qemu_test.py <case>` — boots the kernel in QEMU; asserts on serial
|
||||
markers. Thread cases set `smp: true` (real parallelism) and bump `mem` (they boot
|
||||
the process/scheduler stack); each milestone **adds its case to `CASES`** so its gate
|
||||
is runnable.
|
||||
- **Guardrail every milestone:** the concurrency-sensitive existing cases stay green —
|
||||
`smoke`, `sched`, `priority`, `smp`, `affinity`, `process`, `process-kill`,
|
||||
`supervision`, `fault-recovery`, `vfs-client-death`, `ipc`/`ipc-cap`,
|
||||
`display-service`. A threading change that regresses those is rejected.
|
||||
|
||||
## Unattended execution (the loop contract)
|
||||
|
||||
This plan runs to completion **without human input**. Every design choice is already
|
||||
fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration must:
|
||||
|
||||
1. **Resume** at the first milestone that still has an unchecked `- [ ]`. (All earlier
|
||||
milestones are done — do not revisit them.)
|
||||
2. **Work on a branch.** On the first iteration, branch off the current `main` into a new
|
||||
branch (e.g. `threading-phase2` — Phase 1's `threading` is already merged); never
|
||||
commit to `main` directly. Push that **branch** to `origin` after each milestone (step
|
||||
5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into
|
||||
`main` stays a human step.
|
||||
3. **Implement** every unchecked item in that milestone, including adding its
|
||||
`-Dtest-case` to `CASES` in [test/qemu_test.py](../../test/qemu_test.py) (with
|
||||
`smp: true` / a `mem` bump where noted) so the gate is runnable.
|
||||
4. **Run the gate**: `python3 test/qemu_test.py <case>`, then the full **guardrail
|
||||
set**, then `zig build` (clean) and `zig build test` (green).
|
||||
5. **Decide, do not ask:**
|
||||
- **Green** = the milestone's case prints its stated marker(s) and reports `PASS`,
|
||||
the whole guardrail set passes, `zig build` is clean, and host tests are green.
|
||||
→ tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit`
|
||||
(`threads(M<n>): <summary>`, no `Co-Authored-By` trailer per
|
||||
[coding-standards.md](../coding-standards.md)), then **`git push` the working branch to
|
||||
`origin`** (use `-u` on the first push to set upstream). Continue to the next
|
||||
milestone in the same iteration if budget remains; otherwise let the loop re-fire.
|
||||
- **Red** = anything above fails. Diagnose from the captured serial log
|
||||
(`zig-out/qemu-test/<case>-failed-serial.log`) and fix in place, then re-run — up to
|
||||
**3 fix attempts** for that gate. A concurrency case that fails then passes on a
|
||||
bare re-run is **flaky, not green**: re-run it **twice more** and treat green only
|
||||
if it passes all; otherwise fix the race (a real threading bug), don't paper over
|
||||
it.
|
||||
6. **A genuinely ambiguous fork is not a stop.** Pick the option most consistent with
|
||||
[threading.md](threading.md)'s *Locked decisions*, note the choice in the commit
|
||||
message, and continue. Do not pause for confirmation on in-scope, reversible work —
|
||||
this plan is that authorization.
|
||||
|
||||
**The only stop conditions:**
|
||||
|
||||
- **Done** — every milestone box **in this plan** is checked (M1 through M11), `zig build`
|
||||
clean, the whole `thread-*` suite + guardrail green. Phase 1 (M1–M6) is *already*
|
||||
checked, so do **not** read that as Done: the loop's real work is the first plan section
|
||||
that still has unchecked boxes — Phase 2 (M7–M11). Only stop when M7–M11 are all checked
|
||||
too. Update threading.md's status line, push the final branch state to `origin`, and
|
||||
stop. The branch is on `origin` for review; **merging Phase 2 into `main` is the user's
|
||||
step**, not the loop's.
|
||||
- **Blocked** — a gate is still red after 3 fix attempts, or a step needs something
|
||||
outside the repo (a toolchain change, new hardware, a decision no locked decision
|
||||
covers). Append `> **BLOCKED (M<n>):** <what failed, what was tried, the serial
|
||||
marker missing>` under that milestone, commit **and push** the WIP on the branch, and
|
||||
stop. Do not thrash further and do not silently skip the milestone.
|
||||
|
||||
Nothing else warrants stopping — not "should I proceed?", not "is this right?". The
|
||||
checkboxes + git history are the resumable record; the next iteration picks up from the
|
||||
first unchecked box.
|
||||
|
||||
---
|
||||
|
||||
## M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅
|
||||
|
||||
The one invariant change threads require, landed and proven **before** anything shares
|
||||
an address space. Today address space is 1:1 with a task and teardown destroys it on any user
|
||||
task's exit; make destruction happen on the **last** exit.
|
||||
|
||||
- [x] A refcount keyed by the address-space root, held in `scheduler.zig`
|
||||
(`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the
|
||||
success path, after the slot + stack are secured), all under the big kernel lock.
|
||||
- [x] Both task-teardown paths ([scheduler.zig](../../system/kernel/scheduler.zig):
|
||||
`exitUserLocked` and `destroyTaskLocked`) call `releaseAddressSpace`, which decrements
|
||||
and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test
|
||||
spaces) is destroyed directly, preserving prior behaviour.
|
||||
- [x] `-Dtest-case=address-space-refcount`: spawn and reap several ring-3 processes in sequence
|
||||
and assert (via test-observable `liveAddressSpaceCount`/`addressSpaceDestroyCount`) that the
|
||||
live-space count returns to **baseline** and destructions advance by exactly that
|
||||
many — each space destroyed exactly once, no leak, no double-free. (Refcount
|
||||
observables, not raw frame counts, since kernel stacks are still leaked on exit.)
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py address-space-refcount` passes
|
||||
(`address-space-refcount: spaces released to baseline ok` → `DANOS-TEST-RESULT: PASS`), and the
|
||||
full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `smp`,
|
||||
`affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`,
|
||||
`vfs-client-death`, `ipc`, `ipc-cap`, `display-service`); default `zig build` clean,
|
||||
`zig build test` green. The reframing is invisible until an address space is actually shared.
|
||||
|
||||
## M2 — `thread_spawn` + `thread_exit`: a thread runs in the shared address space ✅
|
||||
|
||||
Spawn only — no join yet. Prove a second task executes in the **caller's** address
|
||||
space and exits cleanly.
|
||||
|
||||
- [x] [abi.zig](../../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in
|
||||
process.zig; `thread_spawn` calls `scheduler.spawnThread` (today, after M3, the
|
||||
handler goes `spawnThreadSupervised` → `scheduler.spawnUserLocked`; shares the caller's
|
||||
address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)`
|
||||
(`terminateCurrent` → `releaseAddressSpace`). The closure pointer is delivered in the new
|
||||
thread's **rdi** via a new `jump_to_user_arg` asm path (`t.user_arg`, 0 for a
|
||||
process) — no naked runtime asm.
|
||||
- [x] `library/runtime/thread.zig` (barrel-exported as `runtime.Thread`): `spawn` maps a
|
||||
stack (`mmap`), heap-allocates the `{args}` closure, and calls
|
||||
`thread_spawn(&Closure.entry, stack_top, closure)`; `Closure.entry` (a plain C-ABI
|
||||
Zig fn, closure in rdi) runs the function and calls `thread_exit`. Stack top is
|
||||
16-aligned-minus-8 for the C entry.
|
||||
- [x] A `threaded` flag on the user-binary recipe (`addThreadedUserBinary` →
|
||||
`single_threaded = false`); `thread-test` is the first opt-in binary.
|
||||
- [x] `-Dtest-case=thread-spawn`: `thread-test` spawns a worker that writes a sentinel to
|
||||
a **shared** global and release-stores `done`; the main thread acquire-polls `done`
|
||||
and asserts the shared global holds the sentinel — proof the worker ran in the same
|
||||
address space.
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py thread-spawn` passes
|
||||
(`thread-test: child ran in shared address space ok` → `DANOS-TEST-RESULT: PASS`); guardrail set
|
||||
16/16 green (incl. `args`/`init`/`process`, which exercise the new `jump_to_user_arg`
|
||||
process path with arg 0) plus `address-space-refcount`; `zig build` clean, `zig build test`
|
||||
green.
|
||||
|
||||
> **Note (deferred to M3+):** the mmap arena is per-*task* (`heap_next`), so two threads
|
||||
> in one address space that both `mmap` would collide. Fine for M2 (only the parent maps, for the
|
||||
> child's stack); make the arena per-address-space and the runtime heap thread-safe alongside the
|
||||
> `Mutex` work (M5).
|
||||
|
||||
## M3 — `join` + `detach` + real parallelism ✅
|
||||
|
||||
- [x] `join` over the existing exit-notification path
|
||||
([process-lifecycle.md](process-lifecycle.md)): `thread_spawn` gained a 4th arg, an
|
||||
`exit_endpoint` handle (resolved + refcounted like `spawnProcessSupervised`, via
|
||||
`spawnThreadSupervised`); `join` blocks in `ipc_reply_wait` on that endpoint until
|
||||
the child-exit notice for its `tid`, then `munmap`s the stack. `detach` relinquishes
|
||||
the join right (its stack is reclaimed at process exit — kernel-reaper reclaim for
|
||||
detached threads is deferred; see note).
|
||||
- [x] `runtime.Thread.join` / `detach`, plus `Thread.currentCore()` (a new `current_core`
|
||||
= 39 syscall) for the parallelism proof. `getCurrentId` deferred to M6 (TLS), where
|
||||
a lighter self-id fits. The closure now rides the **thread's own stack** (not the
|
||||
heap) — private per thread, so spawn/join touch no shared heap.
|
||||
- [x] `-Dtest-case=thread-join` (`smp: 4`): `thread-test` join mode spawns N=4 workers
|
||||
that each do K=100k `@atomicRmw`-increments on a shared counter and stamp the core
|
||||
they ran on; the main thread joins all N and asserts `counter == N*K` **and**
|
||||
`@popCount(cores_seen) > 1` (genuine cross-core parallelism), then a detached worker
|
||||
proves `detach` runs without a join.
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py thread-join` passes (`thread-test: join ok` →
|
||||
`DANOS-TEST-RESULT: PASS`), robust across 4 runs; guardrail 17/17 green (incl. `smp`,
|
||||
`affinity`, `process-kill`, and `args`/`init`/`process` on the exit-endpoint spawn path)
|
||||
plus `address-space-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green.
|
||||
|
||||
> **Note (deferred):** a detached thread's stack is freed only at process exit (not by the
|
||||
> reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And
|
||||
> the runtime heap is still not thread-safe: threads that both allocate concurrently would
|
||||
> race (the thread *machinery* avoids the heap, but worker code sharing an allocator does
|
||||
> not). Both fold into the M5 `Mutex`/allocator work.
|
||||
|
||||
## M4 — Futex: the one blocking primitive ✅
|
||||
|
||||
- [x] [abi.zig](../../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a
|
||||
`.blocked` task tagged with `Task.futex_addr` (no queue linkage);
|
||||
`futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock,
|
||||
parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr,
|
||||
count)` scans the task table and readies up to `count` matching waiters (same
|
||||
address space). No spinning — a parked waiter leaves its core free to `hlt`. A
|
||||
timed wait also sets `wake_at`, so the timer's `wakeExpired` wakes it; `futex_addr`
|
||||
staying non-zero (only `futex_wake` clears it) is how the waiter tells timeout from
|
||||
a real wake.
|
||||
- [x] `runtime.Thread.Futex` (`wait` / `timedWait` / `wake`) over the syscall wrappers.
|
||||
- [x] `-Dtest-case=thread-futex` (`smp: 4`): a waiter thread prints `waiting` and
|
||||
`futex_wait`s on a word; the main thread publishes it, prints `waking`, and
|
||||
`futex_wake`s; the waiter prints `woke`. Then a `timedWait` on an unwoken word
|
||||
reports `error.Timeout`.
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py thread-futex` passes, robust across 3 runs —
|
||||
the case's **ordered** regex asserts `waiting → waking → woke → PASS` on the serial
|
||||
stream (the handoff proof), and `thread-futex: timeout ok` confirms the timeout.
|
||||
Guardrail 18/18 green (incl. `sleep`/`event`/`ipc` blocking paths) + `address-space-refcount`,
|
||||
`thread-spawn`, `thread-join`; `zig build` clean, `zig build test` green.
|
||||
|
||||
> **Note:** the kernel test checks only the freshest verdict marker via `bufferHas` (the
|
||||
> in-memory log ring buffer evicts older lines); ordering is asserted against the full
|
||||
> serial stream by the qemu regex instead.
|
||||
|
||||
## M5 — `Mutex` + `Condition` + `Semaphore` ✅
|
||||
|
||||
- [x] `runtime.Thread.Mutex` (three-state futex mutex: CAS fast path, `futex_wait`/`wake`
|
||||
slow path), `Condition` (`wait`/`timedWait`/`signal`/`broadcast`, a futex sequence
|
||||
counter), `Semaphore` (permits over `Mutex`+`Condition`) — the same state machines
|
||||
`std.Thread` uses, ported onto our `Futex`.
|
||||
- [x] `-Dtest-case=thread-mutex` (`smp: 4`): a bounded producer/consumer — 2 producers +
|
||||
2 consumers over one `Mutex` and two `Condition`s move N=2000 unique items through
|
||||
an 8-slot ring; the consumed checksum and tally match exactly (no lost/duplicated
|
||||
item, no overrun) under real cross-core contention. The small ring forces producers
|
||||
to block on full and consumers on empty, exercising `Condition.wait`.
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py thread-mutex` passes (`thread-mutex: ok` →
|
||||
`DANOS-TEST-RESULT: PASS`), robust across 3 runs; guardrail 17/17 green (incl.
|
||||
`sleep`/`event`/`ipc`) + all M1–M4 thread cases; `zig build` clean, `zig build test`
|
||||
green.
|
||||
|
||||
> **Deferred (with rationale):**
|
||||
> - **`join` → futex completion word** — the exit-endpoint join (M3) is correct and
|
||||
> tested. A futex-completion join needs the *kernel* to clear+wake a word after the
|
||||
> thread is fully off its stack (a CLONE_CHILD_CLEARTID-style mechanism); doing it in
|
||||
> the thread's own trampoline would let `join` `munmap` the stack while the thread still
|
||||
> runs on it (use-after-free). Left on the exit-endpoint path; the kernel clear-on-exit
|
||||
> is a later, separate refinement.
|
||||
> - **Host unit tests for the state machines** — `Mutex`/`Condition` bottom out in the
|
||||
> `futex_*` syscalls, unavailable on the host without a mockable `Futex` seam. The QEMU
|
||||
> `thread-mutex` gate exercises them under real concurrency instead; a host-side mock is
|
||||
> future work.
|
||||
|
||||
## M6 — `getCurrentId`, docs, and CI wiring ✅
|
||||
|
||||
- [x] `getCurrentId` via a small `thread_self = 42` syscall (`runtime.Thread.getCurrentId`
|
||||
returns the kernel task id). **Per-thread `threadlocal` TLS is deferred** — no
|
||||
consumer needs it, and it would require context-switching the thread pointer per task
|
||||
(real kernel + per-switch cost) for an unused feature; threaded binaries have run fine
|
||||
without it through M2–M5. threading.md's TLS reasoning already scoped it as
|
||||
deferred-unless-needed. When a consumer appears, the shape is: `thread_spawn`
|
||||
allocates a per-thread TLS block, sets the thread pointer, and the context switch saves/
|
||||
restores it.
|
||||
- [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same
|
||||
`Futex`/`Mutex`/`Condition` when wanted.
|
||||
- [x] All `thread-*` cases wired into [test/qemu_test.py](../../test/qemu_test.py)
|
||||
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md
|
||||
status updated to **built**; the worked example is threading.md's win-condition.
|
||||
- [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main
|
||||
thread confirms all three ids are non-zero and distinct — each thread has its own
|
||||
kernel identity. (Renamed from `thread-tls`, which implied `threadlocal`.)
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py thread-id` passes; the whole `thread-*` suite
|
||||
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`) plus the full guardrail set pass; default
|
||||
`zig build` clean, `zig build test` green.
|
||||
|
||||
---
|
||||
|
||||
## Status
|
||||
|
||||
**Phase 1 (M1–M6): built.** danos has `runtime.Thread` — `spawn`/`join`/`detach`,
|
||||
cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private
|
||||
thread ABI behind the runtime.
|
||||
|
||||
**Phase 2 (M7–M11): built.** Thread-safe allocation (M7), a task reaper that reclaims dead
|
||||
tasks' kernel stacks (M8), endpoint-free `thread_join` (M9), the per-thread thread pointer (M10),
|
||||
and `RwLock`/`WaitGroup` + host-testable sync (M11). Two things stay deferred by design
|
||||
(no consumer): the Zig `threadlocal` *compiler* layer (M10) and detached-thread user-stack
|
||||
reclaim (M9) — both noted in place.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — hardening (M7–M11)
|
||||
|
||||
The organising principle, so Phase 2 reinforces danos's goals rather than eroding them:
|
||||
|
||||
- **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks,
|
||||
and futex words live in the process's **address space**, and the kernel's per-process
|
||||
state is keyed by the address-space root — so the M1 refcount + `destroyAddressSpace` already
|
||||
free all of it when the last thread exits. A crashed or killed threaded process leaves
|
||||
**nothing** behind. Phase 2 closes the one thing that is *not* address-space-owned — the
|
||||
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
|
||||
[resilience](resilience.md) restart guarantee, extended to threads.
|
||||
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
|
||||
restores the thread pointer, and reaps dead tasks; the runtime decides allocation, TLS layout,
|
||||
and lock algorithms. Every new kernel entry stays a private syscall behind the runtime
|
||||
([syscall.md](syscall.md)) — the ABI stays renumberable.
|
||||
- **The process is still the isolation and restart boundary.** Threads share fate within
|
||||
one process; Phase 2 never adds a way for one process to reach into another (the
|
||||
cross-process futex stays explicitly out of scope, below).
|
||||
|
||||
### M7 — Thread-safe allocation (the correctness gap) ✅
|
||||
|
||||
Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two
|
||||
threads in one process that both allocate corrupt each other. The thread *machinery*
|
||||
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
|
||||
multi-threaded code would hit it. Closed it:
|
||||
|
||||
- [x] **Kernel — per-address-space mmap arena.** Grew M1's `address_space_refs` entry into the
|
||||
per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`);
|
||||
`scheduler.addressSpaceMmapNextPtr`/`addressSpaceDeviceMapNextPtr` expose them. `systemMmap`
|
||||
reserves a disjoint range under a *brief* lock, then maps **per page** under a
|
||||
short-held lock — not the whole grant — because the big lock is held with interrupts
|
||||
disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed
|
||||
the `affinity` scenario out mid-bring-up). Freed at refcount zero, so the cursors
|
||||
vanish with the process.
|
||||
- [x] **Runtime — thread-safe heap.** The allocator's two free-list mutators
|
||||
(`rawAlloc`/`rawFree`) take a `Thread.Mutex`, gated on
|
||||
`!@import("builtin").single_threaded` so single-threaded binaries compile it out and
|
||||
pay nothing. Uncontended acquisition is a single CAS (no syscall).
|
||||
- [x] `-Dtest-case=thread-alloc` (`smp: 4`): 4 threads each do 500 `alloc`/fill/verify/
|
||||
`free` cycles of varied sizes; each block is filled with a per-thread pattern and
|
||||
verified before free, so any overlap between concurrent allocations is caught.
|
||||
|
||||
**Gate (met):** `thread-alloc` passes (3× non-flaky); full guardrail 23/23 green,
|
||||
`zig build`/`zig build test` clean.
|
||||
|
||||
> **Also fixed here:** the `affinity` guardrail's fixed-count busy-loop (`while (spins <
|
||||
> 3e9)`) had codegen-dependent wall-time — adding a function to `tests.zig` flipped how
|
||||
> the optimiser compiled it, swinging affinity from ~4 s to ~63 s and timing it out.
|
||||
> Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is
|
||||
> independent of unrelated code changes.
|
||||
|
||||
### M8 — The task reaper (cleanup + resilience) ✅
|
||||
|
||||
A dead task's **kernel** stack was leaked ("no reaper yet") — every process *and* thread
|
||||
death lost one, so a crash loop bled kernel memory. The reaper fixes it and serves the
|
||||
[resilience](resilience.md) restart goal directly:
|
||||
|
||||
- [x] A dying task cannot free the kernel stack it runs on, so `exit()`/`exitUserLocked`
|
||||
record it in a **per-core `reap_after_switch` slot** and switch away; the task that
|
||||
resumes on that core frees the stack in `switchTo`'s tail (it's on its own stack, the
|
||||
big lock is still held so the slot can't have been reused). A **tick-time drain**
|
||||
(`reapKillPendingLocked`) is the safety net for the case where the next task is
|
||||
*fresh* (enters via the trampoline, bypassing `switchTo`'s tail). A task killed while
|
||||
*not* running is freed immediately in `destroyTaskLocked`. A `live_stack_bytes`
|
||||
counter is the observable. *(Detached-thread user-stack reclaim moves to M9, which
|
||||
adds the joinable/detached flag.)*
|
||||
- [x] `-Dtest-case=task-reap` (`smp: 4`): spawn and kill 12 processes; poll the
|
||||
test-observable `scheduler.liveStackBytes()` until it returns to **baseline** (a
|
||||
correct reaper gets there in a few ms; a genuine leak times out) — every kernel
|
||||
stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered.
|
||||
|
||||
**Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`,
|
||||
`supervision`, `process-kill`, `address-space-refcount`, `smp`, `affinity` all still green (24/24
|
||||
full guardrail); `zig build`/`zig build test` clean.
|
||||
|
||||
> **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap
|
||||
> first read the `pc` **parameter**, but a task that migrated cores carries a *stale* `pc`
|
||||
> in its saved `switchTo` frame — so it read the wrong core's slot and freed a live stack
|
||||
> (a #GP under SMP). Fixed to re-fetch `thisCpu()` after the switch (the switch only swaps
|
||||
> stacks on the current core).
|
||||
|
||||
### M9 — Futex-completion join (retire the per-thread endpoint)
|
||||
|
||||
With the reaper (M8) able to act *after* a thread is fully off its stack, migrate `join`
|
||||
to the std shape and drop M3's per-thread exit endpoint:
|
||||
|
||||
- [x] A **`thread_join(tid)` syscall** (not a user futex word): it blocks the caller until
|
||||
the task with id `tid` exits, and the exit paths call `wakeJoinersLocked`. `join`
|
||||
only reclaims the joined thread's **user** stack, which the thread vacates the moment
|
||||
it enters the kernel to exit — so waking at *exit* time (not reap time) is safe, and
|
||||
no reaper/address-space juggling or user-memory write is needed. This is equally
|
||||
std-shaped (like `pthread_join`) and much simpler/safer than the planned
|
||||
reaper-written completion word. The runtime no longer passes `thread_spawn` an exit
|
||||
endpoint (it passes `no_cap`; the kernel's 4th `exit_endpoint` arg remains and is
|
||||
still honored); the runtime's per-thread IPC endpoint is gone.
|
||||
- [x] `thread-join` passes on the new path, and its join mode now runs **40 spawn+join
|
||||
cycles** — under the old per-thread-endpoint scheme those leaked handles would
|
||||
exhaust the 16-slot handle table; here they all succeed, proving join is endpoint-free.
|
||||
|
||||
**Gate (met):** `thread-join` passes (3× isolated) on the `thread_join` path; full
|
||||
guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-reap`);
|
||||
`zig build`/`zig build test` clean.
|
||||
|
||||
> **Reaper hardened here (fixes an M8 flake).** M8's single per-core reap slot could be
|
||||
> *overwritten* by a second death on that core before the first drained (a fresh-task/SMP
|
||||
> timing window) — an intermittent one-stack leak (`task-reap` flaked ~20%). Replaced it
|
||||
> with a per-core reap **list** plus a `.reaping` task state so a pending slot can't be
|
||||
> reused before its stack is freed. `task-reap` now 11/11 isolated + 2× in the batch.
|
||||
|
||||
> **Deferred:** detached-thread **user-stack** reclaim (still freed at process exit, as in
|
||||
> M3). Doing it in the reaper needs the saved address space + stack range and a
|
||||
> translate/unmap in a not-currently-loaded address space — real complexity for a bounded leak.
|
||||
> A follow-up when a consumer needs it.
|
||||
|
||||
### M10 — Per-thread TLS: the thread-pointer mechanism ✅
|
||||
|
||||
Give each thread its own thread pointer and private TLS storage — the foundation
|
||||
self-hosting Zig ([zig-self-hosting.md](../zig-self-hosting.md)) will build `threadlocal` on.
|
||||
|
||||
- [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch
|
||||
**only when it changes** (the same conditional-load discipline as CR3;
|
||||
`architecture.setThreadPointer` → `wrmsr IA32_FS_BASE` on x86_64). A
|
||||
`set_thread_pointer(addr)` = 44 syscall sets the caller's `thread_pointer` and loads it
|
||||
now. The kernel never touches FS, so there is no swapgs complication.
|
||||
- [x] **Runtime** lays a small per-thread TLS block at the top of each thread's stack
|
||||
(self-pointer at `%fs:0` + scratch slots) and the thread trampoline calls
|
||||
`set_thread_pointer` before any user code — so every spawned thread has a private,
|
||||
switch-stable thread pointer. Reclaimed with the stack.
|
||||
- [x] `-Dtest-case=thread-tls` (`smp: 4`): two threads each write a unique marker to their
|
||||
own `%fs:8` slot and — after both have written — read it back; a shared (non-per-thread)
|
||||
FS base would clobber one and cause cross-talk. Both read their own marker → pass.
|
||||
|
||||
**Gate (met):** `thread-tls` passes (3×); full guardrail 25/25 (the switch-time thread-pointer
|
||||
restore touches every context switch); `zig build`/`zig build test` clean.
|
||||
|
||||
> **Deferred: the Zig `threadlocal` *compiler* layer.** Real `threadlocal` variables need
|
||||
> the ELF **variant-II TLS** surface — `.tdata`/`.tbss` sections + a `PT_TLS` program header
|
||||
> in `user.ld`, a runtime that copies the template with exact negative-offset layout, and
|
||||
> the `.large`-code-model TLS section names — a high-uncertainty lift for a feature with
|
||||
> **no consumer today** (threading.md scopes it "only if a consumer needs it"). What lands
|
||||
> here is the load-bearing piece — the per-thread thread pointer, context-switched — so adding the
|
||||
> compiler layer later is purely runtime+linker work on top, no kernel change. `getCurrentId`
|
||||
> stays the `thread_self` syscall (M6) rather than an fs self-slot (which would need the
|
||||
> main thread's TLS set up in `_start` too).
|
||||
|
||||
**Gate:** `thread-tls` passes; full `thread-*` suite + guardrail green.
|
||||
|
||||
### M11 — `RwLock`, `WaitGroup`, and host-testable sync ✅
|
||||
|
||||
- [x] `runtime.Thread.RwLock` (reader-preferring: `>0` readers / `-1` writer / `0` free,
|
||||
with `lock`/`tryLock`/`unlock` + `lockShared`/`tryLockShared`/`unlockShared`) and
|
||||
`WaitGroup` (`start`/`finish`/`wait`), both on the existing `Mutex`/`Condition`.
|
||||
- [x] A compile-time `Futex` seam gated on `builtin.os.tag == .freestanding`: the futex
|
||||
syscalls on danos, a spin+yield mock off-target (Zig 0.16 has no `std.Thread.Futex`;
|
||||
`wake` is a no-op since the state machines re-check). `thread.zig` is wired into
|
||||
`zig build test`, so `Mutex`/`RwLock`/`WaitGroup` run as **host unit tests** with real
|
||||
`std.Thread` threads (`test` blocks only compile under test).
|
||||
- [x] `-Dtest-case=thread-rwlock` (`smp: 4`): 2 writers set both halves of a value under
|
||||
the exclusive lock while 3 readers check the halves match under the shared lock —
|
||||
zero half-write observations across ~150k reads. Host tests cover the Mutex,
|
||||
RwLock, and WaitGroup state machines.
|
||||
|
||||
**Gate (met):** `zig build test` covers the sync primitives (host threads); `thread-rwlock`
|
||||
passes (3×); full Done gate **26/26** (whole `thread-*` suite + guardrail); `zig build`
|
||||
clean.
|
||||
|
||||
---
|
||||
|
||||
## Deferred (explicitly not in this plan)
|
||||
|
||||
- **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a
|
||||
physical-address key so two processes share a futex through a [shared-memory](../device-driver-development-guide/display-v2.md)
|
||||
region. Not needed for intra-process threads.
|
||||
- **Per-thread priorities / affinity distinct from the process** — threads inherit the
|
||||
process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep.
|
||||
- **Per-thread signal delivery** — signals stay process-scoped
|
||||
([process-lifecycle.md](process-lifecycle.md)).
|
||||
- **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more.
|
||||
- **A real `std.Thread` backend** — arrives with self-hosting
|
||||
([zig-self-hosting.md](../zig-self-hosting.md)); it sits on these same primitives, so it
|
||||
swaps the impl under `runtime.Thread`, not the call sites.
|
||||
@@ -0,0 +1,349 @@
|
||||
# Threading: `Thread`, a std-shaped API over a private thread ABI
|
||||
|
||||
A note on danos **threads** — several tasks sharing one address space — provided by a
|
||||
`Thread` type that mirrors the shape of Zig's `std.Thread` while keeping every
|
||||
kernel entry behind the [runtime](../../library/kernel). **Built** (M1–M11, see
|
||||
[threading-plan.md](threading-plan.md)): `spawn`/`join`/`detach`, cross-core parallelism,
|
||||
a futex, `Mutex`/`Condition`/`Semaphore`/`RwLock`/`WaitGroup`, `getCurrentId`/`currentCore`,
|
||||
per-thread thread-pointer TLS, thread-safe allocation, and a task reaper that reclaims dead
|
||||
tasks' kernel stacks. Deferred by design (no consumer yet): the Zig `threadlocal`
|
||||
*compiler* layer (the per-thread thread pointer is in place, so it's runtime+linker work on top) and
|
||||
detached-thread user-stack reclaim — see the plan's M9/M10 notes. The analysis is against
|
||||
**Zig 0.16** (the pinned toolchain); `std.Thread`'s internals move between releases, so
|
||||
treat upstream shapes as "0.16.x."
|
||||
|
||||
## The win condition
|
||||
|
||||
A danos service can write
|
||||
|
||||
```zig
|
||||
const t = try Thread.spawn(.{}, worker, .{ctx});
|
||||
// ... do other work concurrently ...
|
||||
t.join();
|
||||
```
|
||||
|
||||
and get real parallelism across cores — with `Thread.Mutex`,
|
||||
`Thread.Condition`, and `Thread.Semaphore` available for
|
||||
coordination — **without any code path reaching the kernel except through the
|
||||
runtime**. The call sites read exactly like `std.Thread`, so the day danos becomes a
|
||||
real Zig target (see [self-hosting](#the-self-hosting-endgame)) we swap the
|
||||
implementation underneath, not the API above.
|
||||
|
||||
## Locked decisions (do not relitigate)
|
||||
|
||||
- **We build `Thread`, not literal `std.Thread`.** It mirrors std's *API and
|
||||
features*; the implementation underneath is danos-native. See
|
||||
[Why not literal std.Thread](#why-not-literal-stdthread).
|
||||
- **Threads are a narrow, opt-in capability — not the default concurrency tool.** The
|
||||
default for resilience stays **process + IPC** ([resilience.md](resilience.md),
|
||||
[ipc.md](../device-driver-development-guide/ipc.md)). See [Where threads fit](#where-threads-fit-the-resilience-tension).
|
||||
- **Blocking synchronization is futex-backed, never spin-backed.** Waiters sleep in
|
||||
the kernel so an idle core still halts ([halting.md](halting.md)).
|
||||
- **Per-binary opt-in to multi-threaded codegen.** Only a service that asks for
|
||||
threads is built `single_threaded = false`; the rest stay lean and single-threaded.
|
||||
- **The thread ABI is private.** New syscalls extend [abi.zig](../../system/abi.zig)
|
||||
`SystemCall` and are reached only through `library/kernel` wrappers, exactly like
|
||||
every other danos syscall ([syscall.md](syscall.md)) — numbers stay renumberable.
|
||||
|
||||
## Why not literal `std.Thread`
|
||||
|
||||
danos's ABI invariant is that the **runtime is the sole holder of the syscall ABI**,
|
||||
and that ABI is private and renumberable ([syscall.md](syscall.md) — "unstable
|
||||
private ABI"). That is a security and evolvability asset: no compiled binary can
|
||||
hardcode a syscall number, and the kernel can renumber freely because only the
|
||||
runtime — rebuilt in lockstep — knows the mapping.
|
||||
|
||||
`std.Thread` is incompatible with that invariant on two counts:
|
||||
|
||||
1. **It selects its backend from `builtin.os.tag`, and issues syscalls directly.**
|
||||
danos targets `.os_tag = .freestanding` ([build.zig](../../build.zig)), for which
|
||||
`std.Thread` resolves to an unsupported stub that `@compileError`s. Adding a real
|
||||
backend would either bake danos syscall numbers into std (breaking ABI privacy and
|
||||
renumbering) or fork std to route back through the runtime — a permanent rebase
|
||||
cost that buys nothing the native type doesn't.
|
||||
2. **Our user binaries are built `single_threaded = true`** ([build.zig](../../build.zig)
|
||||
`addUserBinary`), which compiles threading out entirely and makes atomics and TLS
|
||||
single-threaded. Threads need this flipped per binary regardless.
|
||||
|
||||
So we take the *shape* of `std.Thread`, not the *type*. The cost of replicating the
|
||||
surface (spawn/join/Mutex/Condition) is small; the cost of the std type is the ABI
|
||||
invariant.
|
||||
|
||||
## Where threads fit: the resilience tension
|
||||
|
||||
Threads are in genuine tension with a resilience-first microkernel, and it is worth
|
||||
being explicit so we do not reach for them by reflex.
|
||||
|
||||
The reason danos pays for a microkernel is **fault isolation**
|
||||
([resilience.md](resilience.md)): a component corrupts its own address space, faults,
|
||||
and is **restarted** without touching anyone else — because the boundary *is* the
|
||||
address space. Threads deliberately remove that boundary *within* a process:
|
||||
|
||||
- Threads share one address space, so one thread's stray write corrupts them all —
|
||||
there is no isolation **between** threads.
|
||||
- Threads share fate — by contract: a fault in any thread, or a "kill the process"
|
||||
decision, takes down **all** of them, so restartability lives at the process level,
|
||||
not the thread level. (The kernel does not yet enforce this fan-out — see the
|
||||
Lifecycle note under
|
||||
[Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).)
|
||||
- Shared mutable state reintroduces data races — the failure class the
|
||||
isolate-and-message model was chosen to avoid.
|
||||
|
||||
**Therefore:** the default answer to "make X concurrent" stays *another process over
|
||||
IPC* (isolated, independently restartable) or a single event loop with several
|
||||
message sources. Reach for a thread only inside **one** service that needs genuine
|
||||
**shared-memory, low-latency parallelism** and can accept intra-service fate-sharing —
|
||||
e.g. a compositor splitting tile compositing across cores, where per-tile IPC would be
|
||||
too chatty. "Input on one thread, display on another" is *not* that case; it wants two
|
||||
processes. The isolation boundary stays at process granularity.
|
||||
|
||||
## The API surface (mirrors `std.Thread`)
|
||||
|
||||
Lives in `library/kernel/thread.zig`, re-exported as `Thread`.
|
||||
|
||||
```zig
|
||||
pub const Thread = struct {
|
||||
pub const Id = u32; // the kernel task id
|
||||
pub const SpawnConfig = struct {
|
||||
stack_size: usize = default_stack_size, // no allocator: the closure lives at the top of the thread's own stack
|
||||
};
|
||||
pub const SpawnError = error{SystemResources};
|
||||
|
||||
pub fn spawn(config: SpawnConfig, comptime function: anytype, args: anytype) SpawnError!Thread;
|
||||
pub fn join(self: Thread) void; // block until the thread ends, reclaim its stack
|
||||
pub fn detach(self: Thread) void; // give up the right to join; stack reclaimed at process exit
|
||||
pub fn getCurrentId() Id;
|
||||
pub fn currentCore() Id; // danos extension: the calling core's dense index
|
||||
|
||||
pub const Mutex = struct { pub fn lock(*Mutex) void; pub fn tryLock(*Mutex) bool; pub fn unlock(*Mutex) void; };
|
||||
pub const Condition = struct { pub fn wait(*Condition, *Mutex) void; pub fn timedWait(*Condition, *Mutex, u64) error{Timeout}!void; pub fn signal(*Condition) void; pub fn broadcast(*Condition) void; };
|
||||
pub const Semaphore = struct { pub fn wait(*Semaphore) void; pub fn post(*Semaphore) void; };
|
||||
pub const Futex = struct { pub fn wait(*const atomic.Value(u32), u32) void; pub fn timedWait(...) error{Timeout}!void; pub fn wake(*const atomic.Value(u32), u32) void; };
|
||||
// RwLock / ResetEvent / WaitGroup follow the same pattern, added as needed.
|
||||
};
|
||||
```
|
||||
|
||||
Deviations from `std.Thread`, called out honestly:
|
||||
|
||||
- **The thread function's return value is discarded** (as `std.Thread.join` returns
|
||||
`void`). Return data through shared state or a `Semaphore`/`Condition`, not the
|
||||
return.
|
||||
- No `getCpuCount()` (a service rarely needs it) and no `Thread.yield()` — `yield`
|
||||
lives in the `process` module. Instead `currentCore()` exposes the calling core's dense
|
||||
index ([smp.md](smp.md)), used to observe genuine cross-core parallelism.
|
||||
|
||||
## Kernel primitives (new private syscalls)
|
||||
|
||||
Five core entries extend [abi.zig](../../system/abi.zig) `SystemCall` after
|
||||
`shared_memory_physical = 36` (plus small helpers `current_core`, `thread_self`, and
|
||||
`set_thread_pointer`), each with a `library/kernel` wrapper:
|
||||
|
||||
| Syscall | Signature | Purpose |
|
||||
|---|---|---|
|
||||
| `thread_spawn` | `(entry, stack_top, arg, exit_endpoint) -> tid` | create a task sharing the **caller's** address space; the runtime passes `no_cap` for `exit_endpoint` (join is a syscall, not an endpoint) |
|
||||
| `thread_exit` | `()` | end the calling thread; its stack is reclaimed by the joiner's `munmap`, not the kernel |
|
||||
| `thread_join` | `(tid) -> 0` | block until the thread with id `tid` has exited |
|
||||
| `futex_wait` | `(addr, expected, timeout_ns) -> status` | block if `*addr == expected`, until woken or timeout |
|
||||
| `futex_wake` | `(addr, count) -> woken` | wake up to `count` waiters on `addr` |
|
||||
|
||||
Plus one invariant change with no new syscall: **address-space reference counting**.
|
||||
|
||||
## Mechanics
|
||||
|
||||
### Address-space reference counting
|
||||
|
||||
Before this work an address space was 1:1 with a task: `spawnUserLocked` records
|
||||
`address_space` on the Task (as it still does), and teardown did
|
||||
`destroyAddressSpace(t.address_space)` when **any** user task exited
|
||||
([scheduler.zig](../../system/kernel/scheduler.zig)). With threads, several tasks share
|
||||
one `address_space`, so the first to exit would rip the address space out from under its
|
||||
siblings.
|
||||
|
||||
Fix: a small refcount keyed by the address-space root, kept in
|
||||
[scheduler.zig](../../system/kernel/scheduler.zig): `retainAddressSpace` takes a
|
||||
reference for every user task `spawnUserLocked` starts (count 1 on the first take, so
|
||||
a thread sharing the caller's space increments it); task teardown calls
|
||||
`releaseAddressSpace`, which only calls `destroyAddressSpace` at **zero**. All of
|
||||
this is already under the big kernel lock, so no new locking. This is the one piece
|
||||
that must land and be proven before anything shares an address space.
|
||||
|
||||
### `thread_spawn` and the trampoline
|
||||
|
||||
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle
|
||||
values through scratch registers — `startUserTask` reads the entry/stack (and the
|
||||
thread's closure arg, delivered in `rdi` via `jumpToUserArg`) from the Task
|
||||
([scheduler.zig](../../system/kernel/scheduler.zig)). That makes the thread path clean:
|
||||
|
||||
1. The runtime's `spawn` `mmap`s a stack (syscall `4`) and writes the closure —
|
||||
`{ tls_base, args }`, the std "Instance" pattern — at the **top of the new stack
|
||||
itself** (no heap allocation), with a small per-thread TLS block just below it.
|
||||
2. It calls `thread_spawn(entry = &Closure.entry, stack_top, arg = closure_ptr,
|
||||
exit_endpoint = no_cap)`. The kernel calls the same `spawnUserLocked` path with the
|
||||
**caller's address space** (refcount++), `entry`, and `user_sp = stack_top`.
|
||||
3. `Closure.entry` (a small runtime shim) receives the closure pointer in `rdi` — the
|
||||
kernel delivers `arg` as the entry's first C-ABI argument — sets the thread
|
||||
pointer, calls the user function, then calls `thread_exit`.
|
||||
|
||||
Unlike a process start, there is **no** System V argc/argv/auxv block
|
||||
([sysv.md](sysv.md)) — a thread stack carries only the closure and its TLS block.
|
||||
|
||||
### Lifetime: exit, join, detach, stack reclaim
|
||||
|
||||
- **`thread_exit`** (no arguments) marks the task dead. The kernel releases the
|
||||
task's resources, decrements the address-space refcount, and frees the task slot —
|
||||
the user stack is not the kernel's to unmap; the joiner reclaims it.
|
||||
- **`join`** is a dedicated `thread_join(tid)` syscall: the caller blocks in the
|
||||
kernel (`joinThreadLocked`, woken by `wakeJoinersLocked` when the thread exits),
|
||||
then `munmap`s the stack. (The plan staged join over a per-thread `exit_endpoint`
|
||||
first, with a futex `completion` word as a Stage-2 refinement; neither shipped — the
|
||||
dedicated syscall replaced both. `thread_spawn` still accepts an `exit_endpoint`
|
||||
argument, which the runtime passes as `no_cap`.)
|
||||
- **`detach`** relinquishes the join right: no one waits for the thread, and its
|
||||
stack is reclaimed at process exit — kernel-side reclaim of a detached thread's
|
||||
user stack stays deferred (as the intro notes), since `thread_exit` passes no stack
|
||||
range.
|
||||
|
||||
### Futex, and the sync primitives on top
|
||||
|
||||
`futex_wait`/`futex_wake` are the one blocking primitive; `Mutex`, `Condition`, and
|
||||
`Semaphore` are ordinary user-space state machines over an `atomic.Value(u32)` that
|
||||
call the futex wrappers on the slow path — the same construction `std.Thread` uses,
|
||||
so the algorithms port directly.
|
||||
|
||||
Keying: threads share an address space, so a **virtual address within that address space**
|
||||
identifies a futex uniquely; the kernel keys its wait queue by `(address_space_root, virtual_address)`.
|
||||
Keying by the **physical** address instead (translate `virtual_address -> physical_address` on entry) is a
|
||||
deliberate forward door: it lets two *processes* share a futex through an
|
||||
[shared-memory](../device-driver-development-guide/display-v2.md) region later, without changing the API. We start with the
|
||||
private-per-address-space key and note the physical-key upgrade.
|
||||
|
||||
No spinning: a contended lock parks the task in the kernel and the core is free to run
|
||||
other work or `hlt` ([halting.md](halting.md)). This is why futex is a locked
|
||||
decision, not a "maybe later."
|
||||
|
||||
### TLS and `getCurrentId`
|
||||
|
||||
Per-thread thread-pointer TLS is in place (the `threadlocal` *compiler* layer is not —
|
||||
see the intro). Two scoped pieces, as built:
|
||||
|
||||
- **`getCurrentId`** returns the kernel task id via the trivial `thread_self` syscall.
|
||||
- **The thread pointer** is per-thread: `spawn` carves a small TLS block (an `fs:0`
|
||||
self-pointer plus scratch) from the top of the thread's own stack, the trampoline
|
||||
calls `set_thread_pointer` before any user code runs, and the scheduler saves and
|
||||
restores the pointer per task across context switches. Full `threadlocal` support is
|
||||
runtime+linker work on top of this, only if a consumer needs it. Nothing in the core
|
||||
spawn/join/mutex path requires `threadlocal`.
|
||||
|
||||
### Build: multi-threaded codegen, opt-in
|
||||
|
||||
A binary opts in by being added with `addThreadedUserBinary` — as `addUserBinary`,
|
||||
but the shared implementation builds it `single_threaded = false` — so atomics and
|
||||
(later) TLS are real. Threads and atomics are unsound in a `single_threaded` image,
|
||||
so a binary must opt in **before** it may call `Thread.spawn`. Everyone else
|
||||
stays single-threaded and lean.
|
||||
|
||||
## Interaction with the rest of the kernel
|
||||
|
||||
- **Scheduler / SMP** ([scheduling.md](scheduling.md), [smp.md](smp.md)): a thread is
|
||||
just another `Task` with an `address_space` shared with its siblings; the existing
|
||||
per-core ready queues, priorities, and affinity apply unchanged. Threads of one
|
||||
process can run on different cores simultaneously — that is the point.
|
||||
- **Halting** ([halting.md](halting.md)): futex-parked waiters keep the "idle core
|
||||
halts" property intact under lock contention — no busy-wait.
|
||||
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): the contract is that
|
||||
killing a process kills *all* its threads and only then drops the last address-space
|
||||
ref — and the kernel now implements exactly that
|
||||
([shared-fate-plan.md](shared-fate-plan.md)): every death path (`exit` from any
|
||||
thread, a fault, `process_kill` aimed at any member id) fans out through the whole
|
||||
group via a `dying` latch on the address space; the supervisor's one exit
|
||||
notification — badged with the leader — fires only when the last member is gone.
|
||||
A worker's voluntary `thread_exit` stays per-thread; the leader's is refused
|
||||
(`-EPERM`).
|
||||
- **Resilience** ([resilience.md](resilience.md)): by the same contract, a faulting
|
||||
thread kills its whole process (shared fate); the supervisor restarts the
|
||||
**process**, which respawns its threads from a known-good state — restart
|
||||
granularity stays the process. The leader's recorded exit reason carries the fault
|
||||
class even when a worker faulted, so restart policy is unchanged.
|
||||
- **IPC — two consequences threads forced ([ipc.md](../device-driver-development-guide/ipc.md)):**
|
||||
- *Handles do not cross threads.* The handle table lives on the `Task`
|
||||
([scheduler.zig](../../system/kernel/scheduler.zig)), so a handle number is meaningful
|
||||
only to the thread that created it — thread A's endpoint handle `3` is not thread B's.
|
||||
A thread that needs to reach an endpoint another thread owns looks it up
|
||||
(`ipc.lookup(service)`) to install its **own** handle to the same underlying endpoint.
|
||||
This is how the display's mouse-listener thread reaches the compositor loop's endpoint
|
||||
to poke it awake (docs/display.md).
|
||||
- *IPC syscalls that touch shared kernel state now serialize under the big kernel lock.*
|
||||
`create_ipc_endpoint`/`ipc_register`/`ipc_lookup` allocate from the kernel heap and
|
||||
mutate the global service registry, endpoint refcounts, and handle tables. Those paths
|
||||
were unlocked because a single-threaded process could not race itself; a multi-threaded
|
||||
one can, from two cores at once. They now take `sync.enter()` like `call`/`reply_wait`/
|
||||
`send` already did — the kernel heap has no lock of its own (heap.zig: "every kernel
|
||||
entry takes the big kernel lock"), so the big lock is what keeps its callers serialized.
|
||||
|
||||
## Build-out plan (staged, each gate serial-checkable)
|
||||
|
||||
The ordered, `/loop`-runnable milestones live in
|
||||
**[threading-plan.md](threading-plan.md)** (shaped like
|
||||
[display-v2-plan.md](../device-driver-development-guide/display-v2-plan.md)): every milestone lands on its own and ends in
|
||||
a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
|
||||
`zig build test` for host unit tests). The stages below are the shape it expands.
|
||||
|
||||
- **Stage 0 — address-space refcount.** Refcount on the address-space root; teardown destroys
|
||||
at zero. No API yet; nothing shares an address space, so refcount is 1 everywhere.
|
||||
*Gate:* the full QEMU suite stays green (no regression) — proves the reframing is
|
||||
invisible until used.
|
||||
- **Stage 1 — spawn / join / detach.** `thread_spawn` + `thread_exit`, the trampoline,
|
||||
stacks via `mmap`, join over the exit-endpoint (as built, join became the dedicated
|
||||
`thread_join` syscall instead), the `addThreadedUserBinary` build opt-in.
|
||||
*Gate:* two cases as built — `-Dtest-case=thread-spawn`, where a worker thread runs
|
||||
in the caller's address space (a shared-memory write, observed by the main thread),
|
||||
and `-Dtest-case=thread-join`, where N workers each atomically increment a shared
|
||||
counter K times, the parent joins all N and asserts the total is exactly N × K —
|
||||
the join case running multi-core (`smp` 4) to prove real parallelism.
|
||||
- **Stage 2 — blocking synchronization.** `futex_wait`/`futex_wake` + `Futex`,
|
||||
`Mutex`, `Condition`, `Semaphore`; optionally migrate join to a futex completion
|
||||
word. *Gate:* `-Dtest-case=thread-mutex` — a bounded producer/consumer over a
|
||||
`Mutex` + `Condition` moves K items with no lost wakeups and no busy-wait (assert
|
||||
the consumer blocked, e.g. via a low idle tick count).
|
||||
- **Stage 3 — polish.** Per-thread TLS / thread pointer and `threadlocal` (only if a
|
||||
consumer needs it), `RwLock`/`WaitGroup` as demanded, and this doc's cases wired
|
||||
into [test/qemu_test.py](../../test/qemu_test.py).
|
||||
|
||||
## Conventions
|
||||
|
||||
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym
|
||||
abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New syscalls
|
||||
extend [abi.zig](../../system/abi.zig) `SystemCall` + a `library/kernel` wrapper
|
||||
([syscall.md](syscall.md)). `Thread` is a first-class runtime module, the same
|
||||
way `process` ([process-lifecycle.md](process-lifecycle.md)) and `ipc`
|
||||
are — user code never names a syscall.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- **No preemptive user-space signals delivered to a specific thread.** Signals stay
|
||||
process-scoped ([process-lifecycle.md](process-lifecycle.md)).
|
||||
- **No thread priorities distinct from the process.** Threads inherit the process
|
||||
priority; per-thread priority is a later question if it ever earns its keep.
|
||||
- **No cross-process shared-memory futex yet** — the physical-address key leaves the
|
||||
door open, but the first cut is private-per-address-space.
|
||||
- **No `pthread`/POSIX surface.** The API is `std.Thread`-shaped Zig, nothing more.
|
||||
|
||||
## The self-hosting endgame
|
||||
|
||||
When danos becomes a real Zig target and we (eventually) add a danos backend to std
|
||||
([zig-self-hosting.md](../zig-self-hosting.md)), `std.Thread` can sit *on top of* these
|
||||
same kernel primitives — the danos `std.Thread.Impl` would call the very
|
||||
`thread_spawn`/`futex_*` wrappers `Thread` already uses. Because
|
||||
`Thread` was built API-compatible from day one, that transition swaps the
|
||||
implementation, not a single call site. Designing to the std shape now is what makes
|
||||
the later self-hosting lift cheap.
|
||||
|
||||
## Further reading
|
||||
|
||||
- [scheduling.md](scheduling.md), [smp.md](smp.md) — the task model these threads join.
|
||||
- [resilience.md](resilience.md), [vision.md](../vision.md) — why isolation is the default
|
||||
and threads are the exception.
|
||||
- [syscall.md](syscall.md), [ipc.md](../device-driver-development-guide/ipc.md) — the private ABI and the messaging model
|
||||
threads sit beside.
|
||||
- [halting.md](halting.md) — the idle/halt property futex-backed blocking preserves.
|
||||
- [zig-self-hosting.md](../zig-self-hosting.md) — the target this bends toward.
|
||||
@@ -0,0 +1,118 @@
|
||||
# Timers and time
|
||||
|
||||
Two different needs hide under the word "timer", and danos keeps them apart:
|
||||
|
||||
- **Reading the clock** — *what time is it?* A read of a free-running counter.
|
||||
- **Waiting** — *wake me in N milliseconds*, or *notify me when a deadline passes.*
|
||||
|
||||
Both are answered by the **kernel**, because the kernel already owns a timer: it has
|
||||
to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this
|
||||
are built in [device-interrupts.md](../device-driver-development-guide/device-interrupts.md); the scheduler's blocking and
|
||||
wait queues are in [scheduling.md](scheduling.md). This page is about the surface a
|
||||
ring-3 program actually uses, and one deliberate absence: **there is no user-space time
|
||||
service.**
|
||||
|
||||
## Why time is a syscall, not a service
|
||||
|
||||
The tempting microkernel move is to put a timer *driver* in user space and have
|
||||
applications ask it for the time over IPC. For a **monotonic clock that is wrong** —
|
||||
reading `now()` should never cost an IPC round trip. The kernel is already holding the
|
||||
answer: it computes the current time every time it schedules, from the TSC, in a couple
|
||||
of instructions. Surfacing that as a system call is pure mechanism; routing it through a
|
||||
message to another process would be slower *and* redundant, and a device like the HPET
|
||||
(uncacheable MMIO reads) is a particularly bad thing to read on every `now()`.
|
||||
|
||||
This is the same conclusion every serious system reaches: Linux and Zircon read the
|
||||
counter in the vDSO, L4 exposes a clock field in a shared kernel page, seL4 reads the
|
||||
cycle counter directly. None of them make a clock read an IPC. danos makes it a syscall.
|
||||
|
||||
That "from the TSC" hides a portability question, because the TSC is only a valid clock
|
||||
when the CPU guarantees it is *invariant* and when every core's TSC is *synchronized*.
|
||||
danos checks both — the invariant-TSC CPUID bit (`0x80000007` EDX[8], set on Intel and
|
||||
AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET
|
||||
counter when either fails. So `now()` stays accurate on a real Intel box, a real AMD box,
|
||||
and inside a VM alike; only the source behind it differs. The mechanism is in
|
||||
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md).
|
||||
|
||||
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
|
||||
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
|
||||
model; that role now lives in [drivers.md](../device-driver-development-guide/drivers.md), as documentation.) The one place
|
||||
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
|
||||
at the end; it is deliberately not built yet.
|
||||
|
||||
## The three system calls
|
||||
|
||||
Time and waiting are three entries in the small syscall table ([syscall.md](syscall.md)):
|
||||
|
||||
- **`clock` (#23)** → monotonic nanoseconds since boot. It only moves forward. Not
|
||||
wall-clock: no date, no timezone. Backed by `architecture.nanos()` (TSC, scaled with a
|
||||
128-bit intermediate so a long uptime can't overflow) — a few nanoseconds of
|
||||
resolution, and just an `rdtsc` plus a multiply.
|
||||
- **`sleep` (#3)** → block the caller for N milliseconds. The scheduler records a wake
|
||||
deadline and the tick sweep wakes it (`scheduler.sleep`).
|
||||
- **`timer_bind` (#31)** → arm a one-shot timer that, after N milliseconds, posts a
|
||||
**timer notification** to an IPC endpoint. Unlike `sleep` it does **not** block: a
|
||||
service can keep answering messages on the same endpoint while a deadline is pending.
|
||||
This is the timed wait that stop-sequence escalation, hello deadlines, and restart
|
||||
backoff are built from ([process-lifecycle.md](process-lifecycle.md),
|
||||
[device-manager.md](../device-driver-development-guide/device-manager.md)).
|
||||
|
||||
The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space;
|
||||
programs read the TSC through `clock` and get timed wakeups through `sleep`/`timer_bind`,
|
||||
both riding the scheduler tick.
|
||||
|
||||
## `time` — the generic interface
|
||||
|
||||
Applications don't call the syscalls directly; they use `time`
|
||||
(`library/kernel/time.zig`), a thin `Instant`/`Duration` layer over them — an ergonomic
|
||||
front door, not new mechanism.
|
||||
|
||||
```zig
|
||||
const time = @import("time");
|
||||
|
||||
const start = time.now(); // Instant — monotonic
|
||||
doWork();
|
||||
const took = start.elapsed(); // Duration
|
||||
time.sleep(time.Duration.fromMillis(5)); // block ~5 ms
|
||||
|
||||
// A deadline delivered as a notification, so a service keeps serving meanwhile:
|
||||
_ = time.after(endpoint, time.Duration.fromMillis(200));
|
||||
```
|
||||
|
||||
- `Duration` is nanoseconds under the hood, with `fromNanos/fromMicros/fromMillis/
|
||||
fromSeconds` and `asNanos/asMillis`. `ceilMillis` rounds *up* to the kernel's
|
||||
millisecond granularity, so a sub-millisecond `sleep` never rounds down to zero and
|
||||
returns early. All arithmetic saturates rather than wraps.
|
||||
- `Instant` is a point on the monotonic clock: `since`, `elapsed`, `plus`, `reached` —
|
||||
built for deadline loops (`while (!deadline.reached()) …`).
|
||||
- `now()` / `monotonicNanos()` wrap `clock`. `available()` reports whether the clock is
|
||||
calibrated at all (the kernel returns 0 until the TSC frequency is known, so a caller
|
||||
that needs real time can treat 0 as "unavailable" rather than assume it advances).
|
||||
- `sleep(d)` wraps `sleep`; `spin(d)` busy-polls `now()` for the sub-millisecond delays
|
||||
the millisecond tick can't express; `after(endpoint, d)` wraps `timer_bind`.
|
||||
|
||||
The raw wrappers (`clock`, `sleepMillis`, `timerOnce`) and the ergonomic
|
||||
`Instant`/`Duration` layer both live in the `time` module
|
||||
(`library/kernel/time.zig`); the latter is what everyday code uses.
|
||||
|
||||
## Wall-clock time (not built)
|
||||
|
||||
Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and
|
||||
measurement, useless for "what is the date?" Calendar time — a real-time clock, time
|
||||
zones, leap seconds — is genuinely a **user-space** concern, and it *is* the case a time
|
||||
service is for. It would be backed by an **RTC** driver (the CMOS real-time clock), not
|
||||
the HPET, and exposed as a `CLOCK_REALTIME`-style service alongside the monotonic
|
||||
syscall. It is deferred until something needs it; the monotonic clock the kernel already
|
||||
owns covers every current use.
|
||||
|
||||
## Verifying it
|
||||
|
||||
`time`'s `Instant`/`Duration` arithmetic has unit tests that run on the host:
|
||||
|
||||
```
|
||||
$ zig build test # includes library/kernel/time.zig
|
||||
```
|
||||
|
||||
End to end, the proof the clock is real is that it *advances*: read `now()`, `sleep` a
|
||||
`Duration`, read `now()` again, and the second reading is later — the kernel's timer
|
||||
driving a ring-3 program with no service in between.
|
||||
@@ -0,0 +1,216 @@
|
||||
# The vDSO — the public system-call boundary
|
||||
|
||||
> **Status:** design note, not built. The runtime today issues raw `syscall`
|
||||
> instructions from `library/kernel/system-call.zig` using the numbers in
|
||||
> `system/abi.zig`. This note designs the layer that replaces that arrangement:
|
||||
> a **kernel-supplied, C-ABI entry library** mapped into every process — the
|
||||
> only supported way into the kernel — so the raw numbers can stay private,
|
||||
> be renumbered at will, and eventually be randomised per boot.
|
||||
|
||||
## Why: the ABI danos promises, and the one it doesn't
|
||||
|
||||
`system/abi.zig` is the **private** kernel ↔ runtime contract. Its header says
|
||||
so: the numbers are an implementation detail the runtime hides and may
|
||||
renumber, the same split as libSystem over the XNU syscalls on macOS or win32
|
||||
over the NT syscalls on Windows. Linux — with its world-visible, frozen
|
||||
syscall table — is the outlier, not the norm.
|
||||
|
||||
That stance has consequences the moment binaries exist that we don't rebuild
|
||||
ourselves:
|
||||
|
||||
1. **Third-party binaries** (docs/zig-self-hosting.md) must keep working across
|
||||
kernel updates. If they contain raw `syscall` instructions with today's
|
||||
numbers baked in, every renumbering breaks the world — the ABI would be
|
||||
*de facto* public no matter what the header says. Go on macOS made exactly
|
||||
this mistake: it issued XNU syscalls directly instead of going through
|
||||
libSystem, and macOS updates repeatedly broke every Go binary until Go
|
||||
switched to the library like everyone else.
|
||||
2. **Not everything is Zig.** A Rust or C program can't import the danos Zig
|
||||
modules. The public boundary has to be expressible in the one calling
|
||||
convention every language speaks: the C ABI.
|
||||
3. **Randomised syscall numbers** — a hardening option we want open — only
|
||||
work if no user binary anywhere knows a number at build time. The binding
|
||||
must happen at *load time*, from something the kernel controls.
|
||||
|
||||
All three point at the same well-known shape: a **vDSO** (virtual dynamic
|
||||
shared object). The kernel carries a small blob of user-mode code, maps it
|
||||
into every process at spawn, and that blob — not the application — contains
|
||||
the `syscall` instructions. Fuchsia works exactly this way: its vDSO is the
|
||||
*only* kernel entry, version-matched by construction because the kernel itself
|
||||
injects it. Because the kernel and the blob ship as one artifact, there is
|
||||
**no version skew, no loader, no search path, and no shared file on disk** —
|
||||
which is what makes this the resilient way to have a private ABI
|
||||
(docs/resilience.md), where a conventional `ld.so` + `/lib/libdanos.so`
|
||||
arrangement would add a loader to every spawn and a single shared point of
|
||||
failure.
|
||||
|
||||
The public danos ABI then has exactly two layers, neither of which is
|
||||
`abi.zig`:
|
||||
|
||||
| Layer | Contract | Spoken by |
|
||||
|-------|----------|-----------|
|
||||
| **vDSO** | C-ABI functions, this note | every language's thin shim (the `system-call` module for Zig, a `-sys` crate for Rust, a header for C) |
|
||||
| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](../file-system-development/vfs-protocol.md) is the first one documented) | any client that can lay out bytes |
|
||||
|
||||
Everything above those — the heap, `file_system`, the service harness — is
|
||||
per-language convenience, compiled into each binary from source, exactly as
|
||||
today. Nothing about the Zig runtime's shape changes; it just stops being the
|
||||
*only* door.
|
||||
|
||||
## The blob
|
||||
|
||||
A single copy of the vDSO code lives in the kernel image (built by
|
||||
`build.zig` as a tiny freestanding object, embedded like the AP trampoline).
|
||||
At boot the kernel finalises it once — this is where randomised numbers would
|
||||
be patched in — and thereafter maps the **same physical pages** read-execute
|
||||
into every process's address space. The blob is:
|
||||
|
||||
- **Position-independent.** It is mapped at a per-process randomised base, so
|
||||
it must be PIC (rip-relative addressing only — no relocations to process).
|
||||
- **Stateless and re-entrant.** No writable data. Anything stateful belongs to
|
||||
the process, not the vDSO.
|
||||
- **Architecture-specific.** The x86-64 blob wraps `syscall`; an aarch64 blob
|
||||
wraps `svc #0`. It lives beside the other per-architecture kernel sources
|
||||
(`system/kernel/architecture/<arch>/`), selected the same way the
|
||||
`architecture` module is (docs/architecture.md).
|
||||
|
||||
### Shape: a function table, not an ELF
|
||||
|
||||
A real `.so` with a dynamic symbol table is the conventional vDSO shape, but
|
||||
linking against one at load time needs a dynamic linker in every binary —
|
||||
machinery danos deliberately doesn't have. Instead the v1 shape is the
|
||||
simplest thing that is still a stable contract — a **function-pointer table**
|
||||
at the vDSO base:
|
||||
|
||||
```
|
||||
offset 0 u64 magic 'danosVDS' — a mapped-the-wrong-thing guard
|
||||
offset 8 u64 api_level incremented when the table grows
|
||||
offset 16 u64 count number of table entries that follow
|
||||
offset 24 u64 table[count] function pointers into the vDSO's own code
|
||||
```
|
||||
|
||||
Table *indices* are the public constants (published in a C header,
|
||||
`danos.h`), assigned once and append-only — the same discipline the IPC
|
||||
protocols use for operation values. The pointers point at stubs inside the
|
||||
blob; what those stubs put in `rax` is nobody's business but the kernel's.
|
||||
A language shim binds in one step: read the base from the init block, check
|
||||
the magic, keep the table pointer. Feature detection for a binary built
|
||||
against older headers is `count`/`api_level` — a kernel never removes or
|
||||
reorders entries.
|
||||
|
||||
(If danos ever grows a real dynamic linker, the same blob can additionally
|
||||
present an ELF `dynsym` without breaking the table — Fuchsia's vDSO is
|
||||
likewise both a mappable blob and a linkable `.so`. That is a later
|
||||
convenience, not a requirement.)
|
||||
|
||||
### Delivery: the auxiliary vector
|
||||
|
||||
The kernel already builds a System V entry block — argc, argv, envp
|
||||
terminator, **auxiliary vector** — on every new process's stack
|
||||
(`buildEntryStack`, read by the `start` module). The vDSO base rides in a new
|
||||
auxv entry, exactly Linux's `AT_SYSINFO_EHDR` move. No new syscall, no magic
|
||||
address, and a language shim finds it the same portable way on every
|
||||
architecture.
|
||||
|
||||
## The function surface
|
||||
|
||||
One table entry per kernel call, C ABI (System V AMD64), names prefixed
|
||||
`danos_`. The current `SystemCall` set maps directly; integer arguments and
|
||||
returns are `u64`, errors return as negative values exactly as today.
|
||||
|
||||
The calls that return two values in `rax:rdx` today — `dma_alloc`
|
||||
(virtual_address + physical_address), `msi_bind` (address + data), `shared_memory_create` (virtual_address + handle),
|
||||
`fs_resolve` (route tag + node token / backend handle) —
|
||||
become functions returning a two-`u64` struct. The System V ABI returns a
|
||||
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
|
||||
C-ABI spelling of the existing convention, at zero cost. The one call that
|
||||
returns *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
|
||||
`rdx`, received capability in `r8`) — exceeds the two-register return: its
|
||||
function returns a three-`u64` struct, which the ABI passes via a hidden
|
||||
result pointer, so that one stub stores `rax`/`rdx`/`r8` through the pointer
|
||||
after the `syscall` — a few instructions rather than one.
|
||||
|
||||
Grouped as `abi.zig` groups them:
|
||||
|
||||
| Group | Functions |
|
||||
|-------|-----------|
|
||||
| process | `danos_exit`, `danos_yield`, `danos_sleep`, `danos_spawn`, `danos_process_enumerate`, `danos_process_kill`, `danos_process_exit_reason`, `danos_process_subscribe`, `danos_process_signal`, `danos_signal_bind` |
|
||||
| threads | `danos_thread_spawn`, `danos_thread_exit`, `danos_current_core`, `danos_futex_wait`, `danos_futex_wake`, `danos_thread_self`, `danos_thread_join`, `danos_set_thread_pointer` |
|
||||
| memory | `danos_mmap`, `danos_munmap`, `danos_dma_alloc`, `danos_dma_free`, `danos_shared_memory_create`, `danos_shared_memory_map`, `danos_shared_memory_physical` |
|
||||
| ipc | `danos_endpoint_create`, `danos_ipc_register`, `danos_ipc_lookup`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` |
|
||||
| devices | `danos_device_enumerate`, `danos_device_claim`, `danos_device_register`, `danos_mmio_map`, `danos_irq_bind`, `danos_irq_ack`, `danos_msi_bind`, `danos_io_read`, `danos_io_write` |
|
||||
| time | `danos_clock`, `danos_wall_clock`, `danos_timer_bind` |
|
||||
| diagnostics | `danos_debug_write` (leveled, kernel-stamped records), `danos_klog_read`, `danos_klog_status` |
|
||||
| filesystem naming | `danos_fs_resolve`, `danos_fs_node`, `danos_fs_mount`, `danos_fs_unmount` (naming only — file DATA still crosses the vfs-protocol IPC, see below) |
|
||||
|
||||
The constants that ride alongside the calls — mmap protection bits, DMA
|
||||
flags, notification badge bits, `ExitReason`, `Signal`, well-known service
|
||||
ids, `page_size`, the IPC message maximum — move to the public header too:
|
||||
they are wire values a Rust program needs verbatim. What stays private in
|
||||
`abi.zig` is exactly the thing the vDSO exists to hide: the `SystemCall`
|
||||
numbers and the trap convention.
|
||||
|
||||
## Enforcement, and an honest threat model
|
||||
|
||||
Renumbering only has teeth if the kernel **refuses syscalls that don't come
|
||||
from the vDSO**. The check is cheap: on kernel entry, the saved user `rip`
|
||||
must lie inside the calling process's vDSO mapping; otherwise the process is
|
||||
killed with a fault-class exit reason (its supervisor restarts or gives up,
|
||||
docs/process-lifecycle.md — a foreign-syscall attempt is a bug or an attack,
|
||||
never something to limp past). Fuchsia enforces exactly this.
|
||||
|
||||
What this buys, precisely:
|
||||
|
||||
- **ABI freedom** — the real prize. The numbers can change per release or per
|
||||
boot and nothing outside the kernel image cares. The private ABI stays
|
||||
actually private, permanently.
|
||||
- **A single audited chokepoint** for kernel entry, per process, at a
|
||||
randomised address.
|
||||
- **Raised bar for exploits**: shellcode can't issue a hard-coded `syscall`;
|
||||
it must first discover the per-process vDSO base (ASLR) and call through
|
||||
it.
|
||||
|
||||
What it does *not* buy: an attacker with arbitrary code execution in a
|
||||
process can still *call* the vDSO functions — they are mapped executable in
|
||||
that process, and return-oriented chains reach them. Syscall randomisation is
|
||||
hardening, not a security boundary; the security boundary remains the
|
||||
capability model (what the process's endpoints and device claims let it do).
|
||||
It is worth building anyway — for the ABI freedom first and the hardening
|
||||
second — but the design should never be sold as more than that.
|
||||
|
||||
## Migration
|
||||
|
||||
Phased so every step ships alone (the M-milestone discipline):
|
||||
|
||||
1. **The blob + the table.** Build the vDSO, map it at spawn, deliver the
|
||||
base via auxv. `library/kernel/system-call.zig` binds through the table when the
|
||||
auxv entry is present, falls back to raw `syscall` when absent — the whole
|
||||
tree keeps booting during the transition.
|
||||
2. **Cut the system library over.** Delete the raw stubs; the `system-call`
|
||||
module no longer imports the `SystemCall` numbers at all (`abi.zig`'s enum becomes
|
||||
kernel-internal). The QEMU suite passing proves the table carries the
|
||||
whole system.
|
||||
3. **Enforce + randomise.** Add the `rip`-range check, then per-boot number
|
||||
randomisation patched into the blob at kernel init. A test boots with
|
||||
randomisation on and runs the full suite.
|
||||
4. **The other languages.** Publish `danos.h`; a Rust `danos-sys` crate wraps
|
||||
the table. This is also the seam `std.os.danos` calls through when the Zig
|
||||
self-hosting fork lands (docs/zig-self-hosting.md) — the vDSO is what
|
||||
makes that seam stable across kernel versions.
|
||||
|
||||
## What deliberately stays out
|
||||
|
||||
- **No dynamic linker, no `/lib/*.so`.** The vDSO is kernel-injected precisely
|
||||
so danos binaries can stay fully static above it. Sharing *library code*
|
||||
across processes stays what it is today: a service behind IPC, or source
|
||||
compiled into each binary.
|
||||
- **No file/device I/O in the vDSO.** The microkernel line doesn't move: the
|
||||
vDSO wraps the same deliberately tiny table (docs/syscall.md). The kernel
|
||||
resolves file NAMES (`fs_resolve` — the mount table moved in-kernel), but
|
||||
file data is still the filesystem server's business over the vfs-protocol
|
||||
IPC; the kernel never blocks on a userspace filesystem.
|
||||
- **No fast-path user-mode implementations yet.** Linux's vDSO exists mostly
|
||||
to answer `gettimeofday` without a kernel entry. `danos_clock` could one
|
||||
day read the calibrated TSC in user mode the same way — the blob is where
|
||||
such an optimisation would live — but that is an optimisation, not part of
|
||||
this design's contract.
|
||||
Reference in New Issue
Block a user