documenting research for arm and device discovery
This commit is contained in:
@@ -51,6 +51,12 @@ Cutting across all of these:
|
||||
- **[arch.md](arch.md) — the architecture split.** How CPU-specific code is kept
|
||||
behind a build-time `arch` module so the generic kernel never names x86_64,
|
||||
leaving room for other systems (e.g. an AArch64 Raspberry Pi) later.
|
||||
- **[arm.md](arm.md) — ARM targets.** The Raspberry Pi landscape the arch split is
|
||||
aiming at: `arm` (32-bit, Pi Zero W) vs `aarch64` (64-bit, Pi 3-5), UEFI vs
|
||||
device-tree boot, and what each layer needs.
|
||||
- **[discovery.md](discovery.md) — device discovery.** A design note (not built yet)
|
||||
on learning what hardware exists via ACPI (x86) or device tree (ARM) behind one
|
||||
neutral device model — when to build it, and how to keep it architecture-agnostic.
|
||||
- **[sysv.md](sysv.md) — the calling convention.** What "the kernel is SysV" means,
|
||||
and why the loader→kernel boundary has to pin it (the RDI-vs-RCX handoff).
|
||||
- **[testing.md](testing.md) — testing.** How the kernel is tested by booting it in
|
||||
|
||||
+120
@@ -0,0 +1,120 @@
|
||||
# ARM targets (`arm` and `aarch64`)
|
||||
|
||||
danos aims to run on Raspberry Pi hardware eventually. "ARM" isn't one target,
|
||||
though — the Pis span **two different CPU architectures** (32-bit `arm` and 64-bit
|
||||
`aarch64`) and (stock) a different boot protocol from x86-64's UEFI. **danos targets
|
||||
`aarch64` only** (see the decision below); the `arm`/`aarch64` distinction still
|
||||
matters for understanding why. This page maps the landscape so the
|
||||
[arch split](arch.md) and build system can be planned for it.
|
||||
|
||||
## `arm` vs `aarch64` — 32-bit vs 64-bit
|
||||
|
||||
- **`arm`** = **32-bit** ARM (the *AArch32* state, A32/T32 instruction sets).
|
||||
ARMv7 and earlier, plus the 32-bit compatibility mode of newer cores. 16 × 32-bit
|
||||
registers.
|
||||
- **`aarch64`** = **64-bit** ARM (the *AArch64* state, A64 instruction set), from
|
||||
**ARMv8-A** on. Also called **arm64**. 31 × 64-bit registers, a fixed 32-bit
|
||||
instruction width, a redesigned exception model — *not* a widening of A32, a clean
|
||||
new ISA.
|
||||
|
||||
They are as different from each other as either is from x86-64: separate registers,
|
||||
page-table formats, and calling conventions. Each needs its own `src/arch/<name>/`.
|
||||
|
||||
## The Raspberry Pi models
|
||||
|
||||
| Model | SoC | Core | Architecture | danos target |
|
||||
|-------|-----|------|--------------|--------------|
|
||||
| Pi Zero / Zero W | BCM2835 | ARM1176JZF-S | ARMv6, 32-bit only | `arm` (not planned) |
|
||||
| **Pi Zero 2 W** ← target | BCM2710 | Cortex-A53 | ARMv8-A, 64-bit | **`aarch64`** |
|
||||
| **Pi 3 / 3B+** | BCM2837 | Cortex-A53 | ARMv8-A, 64-bit | **`aarch64`** |
|
||||
| **Pi 4** | BCM2711 | Cortex-A72 | ARMv8-A, 64-bit | **`aarch64`** |
|
||||
| **Pi 5** | BCM2712 | Cortex-A76 | ARMv8.2-A, 64-bit | **`aarch64`** |
|
||||
|
||||
> **Decision: `aarch64` only.** The target small board is a **Pi Zero 2 W** (BCM2710,
|
||||
> Cortex-A53) — which is **`aarch64`**, *not* the original Zero W's 32-bit ARMv6. So
|
||||
> every ARM board danos targets (Zero 2 W and Pi 3-5) is `aarch64`, and the 32-bit
|
||||
> `arm`/ARMv6 backend is **not planned** — one ARM CPU port, not two. The original
|
||||
> Zero W (ARMv6) would only re-enter scope if that specific older board were ever
|
||||
> needed; the row above is kept only to explain the distinction.
|
||||
|
||||
## Booting: UEFI is not x86-only
|
||||
|
||||
The boot protocol is a **separate axis** from the CPU (see [arch.md](arch.md)):
|
||||
|
||||
- **UEFI** exists for ARM too — ARM servers require it (SBSA/SBBR), QEMU boots it
|
||||
with **AAVMF** (the AArch64 build of the same EDK2 firmware as x86's OVMF), and
|
||||
the Pi can even run it with community UEFI firmware. Under UEFI the handoff is the
|
||||
*same* as x86-64: system table, boot services, memory map, GOP framebuffer — so
|
||||
the loader logic largely carries over.
|
||||
- **Device tree / firmware boot** — the **stock** Raspberry Pi firmware (VideoCore
|
||||
bootloader) is *not* UEFI: it loads the kernel and jumps to it with a **device-tree
|
||||
blob (DTB)** pointer. Both the Zero W and stock Pi 3-5 boot this way.
|
||||
|
||||
Note that even under UEFI on ARM, the OS still gets its hardware description from
|
||||
**ACPI or a device tree** (often the DTB passed via a UEFI configuration table). So
|
||||
"UEFI on ARM" doesn't remove the device tree — UEFI gives you memory + framebuffer;
|
||||
the DTB/ACPI tells you what devices exist.
|
||||
|
||||
## What danos needs, layer by layer
|
||||
|
||||
- **One CPU arch module: `src/arch/aarch64/`** — covering the Zero 2 W and Pi 3-5,
|
||||
providing the same `arch` interface as x86_64: `halt`, context switch,
|
||||
interrupt/exception vectors, page tables, a UART, a timer. No `src/arch/arm/` is
|
||||
planned (see the decision above), so there's a single ARM backend to write.
|
||||
- **A device-tree boot path.** Since stock Pis boot via DTB, danos needs an entry
|
||||
that parses the DTB's `/memory` and `/reserved-memory` into the neutral
|
||||
[`MemoryMap`](memory-map.md) — the same neutral handoff `efi.zig` produces, just
|
||||
from a different source. This is where keeping boot-protocol knowledge on the
|
||||
loader side (as we did for the UEFI memory-map classification) pays off.
|
||||
- **The UEFI loader mostly carries over.** `src/efi.zig` is largely
|
||||
boot-*protocol* code (`std.os.uefi` protocol calls), not x86 code. Its only truly
|
||||
x86-specific bits are the ELF machine check (`.X86_64`) and the SysV calling
|
||||
convention for the kernel jump. So an `aarch64`-UEFI target (QEMU `virt` + AAVMF)
|
||||
can reuse it — which makes **aarch64-UEFI the easiest second target**, easier than
|
||||
the device-tree Pi.
|
||||
|
||||
## Pi hardware quirks (for when we port)
|
||||
|
||||
The Pi is not a "standard" ARM platform — expect Broadcom-specific peripherals:
|
||||
|
||||
- **Peripheral base moves per SoC**: `0x2000_0000` (BCM2835, Zero W),
|
||||
`0x3F00_0000` (BCM2837, Pi 3), `0xFE00_0000` (BCM2711, Pi 4), different again on
|
||||
Pi 5. Everything below is an offset from it.
|
||||
- **UART**: a **PL011** (at base + `0x20_1000`) plus a mini-UART; on some boards the
|
||||
PL011 is wired to Bluetooth, so which one is the console varies. This is the
|
||||
`aarch64`/`arm` equivalent of our x86 [COM1 serial](testing.md).
|
||||
- **Interrupt controller**: *not* a standard ARM GIC on the older parts — the Zero W
|
||||
and Pi 3 use Broadcom's own ARMCTRL controller (Pi 3 adds a per-core "local"
|
||||
controller for timers/mailboxes). The **Pi 4 and 5 do have a GIC-400**. So the
|
||||
interrupt backend differs even within the `aarch64` Pis.
|
||||
- **Timer**: the ARM generic timer (`CNTPCT`/`CNTFRQ`) on ARMv8, or the BCM system
|
||||
timer — the counterpart to our calibrated LAPIC/TSC clock.
|
||||
|
||||
## Building each (intended)
|
||||
|
||||
`build.zig` currently pins the kernel to `x86_64`; supporting these means selecting
|
||||
the target and `arch` module together (e.g. a `-Darch=` option). The Zig target
|
||||
queries would be roughly:
|
||||
|
||||
- **Pi Zero W**: `.cpu_arch = .arm`, `.cpu_model = arm1176jzf_s`, `.os_tag = .freestanding`
|
||||
- **Pi 3**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a53`, `.os_tag = .freestanding`
|
||||
- **Pi 4**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a72`
|
||||
- **Pi 5**: `.cpu_arch = .aarch64`, `.cpu_model = cortex_a76`
|
||||
|
||||
## Testing in QEMU
|
||||
|
||||
Two routes, mirroring how we test x86-64 with OVMF:
|
||||
|
||||
- **Board emulation**: `qemu-system-aarch64 -machine raspi3b` (and `raspi4b` on
|
||||
recent QEMU) for the Pi 3/4; `qemu-system-arm -machine raspi0`/`raspi1ap` for the
|
||||
ARMv6 Zero-class board — closest to real hardware, device-tree boot.
|
||||
- **Generic aarch64-UEFI**: `qemu-system-aarch64 -machine virt` + AAVMF — the
|
||||
cleanest way to bring up the `aarch64` kernel via the reused UEFI loader before
|
||||
tackling Pi-specific boards. A future `run-aarch64` build step would use this.
|
||||
|
||||
## Related
|
||||
|
||||
- [arch.md](arch.md) — the arch-module boundary these targets plug into, and the
|
||||
CPU-arch vs boot-protocol "two axes".
|
||||
- [efi.md](efi.md) — the UEFI loader that carries over to aarch64-UEFI.
|
||||
- [vision.md](vision.md) — why isolated, portable-across-architectures is the goal.
|
||||
@@ -0,0 +1,167 @@
|
||||
# Device discovery (ACPI / device tree), the agnostic way
|
||||
|
||||
"Device discovery" is how the kernel learns **what hardware exists and where** — the
|
||||
MMIO addresses, IRQ numbers, CPU count, and interrupt controller it can't just
|
||||
assume. On x86 that description comes from **ACPI** tables; on ARM from a **device
|
||||
tree** (DTB). This note is a design plan, not built yet: *when* danos should tackle
|
||||
it, and *how* to keep it architecture-agnostic — the same discipline the
|
||||
[memory map](memory-map.md) and [arch split](arch.md) already follow.
|
||||
|
||||
## What the kernel assumes today
|
||||
|
||||
Right now danos discovers almost nothing — it coasts on legacy PC fixtures that are
|
||||
guaranteed to exist under QEMU + UEFI:
|
||||
|
||||
- `arch/x86_64/apic.zig` assumes the **Local APIC** at the default `0xFEE0_0000` and
|
||||
calibrates its timer against the **PIT** (the legacy 8254).
|
||||
- `arch/x86_64/serial.zig` hardcodes **COM1** at I/O port `0x3F8`.
|
||||
- The framebuffer and memory map come from **UEFI** — that *is* discovery, just done
|
||||
by the firmware and handed over, not read from ACPI.
|
||||
|
||||
This works only because PC-compatible hardware promises those legacy pieces exist at
|
||||
those addresses. It is a crutch, and it does not travel.
|
||||
|
||||
## The forcing functions: when to build it
|
||||
|
||||
Two things drive the need, and they set the timing:
|
||||
|
||||
1. **The second architecture makes it mandatory.** ARM has *no* legacy fixtures —
|
||||
no PIT, no fixed serial port, no standard interrupt controller address. You can't
|
||||
find the UART to print a character without reading the device tree. So on x86 we
|
||||
can defer discovery a long time (until we want the IOAPIC, PCIe, or SMP), but on
|
||||
**aarch64 it's required to boot at all**. The [aarch64 port](arm.md) is what
|
||||
forces the issue.
|
||||
|
||||
2. **Isolated user-space drivers need it.** In the [microkernel vision](vision.md),
|
||||
drivers live in user space — but something has to enumerate the hardware and hand
|
||||
each driver its MMIO regions and IRQs. That enumeration *is* device discovery. So
|
||||
discovery is a prerequisite for real drivers, **not** for user mode itself.
|
||||
|
||||
The conclusion on timing: **don't build full discovery before user mode.** User mode
|
||||
+ address-space isolation needs none of it; the current assumptions are fine there.
|
||||
Build the agnostic discovery layer **when the aarch64 port starts** — because that's
|
||||
when a second, real implementation makes "agnostic" honest.
|
||||
|
||||
## Why build it *with* the second arch, not before
|
||||
|
||||
The same lesson as the arch split: an abstraction with only one implementation
|
||||
quietly bends to that implementation. Build "agnostic discovery" x86-only first and
|
||||
you'll get an ACPI-shaped interface with device tree bolted on afterward. Build it
|
||||
when aarch64 lands and the small, clean DTB parser pulls the abstraction toward the
|
||||
right neutral shape, which the x86 side then fills. Design it against two backends or
|
||||
it isn't really agnostic.
|
||||
|
||||
## The one cheap step to take sooner
|
||||
|
||||
Have the **loader capture the description pointer** into `BootInfo` — a neutral
|
||||
handle, no parsing:
|
||||
|
||||
```zig
|
||||
pub const HardwareInfo = extern struct {
|
||||
kind: enum(u32) { none, acpi, device_tree },
|
||||
addr: u64, // ACPI RSDP, or the DTB blob
|
||||
};
|
||||
```
|
||||
|
||||
On x86-UEFI that's the ACPI **RSDP**, read from the UEFI configuration table *before*
|
||||
`ExitBootServices` — the same "grab it before exit" pattern as the framebuffer and
|
||||
memory map (already flagged in [memory-map.md](memory-map.md)). On ARM it's the DTB
|
||||
pointer the firmware passes. This keeps the door open for near-zero cost without
|
||||
committing to the parser.
|
||||
|
||||
## What "agnostic" looks like
|
||||
|
||||
The happy accident: **the device-tree data model is already a good neutral
|
||||
representation.** A DTB is a tree of nodes, each with:
|
||||
|
||||
- a **`compatible`** string — what the device is,
|
||||
- **`reg`** — its MMIO base(s) and size(s),
|
||||
- **`interrupts`** — its IRQ number(s),
|
||||
- other properties.
|
||||
|
||||
Even OSes running on ACPI hardware normalize into a unified device model shaped like
|
||||
this. So the neutral layer is a **device model** — "here are the devices, each with a
|
||||
type, MMIO regions, and IRQs" — fed by two backends behind it:
|
||||
|
||||
- a **DTB parser** (ARM) — a few hundred lines against a well-specified binary format;
|
||||
- **static ACPI + PCI enumeration** (x86).
|
||||
|
||||
### The caveat that sets the effort: "ACPI" ≠ "AML interpreter"
|
||||
|
||||
A full ACPI namespace is **AML** (ACPI Machine Language) bytecode, and writing an AML
|
||||
interpreter is an enormous undertaking. **We skip it.** Interrupt routing and PCIe
|
||||
come entirely from the *static* tables:
|
||||
|
||||
- **MADT** — the APICs and the **IOAPIC** (interrupt routing),
|
||||
- **MCFG** — PCIe configuration space (then walk the PCI bus to enumerate devices),
|
||||
- **FADT**, **HPET** — power/reset and a precise timer.
|
||||
|
||||
Walking PCI config space finds most devices without any AML. So the x86 static-ACPI
|
||||
backend is roughly comparable in scope to the DTB parser; it's full AML that's the
|
||||
monster, and it isn't on the path.
|
||||
|
||||
## Where it sits in the seam
|
||||
|
||||
Discovery layers onto the existing loader↔kernel split cleanly:
|
||||
|
||||
| Layer | Responsibility | x86 | ARM |
|
||||
|-------|----------------|-----|-----|
|
||||
| **Capture** (loader) | grab the description pointer | RSDP from UEFI config table | DTB pointer from firmware |
|
||||
| **Parse** (kernel, per-mechanism) | pointer → neutral device model | static ACPI + PCI | DTB parser |
|
||||
| **Consume** (generic) | use the device model | — same code — | — same code — |
|
||||
|
||||
Boot-protocol-specific capture, mechanism-specific parse, generic consumption —
|
||||
exactly like the memory map, where the loader classifies and the kernel just sees
|
||||
neutral regions.
|
||||
|
||||
## Where it lives in a microkernel
|
||||
|
||||
For isolated drivers, discovery is not one lump — it splits by *who needs it and
|
||||
when*:
|
||||
|
||||
- **Minimal, in-kernel: the interrupt controller and the timer.** The IOAPIC/GIC and
|
||||
the timer are needed *before* user space exists (scheduling and preemption depend
|
||||
on them), so the kernel must parse at least these from ACPI/DTB itself. This is also
|
||||
the **first real consumer** of discovery — the first genuine reason to read MADT or
|
||||
the DTB is "where is the interrupt controller and how do I route an IRQ?" It's the
|
||||
point where x86 finally graduates from the PIT/legacy-LAPIC assumptions.
|
||||
- **User-space enumeration: a device-manager server.** Everything else — PCI devices,
|
||||
peripherals — is parsed (or queried from the kernel's parse) by a privileged
|
||||
user-space server that hands each driver process its MMIO regions and IRQ rights
|
||||
over [IPC](ipc.md). Combined with **interrupts-as-messages** (an IRQ delivered to a
|
||||
driver as a message on a channel — a natural extension of the wait queues and
|
||||
channels already built), that's what makes drivers genuinely isolated.
|
||||
|
||||
So the device-manager server depends on user mode + IPC; only the interrupt-controller
|
||||
slice is unavoidably in-kernel.
|
||||
|
||||
## A nice ARM contrast
|
||||
|
||||
On ARMv8 the generic timer exposes its frequency directly via the `CNTFRQ` register —
|
||||
no calibration needed. That's cleaner than the x86 side, where we measure the LAPIC
|
||||
and TSC against the PIT because nothing tells us their frequency (see
|
||||
[device-interrupts.md](device-interrupts.md)). Discovery on ARM hands you more for
|
||||
free; discovery on x86 is partly about *finding* what ARM just tells you.
|
||||
|
||||
## Suggested ordering
|
||||
|
||||
1. **Now (cheap):** plumb the neutral `HardwareInfo` pointer through the loader into
|
||||
`BootInfo`. No parser yet.
|
||||
2. **Next milestone unchanged:** user mode + address-space isolation — needs no
|
||||
discovery.
|
||||
3. **With the aarch64 port:** build the agnostic discovery layer, **DTB first** (the
|
||||
forcing function), then x86 static-ACPI to fill the same model — starting with the
|
||||
**interrupt controller + timer**.
|
||||
4. **Then:** a user-space device-manager server + interrupts-as-messages → real
|
||||
isolated drivers (keyboard first).
|
||||
|
||||
## Related
|
||||
|
||||
- [arm.md](arm.md) — the aarch64 target that forces genuine discovery (DTB, GIC).
|
||||
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam,
|
||||
and the note about grabbing the RSDP before `ExitBootServices`.
|
||||
- [device-interrupts.md](device-interrupts.md) — the LAPIC/timer bring-up that
|
||||
discovery will eventually feed (IOAPIC, real IRQ routing).
|
||||
- [ipc.md](ipc.md) — the channels that interrupts-as-messages and the device manager
|
||||
will ride on.
|
||||
- [vision.md](vision.md) — why drivers belong in isolated user space at all.
|
||||
Reference in New Issue
Block a user