add device platform module with ACPI support

This commit is contained in:
Daniel Samson
2026-07-08 09:23:41 +01:00
parent 53a33a7332
commit 2ee898a91e
21 changed files with 3545 additions and 5 deletions
+6 -3
View File
@@ -59,9 +59,12 @@ Cutting across all of these:
- **[arm.md](arm.md) — ARM targets.** The Raspberry Pi landscape the arch split is
aiming at: `arm` (32-bit, Pi Zero W) vs `aarch64` (64-bit, Pi 3-5), UEFI vs
device-tree boot, and what each layer needs.
- **[discovery.md](discovery.md) — device discovery.** A design note (not built yet)
on learning what hardware exists via ACPI (x86) or device tree (ARM) behind one
neutral device model — when to build it, and how to keep it architecture-agnostic.
- **[discovery.md](discovery.md) — device discovery.** A design note on learning what
hardware exists via ACPI (x86) or device tree (ARM) behind one neutral device model —
when to build it, and how to keep it architecture-agnostic.
- **[acpi.md](acpi.md) — finding the ACPI tables.** The concrete x86 locator chain:
how the loader captures the **RSDP**, hands its physical address across in `BootInfo`,
and how the platform derives the **RSDT/XSDT** from it and walks the SDTs.
- **[smp.md](smp.md) — multiple cores.** A design/research note on how microkernels
(L4, seL4) handle SMP — big kernel lock vs per-CPU vs multikernel — and how the
right choice depends on whether danos is chasing real-time or resilience.
+118
View File
@@ -0,0 +1,118 @@
# ACPI: finding the tables (RSDP → RSDT/XSDT → SDTs)
ACPI describes the hardware the kernel can't assume — the interrupt controllers, the
PCIe config window, the timer, the power registers — in a set of **system
description tables** (SDTs). But before it can read any of them, danos has to *find*
them, and they aren't at a fixed address. Getting there is a short chain of pointers,
and this note explains it — in particular the question it's easy to trip on: **how
does the [platform / device module](arch.md) know where the RSDT is?**
Short answer: it doesn't receive the RSDT. The firmware hands over the **RSDP**, and
the RSDT's address is a field *inside* the RSDP. The platform follows that pointer.
## The locator chain
```
UEFI configuration table
│ the loader reads the RSDP's physical address
▼
BootInfo.acpi_rsdp (u64, in the shared `danos` module) src/root.zig
│ the kernel forwards the whole BootInfo
▼
platform.discover(boot_info, …) src/device/platform.zig
│ reads boot_info.acpi_rsdp, hands it to the ACPI backend
▼
acpi.discover(rsdp_phys, …) src/device/acpi.zig
│ dereferences the RSDP, reads the pointer it contains
▼
RSDP ──(a field in the struct)──► RSDT / XSDT ──► SDTs (MADT, MCFG, FADT, HPET, DSDT…)
```
The **RSDP** (Root System Description Pointer) is the root of the whole ACPI tree.
Its only job is to point at the root *table* — the **RSDT** (ACPI 1.0) or its 64-bit
successor the **XSDT** (ACPI 2.0+) — which in turn lists every other SDT.
## Step 1 — the loader finds the RSDP
Only the firmware knows where ACPI lives, so the RSDP must be grabbed while UEFI is
still up. `acpiRootSystemDescriptorPointer()` in `src/boot/efi.zig` walks the UEFI
**configuration table** for the ACPI GUID and returns the vendor pointer — the same
"grab it before `ExitBootServices`" pattern as the [framebuffer](framebuffer.md) and
the [memory map](memory-map.md).
## Step 2 — the handoff: a physical address in `BootInfo`
The loader can't just call the device module: the bootloader binary and the kernel
binary are compiled separately, and **the loader isn't linked against the `platform`
module at all** (it imports only the shared `danos` module). So instead of a call, it
deposits a value in the handoff struct:
```zig
// src/boot/efi.zig — while boot services are still up
.acpi_rsdp = if (acpiRootSystemDescriptorPointer()) |p| @intFromPtr(p) else 0,
```
Two things about what crosses the boundary:
- **It's a *physical* address, not a Zig pointer.** The loader and kernel don't share
an address space at the moment of the jump, so a raw `u64` physical address is the
only thing that survives the handoff. `BootInfo.acpi_rsdp` is `0` when the firmware
exposed no ACPI (e.g. a future device-tree machine, which would fill a different
field instead — the kernel never learns which firmware booted it).
- **The kernel can dereference it because it identity-maps ACPI memory.** The RSDP
lives in ACPI-reclaim memory, which [paging.zig](paging.md) identity-maps along with
the rest of RAM, so by the time discovery runs `@ptrFromInt(rsdp_phys)` is a valid
pointer.
This is the concrete form of the "capture the description pointer" step sketched in
[discovery.md](discovery.md) — a plain `acpi_rsdp: u64` rather than a tagged handle,
since x86 is the only backend wired up so far.
## Step 3 — the platform derives the RSDT from the RSDP
`acpi.discover` reinterprets the physical address as the RSDP struct, validates it,
and then reads the root-table pointer *out of it*. Which pointer depends on the ACPI
version, because the RSDP carries **both**:
```zig
const rsdp: *const RootSystemDescriptionPointer = @ptrFromInt(rsdp_phys);
if (!std.mem.eql(u8, &rsdp.signature, "RSD PTR ")) return error.BadRsdpSignature;
if (!checksumOk(@ptrFromInt(rsdp_phys), 20)) return error.BadRsdpChecksum;
if (rsdp.revision >= 2) {
// ACPI 2.0+: use the 64-bit XSDT pointer (the 32-bit RSDT is deprecated)
const xsdp: *const ExtendedSystemDescriptorPointer = @ptrFromInt(rsdp_phys);
try walkRoot(u64, xsdp.extended_system_descriptor_table_address, …);
} else {
// ACPI 1.0: use the 32-bit RSDT pointer
try walkRoot(u32, rsdp.root_system_description_table_address, …);
}
```
- The **`revision`** byte selects the root table. `root_system_description_table_address`
(32-bit, → RSDT) and `extended_system_descriptor_table_address` (64-bit, → XSDT) are
ordinary fields of the RSDP/XSDP structs — the platform never *receives* the RSDT
address, it *reads* it here. On QEMU q35 the RSDP is revision 2, so the XSDT path is
taken.
- The signature (`"RSD PTR "`) and one-byte checksum guard against a bad pointer before
anything downstream trusts it.
## After the root table
`walkRoot` treats the RSDT/XSDT as an array of physical pointers — 32-bit entries for
the RSDT, 64-bit for the XSDT — and hands each SDT to `handleTable`, which dispatches
on its 4-byte signature: **MADT** (CPUs + IOAPIC), **MCFG** (PCIe ECAM), **FADT**
(power registers, and the pointer to the DSDT), **HPET** (timer). That's where the
firmware-agnostic [device model](discovery.md) gets populated; this note stops at the
part that answers "where are the tables?" — everything past the RSDP is just following
more pointers the tables themselves provide.
## Related
- [efi.md](efi.md) — the loader that captures the RSDP before `ExitBootServices`.
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam, and
the ACPI-reclaim memory the RSDP lives in.
- [discovery.md](discovery.md) — the broader (still-evolving) plan for turning these
tables into one neutral device model shared with the ARM device-tree path.
- [arch.md](arch.md) — why the kernel reaches the device code through a `platform`
module and never names ACPI directly.
+2
View File
@@ -157,6 +157,8 @@ free; discovery on x86 is partly about *finding* what ARM just tells you.
## Related
- [acpi.md](acpi.md) — the built x86 side of the "capture" step: how the loader grabs
the RSDP and the platform follows it to the RSDT/XSDT and the SDTs.
- [arm.md](arm.md) — the aarch64 target that forces genuine discovery (DTB, GIC).
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam,
and the note about grabbing the RSDP before `ExitBootServices`.
+80
View File
@@ -0,0 +1,80 @@
# System Calls
System calls (syscalls) are the bridge between your programs and the operating system's restricted core (kernel).
## The Mechanism of a Syscall
A system call follows a highly orchestrated, 6-step sequence to ensure hardware safety and process security:
1. Setting the Arguments: The program places a unique ID corresponding to the requested service (e.g., \(sys\_write\)) into a specific CPU register (like RAX in x86_64) along with its required parameters.
2. Executing the Trap: The program triggers a special CPU instruction, such as syscall or int 0x80. This acts as a software interrupt.
3. Mode Switch: The CPU atomically flips its execution privilege from unprivileged User Mode (Ring 3) to the highly privileged Kernel Mode (Ring 0).
4. Lookup and Execution: The kernel looks up the syscall ID in a dispatch table and runs the designated internal routine to perform the actual work (like fetching data from the hard drive).
5. Returning the Status: The kernel places the result of the operation—or an error code—back into the RAX register.
6. Return to User Mode: The kernel executes a return instruction (such as sysret), the CPU shifts back to User Mode, and your program continues executing.
## Approaches
There are two approaches to system calls, stable public ABI and unstable private ABI. A public ABI uses a standard defined set of numbers. It makes it easy to guess, easy to write compilers and tools for. private ABI tend to change the numbers to obscure the numbers, preventing attackers from bypassing libc using CPU instructions directly. Private ABIs force users to use a library like libSystem, which the kernel can inject the syscall numbers. An alternative solution is to just have 1 system number but make it generic by the address to a struct in memory.
Practical minimum primitives (Monolithic kernel)
| Category | Primitive | Detail |
|----------|--------------|---------------------------------------------------------------------------------------|
| lifecyle | spawn / exec | Loads a binary from storage into memory and begins execution. |
| lifecyle | exit | Terminates the current process and frees its memory back to the kernel. |
| I/O | read | Requests data from a hardware device or file descriptor into user memory. |
| I/O | write | Pushes data from user memory out to a device or file descriptor (like a screen). |
| Memory | brk / mmap | Requests the kernel to allocate more physical or virtual memory pages to the process. |
| Control | ioctl | A catch-all "escape hatch" call to send hardware-specific commands to device drivers. |
The Microkernel Minimum Set
Everything else---including`read()`,`write()`,`malloc()`, and`fork()`---will run in user space as servers (e.g., a VFS server, a memory manager server) that threads communicate with using these three calls:
1. **`IPC_Call(endpoint, message_buffer)`(Synchronous Send + Receive)**
- **What it does:**The calling thread sends a message block to a service endpoint and immediately blocks (sleeps) until that service processes the request and sends a reply back.
- **Why it's minimal:**Combining*Send*and*Receive*into a single atomic atomic system call eliminates the need for separate tracking and prevents a massive amount of context-switching overhead. This is the foundation of high-performance microkernels like[seL4](https://sel4.systems/)and L4.[[1](https://en.wikipedia.org/wiki/L4_microkernel_family),[2](https://www.microchip.com/en-us/products/microprocessors/64-bit-mpus/pic64hx/ecosystem)]
2. **`IPC_ReplyWait(endpoint, reply_buffer)`(Respond + Wait for Next)**
- **What it does:**Used strictly by your background user-space servers (like your disk driver or filesystem). It sends a reply to the last client that called it, and immediately puts the server to sleep until the next request arrives.[[1](https://news.ycombinator.com/item?id=33078441)]
3. **`Yield()`/`Thread_Ctrl()`**
- **What it does:**Allows a thread to voluntarily give up its CPU time slice, or allows a root task to spawn/kill threads.
* * * * *
Hardware Implementation: x86_64 vs. aarch64
Because you are targeting both platforms, you must design a clean**Architecture Abstraction Layer (AAL)**. Each architecture uses completely different assembly instructions, CPU registers, and privilege levels to jump from user space (Ring 3 / EL0) to kernel space (Ring 0 / EL1).[[1](https://android.googlesource.com/kernel/common/+/84d3e59750bbd/arch/arm64/Kconfig),[2](https://blog.codingconfessions.com/p/making-system-calls-in-x86-64-assembly),[3](https://alex.dzyoba.com/blog/os-segmentation/),[4](https://dev.to/ripan030/linux-kernel-interrupt-handling-part-2-fe1)]
Here is how you will map your bare minimum system calls on both architectures:
1\. x86_64 Implementation
On 64-bit Intel and AMD processors, you ignore the old`int 0x80`software interrupts. Instead, you use the high-performance`syscall`and`sysret`instructions.[[1](https://alex.dzyoba.com/blog/os-segmentation/)]
- **The Trap:**The user-space program executes the`syscall`instruction.[[1](https://namastedev.com/blog/kernel-vs-user-space-2/)]
- **The Registers:**The hardware automatically moves the instruction pointer, but you must pass your arguments in specific registers. A common microkernel convention mimics the System V AMD64 ABI:
- `rax`: System Call ID (e.g.,`0`for IPC_Call,`1`for IPC_ReplyWait)
- `rdi`: Argument 1 (Endpoint ID / Destination)
- `rsi`: Argument 2 (Pointer to the message payload buffer)
- `rdx`: Argument 3 (Size of the message)[[1](https://dev.to/kaamkiya/hello-world-in-assembly-x86-64-2kb8)]
- **Kernel Setup:**Your kernel must configure the Model Specific Registers (MSRs)---specifically`IA32_STAR`and`IA32_LSTAR`---during boot to point to your kernel's system call entry assembly code.[[1](https://nfil.dev/kernel/rust/coding/rust-kernel-to-userspace-and-back/),[2](https://johannst.github.io/notes/arch/x86_64.html)]
2\. aarch64 (ARM 64-bit) Implementation
On ARMv8-A and ARMv9-A architectures, privilege levels are called Exception Levels. User space runs at**EL0**, and your microkernel runs at**EL1**.[[1](https://community.nxp.com/pwmxy87654/attachments/pwmxy87654/imx-processors/183079/1/AN12212.pdf),[2](https://developer.arm.com/-/media/Arm%20Developer%20Community/PDF/Learn%20the%20Architecture/Exception%20model.pdf?revision=a62f2bf2-b08a-4a4f-8cbe-38c67ddf4434),[3](https://people.kernel.org/linusw/),[4](https://drewdevault.com/blog/Helios-aarch64/)]
- **The Trap:**The user-space program executes the`svc #0`(Supervisor Call) instruction.[[1](https://hackmd.io/@xlYUTygoRkyuQQlwXuWDWQ/SJWQuIsIZe)]
- **The Registers:**ARM provides a clean, plentiful register set. You typically pass your arguments in the standard parameter registers:
- `x0`: System Call ID
- `x1`: Argument 1 (Endpoint ID / Destination)
- `x2`: Argument 2 (Pointer to message payload buffer)
- `x3`: Argument 3 (Size of the message)[[1](https://github.com/lelegard/arm-cpusysregs),[2](https://medium.com/@vincentcorbee/http-server-in-arm64-assembly-apple-silicon-m1-077a55bbe9ca)]
- **Kernel Setup:**Your kernel must set up an Exception Vector Table and write its base address to the`VBAR_EL1`register. When`svc`is executed, the CPU jumps to the synchronous exception offset in that table.[[1](https://dev.to/ripan030/linux-kernel-interrupt-handling-part-2-fe1),[2](https://eastrivervillage.com/blog/archive/2018/06/),[3](https://www.wadixtech.com/blog/armv8-a-exception-levels-el0-to-el3)]
* * * * *
Managing the Payload Challenge
Because it is a microkernel, performance lives or dies by how fast your`IPC_Call`can move data from Client to Server. You have two minimal choices for handling the`message_buffer`pointer:[[1](https://anazimzada2020.medium.com/microkernel-architectural-pattern-5e4e9184170e)]
- **The Copy Method (Simplest to start):**Your kernel pauses the client, reads the data from the client's memory space, switches page tables to the server, and copies the data into the server's buffer.
- **The Shared Memory Method (Fastest):**The kernel sets up a temporary, shared virtual memory page between the client and server. The client writes to it, calls`syscall`/`svc`, and the server reads it instantly without the kernel copying any bytes