A user-space process can now touch real hardware directly, capability-gated by the device tree — the microkernel driver model. - src/kernel/devsvc.zig: flattens the discovered device tree into an id-indexed snapshot + a claim table at boot (devsvc.init from main.zig). - Syscalls 11-13: dev_enumerate (snapshot the table), dev_claim (take exclusive ownership), mmio_map (map a claimed device's MMIO window into the caller's AS and return the register base). The claim is the capability: mmio_map refuses any device the caller doesn't own. - paging.mapUserDeviceInto: maps device MMIO strong-uncacheable (PCD|PWT) and marks each leaf with a device_grant PTE bit; freeSubtree skips pmm.free on those leaves, so tearing down a driver never returns MMIO frames to the RAM pool (the teardown hazard). MMIO grants live in a distinct arena, PML4[226] (Task.dev_map_next), so device pages widen no kernel mapping. - lib/dev.zig: user enumerate/claim/mmioMap wrappers; shared DeviceDesc/ResDesc in danos (root.zig). sbin/hpetd.zig: finds the HPET, claims it, maps its registers, enables the counter (an MMIO write) and reads it (0xF0) — proving read+write passthrough to real hardware. - Tests: `hpet` (driver reads the counter advancing from ring 3) and `iopass` (device-granted frame survives address-space teardown). Suite 33/33. irq_bind/irq_ack (IRQ-as-message) are stubbed (-1) pending; notifyFromIsr (M7) is the hook they'll use.
238 lines
10 KiB
Zig
238 lines
10 KiB
Zig
//! Shared definitions that form the contract between a bootloader
|
|
//! (src/boot/, e.g. efi.zig built as BOOTX64.efi) and the kernel (src/kernel/main.zig).
|
|
//!
|
|
//! Both binaries import this as the "danos" module, so the handoff layout is
|
|
//! defined in exactly one place.
|
|
|
|
const std = @import("std");
|
|
|
|
/// Calling convention for the bootloader→kernel jump. Pinned to SysV so it does
|
|
/// not depend on each binary's target default: the UEFI bootloader's C
|
|
/// convention is Microsoft x64 (first arg in RCX), the freestanding kernel's is
|
|
/// SysV (first arg in RDI). Both reference this to agree on where `*BootInfo`
|
|
/// is passed.
|
|
pub const kernel_abi: std.builtin.CallingConvention = .{ .x86_64_sysv = .{} };
|
|
|
|
/// Pixel byte order of the linear framebuffer the firmware handed us.
|
|
pub const PixelFormat = enum(u32) {
|
|
/// Byte 0 = Red, 1 = Green, 2 = Blue, 3 = reserved.
|
|
rgbx,
|
|
/// Byte 0 = Blue, 1 = Green, 2 = Red, 3 = reserved.
|
|
bgrx,
|
|
};
|
|
|
|
/// A linear framebuffer: `width`x`height` pixels, each a 32-bit value, with
|
|
/// `pitch` bytes between the start of one row and the next (which may be larger
|
|
/// than `width * 4` due to hardware padding).
|
|
///
|
|
/// A `base` of 0 means **no framebuffer** — the firmware exposed no Graphics
|
|
/// Output Protocol (a headless server, say). The kernel must treat on-screen
|
|
/// output as optional and never assume a framebuffer exists.
|
|
pub const Framebuffer = extern struct {
|
|
base: usize, // the memory address where pixel data starts (0 = none)
|
|
width: u32, // visible pixels per row (e.g. 1920)
|
|
height: u32, // visible rows (e.g. 1080)
|
|
pitch: u32, // bytes from the start of one row to the start of the next
|
|
format: PixelFormat,
|
|
|
|
/// Whether a usable framebuffer was handed over.
|
|
pub fn present(self: Framebuffer) bool {
|
|
return self.base != 0 and self.width != 0 and self.height != 0;
|
|
}
|
|
};
|
|
|
|
/// Page size the memory map is measured in. 4 KiB on every architecture danos
|
|
/// targets so far.
|
|
pub const page_size = 4096;
|
|
|
|
/// The kernel's virtual-memory layout (higher-half). The kernel is linked at
|
|
/// `kernel_virt_base` but loaded at a low physical address; all of RAM (and the
|
|
/// device MMIO windows) is also mapped at `physmap_base + phys`, so the kernel
|
|
/// can reach any physical address by adding a constant. The low half is left
|
|
/// entirely to user space.
|
|
///
|
|
/// user image + stack : 0x0000_7000_0000_0000 (PML4[224], low half)
|
|
/// kernel heap : 0xFFFF_8000_0000_0000 (PML4[256])
|
|
/// physmap : 0xFFFF_8800_0000_0000 (PML4[272]) + phys
|
|
/// kernel image : 0xFFFF_FFFF_8000_0000 (PML4[511])
|
|
pub const physmap_base: u64 = 0xFFFF_8800_0000_0000;
|
|
pub const kernel_virt_base: u64 = 0xFFFF_FFFF_8000_0000;
|
|
|
|
/// The kernel syscall numbers — the single source of truth shared by the kernel
|
|
/// dispatcher (src/kernel/process.zig) and the user runtime library, so the two
|
|
/// can never drift. The set is deliberately microkernel-minimal: file/device I/O
|
|
/// is not here — it lives in user-space servers reached through the IPC calls.
|
|
/// The table grows one milestone at a time; see docs/syscall.md.
|
|
pub const Syscall = enum(u64) {
|
|
exit = 0, // exit(code): end the calling process
|
|
yield = 1, // yield(): give up the rest of this quantum
|
|
debug_write = 2, // debug_write(ptr, len): raw bytes to the kernel log (bring-up only)
|
|
sleep = 3, // sleep(ms): block the caller for ms milliseconds
|
|
mmap = 4, // mmap(len, prot) -> base: grant zeroed, page-aligned user pages
|
|
munmap = 5, // munmap(base, len): release pages from a prior mmap
|
|
create_endpoint = 6, // create_endpoint() -> handle: a new IPC endpoint
|
|
ipc_register = 7, // ipc_register(service_id, handle): publish an endpoint by well-known id
|
|
ipc_lookup = 8, // ipc_lookup(service_id) -> handle: find a published endpoint
|
|
ipc_call = 9, // ipc_call(h, msg, len, reply, cap) -> reply_len: send + block for reply
|
|
ipc_reply_wait = 10, // ipc_reply_wait(h, reply, len, recv, cap) -> recv_len (+badge in rdx)
|
|
dev_enumerate = 11, // dev_enumerate(buf, max) -> count: snapshot the device table
|
|
dev_claim = 12, // dev_claim(id) -> ok: take exclusive ownership of a device
|
|
mmio_map = 13, // mmio_map(id, res_idx) -> vaddr: map a claimed device's MMIO into this AS
|
|
irq_bind = 14, // irq_bind(id, res_idx, endpoint): deliver a device IRQ as an IPC notification
|
|
irq_ack = 15, // irq_ack(id, res_idx): re-arm a bound IRQ after servicing it
|
|
_,
|
|
};
|
|
|
|
/// A device class, mirroring src/device/device.zig's `DeviceClass` **in order**
|
|
/// (its `@intFromEnum` values cross the syscall boundary in `DeviceDesc.class`).
|
|
/// Keep the two in sync.
|
|
pub const DeviceClass = enum(u32) {
|
|
root,
|
|
processor,
|
|
interrupt_controller,
|
|
timer,
|
|
pci_host_bridge,
|
|
pci_device,
|
|
acpi_device,
|
|
unknown,
|
|
};
|
|
|
|
/// A resource kind, mirroring src/device/device.zig's `ResourceKind` in order.
|
|
pub const ResourceKind = enum(u32) {
|
|
memory,
|
|
io_port,
|
|
irq,
|
|
bus_range,
|
|
};
|
|
|
|
/// One device resource, as handed to a user-space driver (flat, extern).
|
|
pub const ResDesc = extern struct {
|
|
kind: u64, // a ResourceKind value
|
|
start: u64,
|
|
len: u64,
|
|
};
|
|
|
|
pub const max_dev_resources = 8;
|
|
|
|
/// A device, as snapshotted for user space by `dev_enumerate`. A driver scans
|
|
/// these to find the hardware it owns, claims it, and maps its MMIO.
|
|
pub const DeviceDesc = extern struct {
|
|
id: u64,
|
|
class: u64, // a DeviceClass value
|
|
hid_len: u64,
|
|
resource_count: u64,
|
|
hid: [8]u8,
|
|
resources: [max_dev_resources]ResDesc,
|
|
};
|
|
|
|
/// Well-known IPC service ids for the bootstrap name registry (create_endpoint +
|
|
/// ipc_register/ipc_lookup). Small integers, so no string interning is needed
|
|
/// during bring-up. The VFS server registers under `vfs`; clients look it up.
|
|
pub const ServiceId = enum(u32) {
|
|
vfs = 1,
|
|
_,
|
|
};
|
|
|
|
/// Protection flags for `mmap` (matching the usual C bit values).
|
|
pub const prot_read: u64 = 1;
|
|
pub const prot_write: u64 = 2;
|
|
pub const prot_exec: u64 = 4;
|
|
|
|
/// Physical address -> its virtual address in the physmap. The single way the
|
|
/// kernel dereferences a physical address once paging is up.
|
|
///
|
|
/// **Hazard:** valid only once the (bootstrap or final) page tables are live.
|
|
/// The bootloader may use the *constant* `physmap_base` to build those tables,
|
|
/// but must not call this to dereference memory before its own CR3 is loaded —
|
|
/// it runs under the firmware's identity map, where these addresses are unmapped.
|
|
pub inline fn physToVirt(phys: u64) u64 {
|
|
return phys + physmap_base;
|
|
}
|
|
|
|
/// Physmap virtual address -> physical. Inverse of `physToVirt`; for producing
|
|
/// the physical address of something the kernel holds a physmap pointer to
|
|
/// (e.g. a page-table frame for CR3, a post-mortem breadcrumb's RAM location).
|
|
pub inline fn virtToPhys(virt: u64) u64 {
|
|
return virt - physmap_base;
|
|
}
|
|
|
|
/// danos's own classification of a span of physical memory — deliberately not
|
|
/// UEFI's vocabulary. Each boot path (UEFI now, device tree later) translates its
|
|
/// native memory description into these kinds, so the kernel never learns what
|
|
/// booted it. [[arch]] keeps the same discipline for CPU code.
|
|
pub const MemoryKind = enum(u32) {
|
|
/// Free RAM the kernel may allocate. Each boot path folds its own transient
|
|
/// memory into this once it's genuinely free (e.g. the UEFI loader classifies
|
|
/// boot-services memory as usable after ExitBootServices), so the kernel never
|
|
/// has to know about boot-protocol-specific "reclaimable" states.
|
|
usable,
|
|
/// Firmware, MMIO, the kernel image, our own boot buffers, the boot stack —
|
|
/// never hand out.
|
|
reserved,
|
|
/// ACPI tables: parse, then reclaim.
|
|
acpi_tables,
|
|
/// ACPI non-volatile storage: preserve across sleep, do not allocate.
|
|
acpi_nvs,
|
|
/// Not backed by RAM: memory-mapped device registers or a reserved
|
|
/// address-space window (e.g. PCIe config space). Kept distinct from
|
|
/// `reserved` so RAM accounting doesn't count device address space.
|
|
mmio,
|
|
};
|
|
|
|
/// One contiguous span of physical memory. Because danos defines this layout
|
|
/// itself (unlike the UEFI descriptor it's built from), `@sizeOf` is
|
|
/// authoritative — the kernel walks a plain `[]MemoryRegion`, with none of the
|
|
/// firmware's variable descriptor-stride to worry about.
|
|
pub const MemoryRegion = extern struct {
|
|
base: u64, // physical start address
|
|
pages: u64, // length in `page_size` units
|
|
kind: MemoryKind,
|
|
_pad: u32 = 0,
|
|
};
|
|
|
|
/// The physical memory layout handed to the kernel: a pointer to an array of
|
|
/// `len` `MemoryRegion`s, in a buffer that outlives the loader.
|
|
pub const MemoryMap = extern struct {
|
|
regions: usize, // address of a `[len]MemoryRegion`
|
|
len: usize,
|
|
};
|
|
|
|
/// One PT_LOAD segment of the kernel image, so the kernel can re-map itself with
|
|
/// correct permissions (code R+X, rodata R, data R+W+NX). `flags` are raw ELF
|
|
/// segment flags: PF_X=1, PF_W=2, PF_R=4. `virt` is the higher-half link address;
|
|
/// `phys` is where the loader actually placed the segment (they differ once the
|
|
/// kernel links high — the loader records the real load address here).
|
|
pub const KernelSegment = extern struct {
|
|
virt: u64,
|
|
phys: u64,
|
|
pages: u64,
|
|
flags: u32,
|
|
_pad: u32 = 0,
|
|
};
|
|
|
|
/// Handoff structure the bootloader fills in and passes to the kernel's
|
|
/// `_start` in RDI (the first argument under the SysV AMD64 C ABI).
|
|
pub const BootInfo = extern struct {
|
|
framebuffer: Framebuffer,
|
|
memory_map: MemoryMap,
|
|
/// The kernel's own PT_LOAD segments (it has three: text, rodata, data).
|
|
kernel_segments: [8]KernelSegment,
|
|
kernel_segment_count: u32,
|
|
/// Physical address of the ACPI RSDP the firmware exposed, or 0 if none. The
|
|
/// kernel's device layer parses the ACPI tables from here to discover hardware.
|
|
/// A device-tree boot path leaves this 0 and (later) fills a `device_tree_blob`
|
|
/// field instead, so the kernel discovers devices without knowing what booted it.
|
|
acpi_rsdp: u64 = 0,
|
|
/// The raw `/sbin/init` ELF image, read off the boot volume by the loader
|
|
/// into memory that survives the handoff (classified reserved, so the kernel
|
|
/// identity-maps it and never allocates over it). 0/0 = no init found — the
|
|
/// kernel boots without user space. Grows into a full initrd handoff later.
|
|
init_base: u64 = 0,
|
|
init_len: u64 = 0,
|
|
/// The initrd image (a bundle of extra user binaries — the VFS server and
|
|
/// device drivers), read off the boot volume into memory that survives the
|
|
/// handoff, same as `init` above. 0/0 = no initrd. See src/user/proto/initrd.zig.
|
|
initrd_base: u64 = 0,
|
|
initrd_len: u64 = 0,
|
|
};
|