M11–M12: IRQ-as-IPC and bus drivers; expand names tree-wide
Two driver-model milestones plus a tree-wide naming pass. Suite 35/35 (QEMU) + host tests green. M11 — IRQ-as-IPC. A ring-3 driver now sleeps until its device interrupts it. New src/kernel/irq.zig: per-GSI endpoint bindings, comptime per-vector trampolines, dispatch = mask GSI -> LAPIC EOI -> notifyLocked, all under one lock region. irq_bind/irq_ack syscalls, gated by the device claim like mmio_map. interruptDispatch no longer EOIs — each handler owns its EOI, because a level line must be masked before it is acknowledged (irq_ack is the unmask). Bindings are keyed on the owning task and released on exit (a shared endpoint's siblings survive). hpetd rewritten interrupt-driven. Tests: hpet (rewritten, reads back the I/O APIC routing) and irqfree. M12 — bus drivers. DeviceDesc gains a parent, making the device table a tree. dev_register (device_register) lets a process publish children below a device it claimed; the kernel enforces resource containment (a child's resources must nest in its parent's), so a descriptor can't fabricate a window over kernel RAM. Descriptor copied in via copyFromUser (physmap walk — an unmapped user pointer fails the call instead of faulting the kernel). Per-parent child cap bounds table exhaustion. sbin/busd.zig is a worked bus driver. Test: bus. Naming — per docs/coding-standards.md: non-acronym abbreviations spelled out (message, descriptor, device_service, scheduler, runtime, physical, interpreter, ...); acronyms kept (IPC, MMIO, DMA, HCD, ...); files are kebab-case (ipc-synchronous.zig, device-service.zig, vfs-protocol.zig, ...). Exceptions: POSIX/C ABI names and Zig idioms (init/len/ptr) kept. Module collisions resolved by specific naming (config -> parameters, device.zig alias -> device_model). AML op/Op disambiguated: op = opcode, Op = operation; per-opcode parse handlers renamed opX -> parseX. New driver docs: drivers.md, driver-model.md (bus/class/HCD shapes + the proposed M13–M16 ABI), coding-standards.md.
This commit is contained in:
+54
-35
@@ -6,10 +6,10 @@
|
||||
|
||||
const std = @import("std");
|
||||
|
||||
/// Calling convention for the bootloader→kernel jump. Pinned to SysV so it does
|
||||
/// Calling convention for the bootloader→kernel jump. Pinned to SystemV so it does
|
||||
/// not depend on each binary's target default: the UEFI bootloader's C
|
||||
/// convention is Microsoft x64 (first arg in RCX), the freestanding kernel's is
|
||||
/// SysV (first arg in RDI). Both reference this to agree on where `*BootInfo`
|
||||
/// SystemV (first arg in RDI). Both reference this to agree on where `*BootInformation`
|
||||
/// is passed.
|
||||
pub const kernel_abi: std.builtin.CallingConvention = .{ .x86_64_sysv = .{} };
|
||||
|
||||
@@ -47,23 +47,23 @@ pub const page_size = 4096;
|
||||
|
||||
/// The kernel's virtual-memory layout (higher-half). The kernel is linked at
|
||||
/// `kernel_virt_base` but loaded at a low physical address; all of RAM (and the
|
||||
/// device MMIO windows) is also mapped at `physmap_base + phys`, so the kernel
|
||||
/// device MMIO windows) is also mapped at `physmap_base + physical`, so the kernel
|
||||
/// can reach any physical address by adding a constant. The low half is left
|
||||
/// entirely to user space.
|
||||
///
|
||||
/// user image + stack : 0x0000_7000_0000_0000 (PML4[224], low half)
|
||||
/// kernel heap : 0xFFFF_8000_0000_0000 (PML4[256])
|
||||
/// physmap : 0xFFFF_8800_0000_0000 (PML4[272]) + phys
|
||||
/// physmap : 0xFFFF_8800_0000_0000 (PML4[272]) + physical
|
||||
/// kernel image : 0xFFFF_FFFF_8000_0000 (PML4[511])
|
||||
pub const physmap_base: u64 = 0xFFFF_8800_0000_0000;
|
||||
pub const kernel_virt_base: u64 = 0xFFFF_FFFF_8000_0000;
|
||||
|
||||
/// The kernel syscall numbers — the single source of truth shared by the kernel
|
||||
/// The kernel system_call numbers — the single source of truth shared by the kernel
|
||||
/// dispatcher (src/kernel/process.zig) and the user runtime library, so the two
|
||||
/// can never drift. The set is deliberately microkernel-minimal: file/device I/O
|
||||
/// is not here — it lives in user-space servers reached through the IPC calls.
|
||||
/// The table grows one milestone at a time; see docs/syscall.md.
|
||||
pub const Syscall = enum(u64) {
|
||||
pub const SystemCall = enum(u64) {
|
||||
exit = 0, // exit(code): end the calling process
|
||||
yield = 1, // yield(): give up the rest of this quantum
|
||||
debug_write = 2, // debug_write(ptr, len): raw bytes to the kernel log (bring-up only)
|
||||
@@ -73,18 +73,26 @@ pub const Syscall = enum(u64) {
|
||||
create_endpoint = 6, // create_endpoint() -> handle: a new IPC endpoint
|
||||
ipc_register = 7, // ipc_register(service_id, handle): publish an endpoint by well-known id
|
||||
ipc_lookup = 8, // ipc_lookup(service_id) -> handle: find a published endpoint
|
||||
ipc_call = 9, // ipc_call(h, msg, len, reply, cap) -> reply_len: send + block for reply
|
||||
ipc_reply_wait = 10, // ipc_reply_wait(h, reply, len, recv, cap) -> recv_len (+badge in rdx)
|
||||
dev_enumerate = 11, // dev_enumerate(buf, max) -> count: snapshot the device table
|
||||
dev_claim = 12, // dev_claim(id) -> ok: take exclusive ownership of a device
|
||||
mmio_map = 13, // mmio_map(id, res_idx) -> vaddr: map a claimed device's MMIO into this AS
|
||||
irq_bind = 14, // irq_bind(id, res_idx, endpoint): deliver a device IRQ as an IPC notification
|
||||
irq_ack = 15, // irq_ack(id, res_idx): re-arm a bound IRQ after servicing it
|
||||
ipc_call = 9, // ipc_call(h, message, len, reply, cap) -> reply_len: send + block for reply
|
||||
ipc_reply_wait = 10, // ipc_reply_wait(h, reply, len, receive, cap) -> receive_len (+badge in rdx)
|
||||
device_enumerate = 11, // device_enumerate(buffer, maximum) -> count: snapshot the device table
|
||||
device_claim = 12, // device_claim(id) -> ok: take exclusive ownership of a device
|
||||
mmio_map = 13, // mmio_map(id, resource_index) -> vaddr: map a claimed device's MMIO into this AS
|
||||
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
|
||||
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
|
||||
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
|
||||
_,
|
||||
};
|
||||
|
||||
/// A device class, mirroring src/device/device.zig's `DeviceClass` **in order**
|
||||
/// (its `@intFromEnum` values cross the syscall boundary in `DeviceDesc.class`).
|
||||
/// Set in the badge returned by `ipc_reply_wait` when what arrived is an
|
||||
/// **asynchronous notification** (today: a device interrupt bound with `irq_bind`)
|
||||
/// rather than a message from a client. There is no payload and no reply owed; the
|
||||
/// low bits carry the source, a GSI. Shared so the kernel's ISR and the driver's
|
||||
/// event loop can't disagree about which bit means "the hardware spoke".
|
||||
pub const notify_badge_bit: u64 = 1 << 63;
|
||||
|
||||
/// A device class, mirroring src/device/device-model.zig's `DeviceClass` **in order**
|
||||
/// (its `@intFromEnum` values cross the system_call boundary in `DeviceDescriptor.class`).
|
||||
/// Keep the two in sync.
|
||||
pub const DeviceClass = enum(u32) {
|
||||
root,
|
||||
@@ -97,7 +105,7 @@ pub const DeviceClass = enum(u32) {
|
||||
unknown,
|
||||
};
|
||||
|
||||
/// A resource kind, mirroring src/device/device.zig's `ResourceKind` in order.
|
||||
/// A resource kind, mirroring src/device/device-model.zig's `ResourceKind` in order.
|
||||
pub const ResourceKind = enum(u32) {
|
||||
memory,
|
||||
io_port,
|
||||
@@ -106,23 +114,34 @@ pub const ResourceKind = enum(u32) {
|
||||
};
|
||||
|
||||
/// One device resource, as handed to a user-space driver (flat, extern).
|
||||
pub const ResDesc = extern struct {
|
||||
pub const ResourceDescriptor = extern struct {
|
||||
kind: u64, // a ResourceKind value
|
||||
start: u64,
|
||||
len: u64,
|
||||
};
|
||||
|
||||
pub const max_dev_resources = 8;
|
||||
pub const maximum_device_resources = 8;
|
||||
|
||||
/// A device, as snapshotted for user space by `dev_enumerate`. A driver scans
|
||||
/// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree.
|
||||
pub const no_parent: u64 = ~@as(u64, 0);
|
||||
|
||||
/// A device, as snapshotted for user space by `device_enumerate`. A driver scans
|
||||
/// these to find the hardware it owns, claims it, and maps its MMIO.
|
||||
pub const DeviceDesc = extern struct {
|
||||
///
|
||||
/// `parent` makes the table a tree rather than a list, which is what a **bus driver**
|
||||
/// needs: it claims the bus, finds the devices below it, and publishes any it
|
||||
/// discovers itself with `device_register`. A registered child's resources must lie
|
||||
/// within its parent's (the kernel enforces this) — that containment is what makes
|
||||
/// delegation safe, since a device descriptor is otherwise a licence to map physical
|
||||
/// memory.
|
||||
pub const DeviceDescriptor = extern struct {
|
||||
id: u64,
|
||||
parent: u64, // a device id, or `no_parent`
|
||||
class: u64, // a DeviceClass value
|
||||
hid_len: u64,
|
||||
resource_count: u64,
|
||||
hid: [8]u8,
|
||||
resources: [max_dev_resources]ResDesc,
|
||||
resources: [maximum_device_resources]ResourceDescriptor,
|
||||
};
|
||||
|
||||
/// Well-known IPC service ids for the bootstrap name registry (create_endpoint +
|
||||
@@ -145,21 +164,21 @@ pub const prot_exec: u64 = 4;
|
||||
/// The bootloader may use the *constant* `physmap_base` to build those tables,
|
||||
/// but must not call this to dereference memory before its own CR3 is loaded —
|
||||
/// it runs under the firmware's identity map, where these addresses are unmapped.
|
||||
pub inline fn physToVirt(phys: u64) u64 {
|
||||
return phys + physmap_base;
|
||||
pub inline fn physicalToVirtual(physical: u64) u64 {
|
||||
return physical + physmap_base;
|
||||
}
|
||||
|
||||
/// Physmap virtual address -> physical. Inverse of `physToVirt`; for producing
|
||||
/// Physmap virtual address -> physical. Inverse of `physicalToVirtual`; for producing
|
||||
/// the physical address of something the kernel holds a physmap pointer to
|
||||
/// (e.g. a page-table frame for CR3, a post-mortem breadcrumb's RAM location).
|
||||
pub inline fn virtToPhys(virt: u64) u64 {
|
||||
return virt - physmap_base;
|
||||
pub inline fn virtualToPhysical(virtual: u64) u64 {
|
||||
return virtual - physmap_base;
|
||||
}
|
||||
|
||||
/// danos's own classification of a span of physical memory — deliberately not
|
||||
/// UEFI's vocabulary. Each boot path (UEFI now, device tree later) translates its
|
||||
/// native memory description into these kinds, so the kernel never learns what
|
||||
/// booted it. [[arch]] keeps the same discipline for CPU code.
|
||||
/// booted it. [[architecture]] keeps the same discipline for CPU code.
|
||||
pub const MemoryKind = enum(u32) {
|
||||
/// Free RAM the kernel may allocate. Each boot path folds its own transient
|
||||
/// memory into this once it's genuinely free (e.g. the UEFI loader classifies
|
||||
@@ -174,7 +193,7 @@ pub const MemoryKind = enum(u32) {
|
||||
/// ACPI non-volatile storage: preserve across sleep, do not allocate.
|
||||
acpi_nvs,
|
||||
/// Not backed by RAM: memory-mapped device registers or a reserved
|
||||
/// address-space window (e.g. PCIe config space). Kept distinct from
|
||||
/// address-space window (e.g. PCIe configuration space). Kept distinct from
|
||||
/// `reserved` so RAM accounting doesn't count device address space.
|
||||
mmio,
|
||||
};
|
||||
@@ -199,20 +218,20 @@ pub const MemoryMap = extern struct {
|
||||
|
||||
/// One PT_LOAD segment of the kernel image, so the kernel can re-map itself with
|
||||
/// correct permissions (code R+X, rodata R, data R+W+NX). `flags` are raw ELF
|
||||
/// segment flags: PF_X=1, PF_W=2, PF_R=4. `virt` is the higher-half link address;
|
||||
/// `phys` is where the loader actually placed the segment (they differ once the
|
||||
/// segment flags: PF_X=1, PF_W=2, PF_R=4. `virtual` is the higher-half link address;
|
||||
/// `physical` is where the loader actually placed the segment (they differ once the
|
||||
/// kernel links high — the loader records the real load address here).
|
||||
pub const KernelSegment = extern struct {
|
||||
virt: u64,
|
||||
phys: u64,
|
||||
virtual: u64,
|
||||
physical: u64,
|
||||
pages: u64,
|
||||
flags: u32,
|
||||
_pad: u32 = 0,
|
||||
};
|
||||
|
||||
/// Handoff structure the bootloader fills in and passes to the kernel's
|
||||
/// `_start` in RDI (the first argument under the SysV AMD64 C ABI).
|
||||
pub const BootInfo = extern struct {
|
||||
/// `_start` in RDI (the first argument under the SystemV AMD64 C ABI).
|
||||
pub const BootInformation = extern struct {
|
||||
framebuffer: Framebuffer,
|
||||
memory_map: MemoryMap,
|
||||
/// The kernel's own PT_LOAD segments (it has three: text, rodata, data).
|
||||
@@ -231,7 +250,7 @@ pub const BootInfo = extern struct {
|
||||
init_len: u64 = 0,
|
||||
/// The initrd image (a bundle of extra user binaries — the VFS server and
|
||||
/// device drivers), read off the boot volume into memory that survives the
|
||||
/// handoff, same as `init` above. 0/0 = no initrd. See src/user/proto/initrd.zig.
|
||||
/// handoff, same as `init` above. 0/0 = no initrd. See src/user/protocol/initrd.zig.
|
||||
initrd_base: u64 = 0,
|
||||
initrd_len: u64 = 0,
|
||||
};
|
||||
|
||||
Reference in New Issue
Block a user