Post-reorg cleanup: POSIX layer, and naming fixes

Follow-up to the monorepo re-org. Suite 35/35 plus host tests green.

POSIX compatibility is now its own library, library/posix/ (unistd, stdio),
layered strictly over the runtime — it calls the runtime's IPC/heap, never
system calls directly. The runtime is now POSIX-free (the danos-native
application ABI). The VFS wire protocol is danos-native throughout
(Stat -> FileStatus, .stat -> .status, O_CREAT -> create); the POSIX layer
maps the POSIX spellings at the boundary. The coding standard's ABI-name
exception is scoped to one place: a file is allowed POSIX spellings only if it
lives under library/posix/ — everywhere else, danos naming with no exception.

Naming fixes, all mechanical:
- initrd -> initial-ramdisk: the source file, the module, the tool
  (make-initial-ramdisk.py), the artifact (initial-ramdisk.img, including the
  bootloader's load path), and the identifiers.
- system/kernel/device-service.zig -> devices-broker.zig: it is ring-0 kernel
  code (the trusted device table + claim capability), not a ring-3 service. The
  future user-space device *manager* (policy) will live in system/services/.
- Dropped the daemon `d` suffix: hpetd -> hpet, busd -> bus. A driver lives in
  system/drivers/, so the folder already says what it is; encoding the role in
  the name too is redundant. The coding standard drops that exception.
- system/devices/aml/interp.zig -> interpreter.zig (the type was already
  Interpreter).
This commit is contained in:
Daniel Samson
2026-07-10 13:33:06 +01:00
parent 8754d4e46a
commit ceacc6b514
27 changed files with 308 additions and 254 deletions
+5 -5
View File
@@ -245,12 +245,12 @@ pub const BootInformation = extern struct {
/// The raw `/sbin/init` ELF image, read off the boot volume by the loader
/// into memory that survives the handoff (classified reserved, so the kernel
/// identity-maps it and never allocates over it). 0/0 = no init found — the
/// kernel boots without user space. Grows into a full initrd handoff later.
/// kernel boots without user space. Grows into a full initial_ramdisk handoff later.
init_base: u64 = 0,
init_len: u64 = 0,
/// The initrd image (a bundle of extra user binaries — the VFS server and
/// The initial_ramdisk image (a bundle of extra user binaries — the VFS server and
/// device drivers), read off the boot volume into memory that survives the
/// handoff, same as `init` above. 0/0 = no initrd. See system/initrd.zig.
initrd_base: u64 = 0,
initrd_len: u64 = 0,
/// handoff, same as `init` above. 0/0 = no initial_ramdisk. See system/initial-ramdisk.zig.
initial_ramdisk_base: u64 = 0,
initial_ramdisk_len: u64 = 0,
};
+5 -5
View File
@@ -3,7 +3,7 @@
//!
//! This module has two stages. `parser.zig` walks the entire byte stream and
//! records every named object into a namespace tree (`namespace.zig`), capturing
//! method bodies and field/region layout. `interp.zig` then *evaluates* control
//! method bodies and field/region layout. `interpreter.zig` then *evaluates* control
//! methods on demand — running operators, control flow, and OperationRegion field
//! access — so callers can resolve device status (`_STA`), current resource
//! settings (`_CRS`), sleep states (`_Sx`), and the like against the live namespace.
@@ -17,10 +17,10 @@ pub const Node = @import("namespace.zig").Node;
pub const NodeKind = @import("namespace.zig").NodeKind;
/// The AML evaluator: interprets control methods (and reads Names/Fields) far
/// enough for device discovery. See `interp.zig`.
pub const Interpreter = @import("interp.zig").Interpreter;
pub const Object = @import("interp.zig").Object;
pub const EvaluateHal = @import("interp.zig").Hal;
/// enough for device discovery. See `interpreter.zig`.
pub const Interpreter = @import("interpreter.zig").Interpreter;
pub const Object = @import("interpreter.zig").Object;
pub const EvaluateHal = @import("interpreter.zig").Hal;
/// The SLP_TYP values written to PM1a/PM1b control to enter a sleep state.
pub const SleepType = struct {
@@ -1,4 +1,4 @@
//! /sbin/busd — a user-space **bus driver**, and the smallest honest example of one.
//! /sbin/bus — a user-space **bus driver**, and the smallest honest example of one.
//!
//! A bus driver owns a device that *contains other devices*, enumerates them by some
//! bus-specific protocol, and publishes each one into the kernel's device table so a
@@ -6,10 +6,10 @@
//! the "bus" is the HPET's register block and the "devices" are its comparators, each
//! a 0x20-byte window at 0x100 + 0x20*n that can be driven independently.
//!
//! It's a toy bus, but nothing about the mechanism is: `busd` reads how many children
//! It's a toy bus, but nothing about the mechanism is: `bus` reads how many children
//! exist from the hardware (GENERAL_CAP bits [12:8]), publishes one `DeviceDescriptor` per
//! child with a sub-window of its own MMIO plus the shared IRQ, and the kernel checks
//! every one of those resources is contained in what `busd` was granted. A comparator
//! every one of those resources is contained in what `bus` was granted. A comparator
//! driver then claims a child and maps only *its* registers — not the whole block.
//!
//! It also proves the negative: registering a child whose window escapes the parent's
@@ -65,30 +65,30 @@ fn firstChildOf(buffer: []device.DeviceDescriptor, total: usize, parent_id: u64)
pub fn main() void {
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("busd: out of memory\n");
_ = runtime.system.write("bus: out of memory\n");
return;
};
const parent = findHpet(buffer) orelse {
_ = runtime.system.write("busd: no HPET\n");
_ = runtime.system.write("bus: no HPET\n");
return;
};
const resource = resourcesOf(parent);
// Claim the bus. Everything below is subdivision of what this claim granted.
//
// Claims are exclusive, and at a normal boot the kernel spawns every initrd
// binary — so hpetd may own the HPET already. That's not an error, it's the
// Claims are exclusive, and at a normal boot the kernel spawns every initial_ramdisk
// binary — so hpet may own the HPET already. That's not an error, it's the
// capability model working: exit quietly and leave the device to its owner. The
// `bus` test spawns busd alone, so there it wins the claim.
// `bus` test spawns bus alone, so there it wins the claim.
if (!device.claim(parent.id)) {
_ = runtime.system.write("busd: HPET already claimed by another driver, nothing to do\n");
_ = runtime.system.write("bus: HPET already claimed by another driver, nothing to do\n");
return;
}
// Enumerate the bus: ask the hardware how many children it has.
const base = device.mmioMap(parent.id, 0) orelse {
_ = runtime.system.write("busd: mmio_map failed\n");
_ = runtime.system.write("bus: mmio_map failed\n");
return;
};
const cap: *volatile u64 = @ptrFromInt(base + register_general_cap);
@@ -112,7 +112,7 @@ pub fn main() void {
}
if (device.register(parent.id, &child) == null) {
_ = runtime.system.write("busd: register failed\n");
_ = runtime.system.write("bus: register failed\n");
return;
}
published += 1;
@@ -133,11 +133,11 @@ pub fn main() void {
.len = 0x1000,
};
if (device.register(parent.id, &rogue) != null) {
_ = runtime.system.write("busd: FAIL out-of-window child was accepted\n");
_ = runtime.system.write("bus: FAIL out-of-window child was accepted\n");
return;
}
if (device.enumerate(buffer) != before) {
_ = runtime.system.write("busd: FAIL rogue child leaked into the table\n");
_ = runtime.system.write("bus: FAIL rogue child leaked into the table\n");
return;
}
@@ -149,18 +149,18 @@ pub fn main() void {
if (d.parent != parent.id) continue;
const w = d.resources[0];
if (w.start < resource.mmio.start or w.len >= resource.mmio.len) {
_ = runtime.system.write("busd: FAIL child window is not inside the bus\n");
_ = runtime.system.write("bus: FAIL child window is not inside the bus\n");
return;
}
seen += 1;
}
if (seen != published) {
_ = runtime.system.write("busd: FAIL child count mismatch\n");
_ = runtime.system.write("bus: FAIL child count mismatch\n");
return;
}
// Delegation, end to end: claim a child and map *it*. A real class driver would be
// a different process; here busd plays both parts, which exercises the same path.
// a different process; here bus plays both parts, which exercises the same path.
// The child's window is 0x20 bytes at parent+0x100, so the register it sees at
// offset 0 must be the same timer-0 configuration register the bus sees at 0x100.
//
@@ -168,21 +168,21 @@ pub fn main() void {
// 4 KiB the HPET lives in — the granularity limit documented in docs/drivers.md.
// The *resource* is narrow even though the page isn't.)
const child_id = firstChildOf(buffer, device.enumerate(buffer), parent.id) orelse {
_ = runtime.system.write("busd: FAIL no child to claim\n");
_ = runtime.system.write("bus: FAIL no child to claim\n");
return;
};
if (!device.claim(child_id)) {
_ = runtime.system.write("busd: FAIL could not claim own child\n");
_ = runtime.system.write("bus: FAIL could not claim own child\n");
return;
}
const child_base = device.mmioMap(child_id, 0) orelse {
_ = runtime.system.write("busd: FAIL child mmio_map refused\n");
_ = runtime.system.write("bus: FAIL child mmio_map refused\n");
return;
};
const via_child: *volatile u64 = @ptrFromInt(child_base);
const via_bus: *volatile u64 = @ptrFromInt(base + 0x100);
if (via_child.* != via_bus.*) {
_ = runtime.system.write("busd: FAIL child window does not alias the bus register\n");
_ = runtime.system.write("bus: FAIL child window does not alias the bus register\n");
return;
}
@@ -195,12 +195,12 @@ pub fn main() void {
_ = runtime.system.munmap(scratch, 0x1000);
const descriptor: *const device.DeviceDescriptor = @ptrFromInt(scratch);
if (device.register(parent.id, descriptor) != null) {
_ = runtime.system.write("busd: FAIL register accepted an unmapped descriptor\n");
_ = runtime.system.write("bus: FAIL register accepted an unmapped descriptor\n");
return;
}
}
_ = runtime.system.write("busd: ok\n");
_ = runtime.system.write("bus: ok\n");
while (true) runtime.system.sleep(1000);
}
@@ -1,4 +1,4 @@
//! /sbin/hpetd — a user-space HPET driver. It proves the whole driver model end to
//! /sbin/hpet — a user-space HPET driver. It proves the whole driver model end to
//! end: enumerate the device table, find the HPET, claim it, map its registers into
//! this ring-3 address space (strong-uncacheable), **bind its interrupt to an IPC
//! endpoint**, then sit blocked in `replyWait` until the hardware wakes it.
@@ -13,8 +13,8 @@
//! the full cycle to be correct:
//!
//! kernel ISR mask the GSI -> EOI -> notify this endpoint
//! hpetd wake, clear GENERAL_INT_STATUS (deasserts the line), re-arm
//! hpetd irq_ack -> kernel unmasks the GSI
//! hpet wake, clear GENERAL_INT_STATUS (deasserts the line), re-arm
//! hpet irq_ack -> kernel unmasks the GSI
//!
//! Clear the status bit *before* acking, or the line is still asserted when the
//! kernel unmasks and the I/O APIC redelivers forever.
@@ -64,7 +64,7 @@ fn findHpet(buffer: []device.DeviceDescriptor) ?Found {
for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(device.DeviceClass.timer)) continue;
// Skip comparator children a bus driver may have published below the block
// (see system/drivers/busd/busd.zig) — we want the register block itself.
// (see system/drivers/bus/bus.zig) — we want the register block itself.
if (d.parent != device.no_parent) continue;
var mmio: ?u64 = null;
var irq: ?u64 = null;
@@ -85,21 +85,21 @@ fn findHpet(buffer: []device.DeviceDescriptor) ?Found {
pub fn main() void {
// Enumerate into a heap buffer (too big for the one-page user stack).
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 32) catch {
_ = runtime.system.write("hpetd: out of memory\n");
_ = runtime.system.write("hpet: out of memory\n");
return;
};
const hpet = findHpet(buffer) orelse {
_ = runtime.system.write("hpetd: no HPET with an IRQ\n");
_ = runtime.system.write("hpet: no HPET with an IRQ\n");
return;
};
if (!device.claim(hpet.device_id)) {
_ = runtime.system.write("hpetd: claim failed\n");
_ = runtime.system.write("hpet: claim failed\n");
return;
}
const base = device.mmioMap(hpet.device_id, hpet.mmio) orelse {
_ = runtime.system.write("hpetd: mmio_map failed\n");
_ = runtime.system.write("hpet: mmio_map failed\n");
return;
};
@@ -108,7 +108,7 @@ pub fn main() void {
const gsi = hpet.gsi;
const endpoint = ipc.createEndpoint() orelse {
_ = runtime.system.write("hpetd: create_endpoint failed\n");
_ = runtime.system.write("hpet: create_endpoint failed\n");
return;
};
@@ -116,7 +116,7 @@ pub fn main() void {
// Counter period, so we can arm the comparator a fixed wall-clock distance out.
const femtos_per_tick = register(base, register_general_cap).* >> 32;
if (femtos_per_tick == 0) {
_ = runtime.system.write("hpetd: bad HPET period\n");
_ = runtime.system.write("hpet: bad HPET period\n");
return;
}
const ticks_per_ms = 1_000_000_000_000 / femtos_per_tick;
@@ -139,10 +139,10 @@ pub fn main() void {
register(base, register_general_configuration).* |= configuration_enable;
if (!device.irqBind(hpet.device_id, hpet.irq, endpoint)) {
_ = runtime.system.write("hpetd: irq_bind failed\n");
_ = runtime.system.write("hpet: irq_bind failed\n");
return;
}
_ = runtime.system.write("hpetd: bound, sleeping until the hardware speaks\n");
_ = runtime.system.write("hpet: bound, sleeping until the hardware speaks\n");
// --- the driver loop -----------------------------------------------------
// Blocked in replyWait. No polling, no spinning: the next line of this function
@@ -170,14 +170,14 @@ pub fn main() void {
register(base, register_timer0_configuration).* &= ~tn_int_enb;
}
_ = runtime.system.write("hpetd: irq\n");
_ = runtime.system.write("hpet: irq\n");
if (!device.irqAck(hpet.device_id, hpet.irq)) {
_ = runtime.system.write("hpetd: irq_ack failed\n");
_ = runtime.system.write("hpet: irq_ack failed\n");
return;
}
}
_ = runtime.system.write("hpetd: ok\n");
_ = runtime.system.write("hpet: ok\n");
while (true) runtime.system.sleep(1000);
}
@@ -1,5 +1,5 @@
//! The initrd (initial ramdisk) container format — shared by the build-time
//! packer (tools/mkinitrd.zig) and the kernel that unpacks it. Deliberately
//! The initial_ramdisk (initial ramdisk) container format — shared by the build-time
//! packer (tools/make-initial-ramdisk.py) and the kernel that unpacks it. Deliberately
//! trivial: a header, a table of fixed-size entries, then the concatenated file
//! blobs. We own both producer and consumer, so it need be no fancier.
//!
@@ -10,7 +10,7 @@
const std = @import("std");
/// "DNRD" — identifies a danos initrd image.
/// "DNRD" — identifies a danos initial_ramdisk image.
pub const magic: u32 = 0x444E5244;
pub const Header = extern struct {
@@ -24,7 +24,7 @@ pub const Entry = extern struct {
len: u64, // blob length in bytes
};
/// A validated view over an initrd image. `init` checks the magic and that the
/// A validated view over an initial_ramdisk image. `init` checks the magic and that the
/// entry table fits; `entry` bounds-checks each blob against the image.
pub const Reader = struct {
image: []const u8,
+2 -2
View File
@@ -22,7 +22,7 @@
//!
//! Binding is capability-gated exactly like `mmio_map`: the caller must have
//! `device_claim`ed the device, and the GSI must come from one of that device's `irq`
//! resources in the discovered device table (system/kernel/device-service.zig). A driver can
//! resources in the discovered device table (system/kernel/devices-broker.zig). A driver can
//! therefore never bind an interrupt it doesn't own — a raw-GSI system_call would let
//! any process steal the keyboard's line.
//!
@@ -152,7 +152,7 @@ pub fn bind(gsi: u32, endpoint: *ipc_sync.Endpoint, owner: u32) BindError!void {
bound_owner[gsi] = owner;
// Level-triggered, active-high. Level is the general case a driver must survive
// (and what hpetd configures its comparator for); an edge source simply never
// (and what hpet configures its comparator for); an edge source simply never
// leaves the line asserted, so the mask/ack cycle is harmless there.
//
// Hardcoded for now: a device whose MADT interrupt-source override declares the
+15 -15
View File
@@ -8,9 +8,9 @@ const pmm = @import("pmm.zig");
const heap = @import("heap.zig");
const scheduler = @import("scheduler.zig");
const process = @import("process.zig");
const device_service = @import("device-service.zig");
const devices_broker = @import("devices-broker.zig");
const irq = @import("irq.zig");
const initrd = @import("initrd");
const initial_ramdisk = @import("initial-ramdisk");
const platform = @import("platform");
const tests = @import("tests.zig");
const build_options = @import("build_options");
@@ -161,10 +161,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// Snapshot the device tree for user-space drivers (device_enumerate/claim/
// mmio_map operate on this flat, id-indexed table + claim map).
device_service.init(&device_tree);
if (device_service.dropped > 0) {
devices_broker.init(&device_tree);
if (devices_broker.dropped > 0) {
// Otherwise entirely silent: drivers would just never see that hardware.
log.print("danos: WARNING {d} device(s) dropped — table full\n", .{device_service.dropped});
log.print("danos: WARNING {d} device(s) dropped — table full\n", .{devices_broker.dropped});
}
// Install the device-IRQ trampolines, so a driver's irq_bind has vectors to
@@ -285,10 +285,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
status("no /sbin/init on the boot volume.\n");
}
// Spawn the extra user binaries the loader ferried in the initrd (the VFS
// Spawn the extra user binaries the loader ferried in the initial_ramdisk (the VFS
// server, and later device drivers). For now the kernel launches them all;
// once init is a real service supervisor it will spawn them itself (system_spawn).
startInitrdBinaries(boot_information);
startInitialRamdiskBinaries(boot_information);
// Become the idle task: drop below every real task and halt until an
// interrupt. The timer keeps preempting into init and any other work.
@@ -297,22 +297,22 @@ fn kmain(boot_information: *const BootInformation) noreturn {
architecture.halt();
}
/// Spawn every program bundled in the initrd as its own ring-3 process. A bad
/// Spawn every program bundled in the initial_ramdisk as its own ring-3 process. A bad
/// image or a program that fails to load is logged and skipped — the rest of the
/// system still runs.
fn startInitrdBinaries(boot_information: *const danos.BootInformation) void {
if (boot_information.initrd_len == 0) return;
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initrd_base)))[0..boot_information.initrd_len];
const rd = initrd.Reader.init(image) orelse {
status("initrd: bad image, skipping\n");
fn startInitialRamdiskBinaries(boot_information: *const danos.BootInformation) void {
if (boot_information.initial_ramdisk_len == 0) return;
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
status("initial_ramdisk: bad image, skipping\n");
return;
};
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
statusPrint("starting /sbin/{s} (from initrd)...\n", .{item.name});
statusPrint("starting /sbin/{s} (from initial_ramdisk)...\n", .{item.name});
process.spawnProcess(item.blob, 4) catch |err| {
statusPrint("initrd: {s} failed to load: {s}\n", .{ item.name, @errorName(err) });
statusPrint("initial_ramdisk: {s} failed to load: {s}\n", .{ item.name, @errorName(err) });
};
}
}
+8 -8
View File
@@ -27,7 +27,7 @@ const pmm = @import("pmm.zig");
const scheduler = @import("scheduler.zig");
const sync = @import("sync.zig");
const ipc = @import("ipc-synchronous.zig");
const device_service = @import("device-service.zig");
const devices_broker = @import("devices-broker.zig");
const irq = @import("irq.zig");
const log = @import("log.zig");
@@ -206,12 +206,12 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
const sz = @sizeOf(danos.DeviceDescriptor);
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
const out: [*]danos.DeviceDescriptor = @ptrFromInt(buffer_ptr);
architecture.setSystemCallResult(state, device_service.enumerate(out[0..@intCast(cap)]));
architecture.setSystemCallResult(state, devices_broker.enumerate(out[0..@intCast(cap)]));
}
/// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process.
fn systemDeviceClaim(state: *architecture.CpuState) void {
if (device_service.claim(architecture.systemCallArg(state, 0), scheduler.current().id))
if (devices_broker.claim(architecture.systemCallArg(state, 0), scheduler.current().id))
architecture.setSystemCallResult(state, 0)
else
fail(state);
@@ -225,9 +225,9 @@ fn systemMmioMap(state: *architecture.CpuState) void {
const resource_index = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const owner = device_service.ownerOf(device_id) orelse return fail(state);
const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
if (owner != t.id) return fail(state); // not claimed by this process
const r = device_service.resourceOf(device_id, resource_index) orelse return fail(state);
const r = devices_broker.resourceOf(device_id, resource_index) orelse return fail(state);
if (r.kind != @intFromEnum(danos.ResourceKind.memory)) return fail(state);
if (t.device_map_next == 0) t.device_map_next = device_arena_base;
@@ -263,7 +263,7 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
var descriptor: danos.DeviceDescriptor = undefined;
if (!ipc.copyFromUser(t.aspace, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
const id = device_service.register(parent_id, t.id, &descriptor) catch return fail(state);
const id = devices_broker.register(parent_id, t.id, &descriptor) catch return fail(state);
architecture.setSystemCallResult(state, id);
}
@@ -281,9 +281,9 @@ fn releaseIrqs(t: *scheduler.Task) void {
/// by discovery. Neither a raw GSI nor an unclaimed device can get through — which
/// is why irq_bind takes a resource index and not an interrupt number.
fn ownedGsi(t: *scheduler.Task, device_id: u64, resource_index: u64) ?u32 {
const owner = device_service.ownerOf(device_id) orelse return null;
const owner = devices_broker.ownerOf(device_id) orelse return null;
if (owner != t.id) return null;
const r = device_service.resourceOf(device_id, resource_index) orelse return null;
const r = devices_broker.resourceOf(device_id, resource_index) orelse return null;
if (r.kind != @intFromEnum(danos.ResourceKind.irq)) return null;
if (r.start >= irq.maximum_gsi) return null;
return @intCast(r.start);
+53 -53
View File
@@ -12,7 +12,7 @@
const std = @import("std");
const danos = @import("danos");
const architecture = @import("architecture");
const device_service = @import("device-service.zig");
const devices_broker = @import("devices-broker.zig");
const platform = @import("platform");
const pmm = @import("pmm.zig");
const heap = @import("heap.zig");
@@ -22,7 +22,7 @@ const ipcsync = @import("ipc-synchronous.zig");
const irq = @import("irq.zig");
const sync = @import("sync.zig");
const process = @import("process.zig");
const initrd = @import("initrd");
const initial_ramdisk = @import("initial-ramdisk");
/// Formatted write straight to serial, independent of the framebuffer console.
fn log(comptime fmt: []const u8, args: anytype) void {
@@ -108,8 +108,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
initTest(boot_information);
} else if (eql(case, "process")) {
processTest(boot_information);
} else if (eql(case, "initrd")) {
initrdTest(boot_information);
} else if (eql(case, "initial-ramdisk")) {
initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) {
vfsTest(boot_information);
} else if (eql(case, "hpet")) {
@@ -952,24 +952,24 @@ fn initTest(boot_information: *const BootInformation) void {
result();
}
/// The initrd path: the bootloader handed over an image bundling extra user
/// The initial_ramdisk path: the bootloader handed over an image bundling extra user
/// binaries; parse it, spawn every program, and confirm one (the vfs stub)
/// reaches ring 3 and heartbeats — proving the whole ferry-parse-spawn pipeline.
fn initrdTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: initrd\n", .{});
check("bootloader handed over an initrd", boot_information.initrd_len != 0);
if (boot_information.initrd_len == 0) {
fn initialRamdiskTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: initial_ramdisk\n", .{});
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
if (boot_information.initial_ramdisk_len == 0) {
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initrd_base)))[0..boot_information.initrd_len];
const rd = initrd.Reader.init(image) orelse {
check("initrd image is valid", false);
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
check("initrd image is valid", true);
check("initrd contains at least one binary", rd.count >= 1);
check("initial_ramdisk image is valid", true);
check("initial_ramdisk contains at least one binary", rd.count >= 1);
process.write_count = 0;
process.write_from_user = false;
@@ -981,7 +981,7 @@ fn initrdTest(boot_information: *const BootInformation) void {
log("DANOS-INITRD-ERR: {s}: {s}\n", .{ item.name, @errorName(err) });
}
}
check("every initrd binary spawned", spawned == rd.count);
check("every initial_ramdisk binary spawned", spawned == rd.count);
// Wait for the spawned programs to run and make syscalls (they write + sleep).
scheduler.setPriority(1);
@@ -989,33 +989,33 @@ fn initrdTest(boot_information: *const BootInformation) void {
while (process.write_count < 2 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4);
check("initrd processes ran and made syscalls (>=2)", process.write_count >= 2);
check("initial_ramdisk processes ran and made syscalls (>=2)", process.write_count >= 2);
check("syscalls came from user mode (CPL 3)", process.write_from_user);
result();
}
/// The full VFS path: spawn the user-space VFS server and a client from the
/// initrd. The client opens a file through the runtime file API, writes, seeks, reads
/// initial_ramdisk. The client opens a file through the runtime file API, writes, seeks, reads
/// it back, and — only if the round trip matched — heartbeats "vfstest: ok". So
/// seeing that marker proves client open/write/read reached the server over IPC
/// and came back correct. (The client retries until the server registers.)
fn vfsTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: vfs\n", .{});
if (boot_information.initrd_len == 0) {
check("bootloader handed over an initrd", false);
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initrd_base)))[0..boot_information.initrd_len];
const rd = initrd.Reader.init(image) orelse {
check("initrd image is valid", false);
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.write_count = 0;
process.write_from_user = false;
// Spawn just the server and its client (other initrd binaries would write to
// Spawn just the server and its client (other initial_ramdisk binaries would write to
// the shared evidence buffer and confuse the marker check).
_ = spawnNamed(rd, "vfs");
_ = spawnNamed(rd, "vfs-test");
@@ -1037,9 +1037,9 @@ fn vfsTest(boot_information: *const BootInformation) void {
result();
}
/// Spawn the initrd binary named `name` as a ring-3 process. Returns false if it
/// Spawn the initial_ramdisk binary named `name` as a ring-3 process. Returns false if it
/// isn't in the image or fails to load.
fn spawnNamed(rd: initrd.Reader, name: []const u8) bool {
fn spawnNamed(rd: initial_ramdisk.Reader, name: []const u8) bool {
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
@@ -1051,36 +1051,36 @@ fn spawnNamed(rd: initrd.Reader, name: []const u8) bool {
}
/// IO passthrough + IRQ-as-IPC: a user-space driver drives real hardware and is
/// *woken by it*. Spawn hpetd, which claims the HPET, maps its registers into its
/// *woken by it*. Spawn hpet, which claims the HPET, maps its registers into its
/// own ring-3 address space, arms a level-triggered comparator, binds the interrupt
/// to an IPC endpoint, and then blocks. It prints "hpetd: ok" only after being woken
/// to an IPC endpoint, and then blocks. It prints "hpet: ok" only after being woken
/// `target_ticks` times — it cannot reach that line by polling, because the loop's
/// only exit is through `replyWait` returning a notification badge.
///
/// The interesting assertion is the last one, which doesn't trust hpetd at all: it
/// The interesting assertion is the last one, which doesn't trust hpet at all: it
/// reads the I/O APIC's redirection entry back and checks the kernel really routed
/// the line (our vector, level-triggered) and really left it unmasked after the
/// driver's final `irq_ack`. hpetd disables its comparator on the last interrupt, so
/// driver's final `irq_ack`. hpet disables its comparator on the last interrupt, so
/// that state is quiescent and not a race.
fn hpetTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: hpet\n", .{});
if (boot_information.initrd_len == 0) {
check("bootloader handed over an initrd", false);
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initrd_base)))[0..boot_information.initrd_len];
const rd = initrd.Reader.init(image) orelse {
check("initrd image is valid", false);
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.write_count = 0;
process.write_from_user = false;
check("hpetd spawned from the initrd", spawnNamed(rd, "hpetd"));
check("hpet spawned from the initial_ramdisk", spawnNamed(rd, "hpet"));
const prefix = "hpetd: ok";
const prefix = "hpet: ok";
scheduler.setPriority(1);
const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
@@ -1114,7 +1114,7 @@ fn hpetRouteOk() bool {
/// The GSI discovery recorded for the HPET, from the same device table the driver saw.
fn hpetGsi() ?u32 {
var buffer: [16]danos.DeviceDescriptor = undefined;
const n = @min(device_service.enumerate(&buffer), buffer.len);
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(danos.DeviceClass.timer)) continue;
if (d.parent != danos.no_parent) continue; // the block, not a comparator child
@@ -1130,34 +1130,34 @@ fn hpetGsi() ?u32 {
/// them from the hardware, and publishes each as a child via `device_register` — the
/// primitive a PCI bridge or USB hub driver is built from.
///
/// `busd` treats the HPET's register block as a bus and its comparators as children,
/// `bus` treats the HPET's register block as a bus and its comparators as children,
/// giving each a 0x20 sub-window. It checks its own work (children come back from the
/// table with the right parent and a strictly narrower window) and, importantly, that
/// the kernel **refuses** a child whose window escapes the parent's — without that,
/// `device_register` would be a system_call for mapping arbitrary physical memory. It prints
/// "busd: ok" only if all of that holds.
/// "bus: ok" only if all of that holds.
///
/// The kernel-side check here is the one busd can't make: that the children really did
/// The kernel-side check here is the one bus can't make: that the children really did
/// land in the device table with the containment invariant intact.
fn busTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: bus\n", .{});
if (boot_information.initrd_len == 0) {
check("bootloader handed over an initrd", false);
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initrd_base)))[0..boot_information.initrd_len];
const rd = initrd.Reader.init(image) orelse {
check("initrd image is valid", false);
const image = @as([*]const u8, @ptrFromInt(danos.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.write_count = 0;
process.write_from_user = false;
check("busd spawned from the initrd", spawnNamed(rd, "busd"));
check("bus spawned from the initial_ramdisk", spawnNamed(rd, "bus"));
const prefix = "busd: ok";
const prefix = "bus: ok";
scheduler.setPriority(1);
const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
@@ -1173,7 +1173,7 @@ fn busTest(boot_information: *const BootInformation) void {
result();
}
/// Every child `busd` registered must have each of its resources inside a parent
/// Every child `bus` registered must have each of its resources inside a parent
/// resource of the same kind — the invariant `device_register` exists to maintain,
/// checked from the kernel's own table rather than the driver's word for it.
///
@@ -1182,7 +1182,7 @@ fn busTest(boot_information: *const BootInformation) void {
/// bridge's `bus_range`, because a bus-number range isn't an address window.
fn childrenContained() bool {
var buffer: [64]danos.DeviceDescriptor = undefined;
const n = @min(device_service.enumerate(&buffer), buffer.len);
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
const bus_id = hpetDeviceId() orelse return false;
const p = buffer[@intCast(bus_id)];
@@ -1205,13 +1205,13 @@ fn childrenContained() bool {
if (!ok) return false;
}
}
return children > 0; // busd must have published at least one
return children > 0; // bus must have published at least one
}
/// Device id of the HPET (the bus busd claims), from the same table drivers see.
/// Device id of the HPET (the bus bus claims), from the same table drivers see.
fn hpetDeviceId() ?u64 {
var buffer: [64]danos.DeviceDescriptor = undefined;
const n = @min(device_service.enumerate(&buffer), buffer.len);
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(danos.DeviceClass.timer)) continue;
if (d.parent != danos.no_parent) continue; // a comparator child, not the block
@@ -1226,7 +1226,7 @@ fn hpetDeviceId() ?u64 {
/// (so a dead driver's device goes quiet instead of storming) and the slot cleared
/// (so an ISR never posts a notification into the endpoint that is about to be freed).
///
/// This is the path `hpetd` never takes — it runs forever — so it gets its own test.
/// This is the path `hpet` never takes — it runs forever — so it gets its own test.
/// Two properties, both read back from the hardware rather than from our own state:
///
/// 1. A bound GSI is routed and unmasked.
+17 -11
View File
@@ -1,18 +1,22 @@
//! The VFS wire protocol — the message format spoken between a client (via the
//! `runtime` file API) and the user-space VFS server over IPC. A request is a fixed
//! `Request` header followed by an inline payload (a path, or write bytes); a
//! reply is a fixed `Reply` header followed by an inline payload (read bytes, or
//! a Stat). Everything fits in one IPC message (<= ipc MESSAGE_MAXIMUM = 256 bytes).
//! The VFS wire protocol — the message format spoken between a client (via the file
//! API) and the user-space VFS server over IPC. A request is a fixed `Request` header
//! followed by an inline payload (a path, or write bytes); a reply is a fixed `Reply`
//! header followed by an inline payload (read bytes, or a FileStatus). Everything fits
//! in one IPC message (<= ipc MESSAGE_MAXIMUM = 256 bytes).
//!
//! This is a danos-native contract, so it uses danos names throughout — the POSIX
//! spellings (`stat`, `O_CREAT`, ...) live only in the POSIX layer
//! (library/posix/unistd.zig), which translates to these.
//!
//! This is user-space only — the kernel knows nothing of files or paths; it only
//! moves the bytes. Shared by library/runtime/unistd.zig (client) and system/services/vfs/vfs.zig (server).
//! moves the bytes. Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server).
pub const Operation = enum(u32) {
open, // open(path) -> node id
close, // close(node)
read, // read(node, offset, len) -> bytes
write, // write(node, offset, bytes) -> count
stat, // stat(node) -> Stat
status, // status(node) -> FileStatus
};
/// Request header. `node` is the server-side open-file id (from a prior open);
@@ -28,7 +32,7 @@ pub const Request = extern struct {
/// Reply header. `status` is 0 on success or a negative errno; `node` is the new
/// open-file id (for `open`); `len` is the payload length (bytes read, or the
/// Stat size).
/// FileStatus size).
pub const Reply = extern struct {
status: i32,
_padding: u32 = 0,
@@ -37,7 +41,9 @@ pub const Reply = extern struct {
_padding2: u32 = 0,
};
pub const Stat = extern struct {
/// A file's metadata (the danos-native answer to a `status` request). The POSIX
/// layer maps this onto `struct stat`.
pub const FileStatus = extern struct {
size: u64,
kind: u32,
_padding: u32 = 0,
@@ -49,5 +55,5 @@ pub const reply_size: usize = @sizeOf(Reply);
/// Largest inline payload that still fits one IPC message alongside a header.
pub const maximum_payload: usize = message_maximum - request_size;
/// Open flags.
pub const O_CREAT: u32 = 1;
/// Open flags (danos-native; the POSIX layer maps `O_CREAT` onto `create`).
pub const create: u32 = 1;
+2 -2
View File
@@ -1,13 +1,13 @@
//! /sbin/vfstest — a client that proves the VFS round trip end to end: open a
//! file through the `runtime` file API, write to it, seek back, read it, and compare.
//! On success it heartbeats "vfstest: ok" so the kernel test can observe it;
//! on failure it reports what went wrong. Shipped in the initrd alongside vfs.
//! on failure it reports what went wrong. Shipped in the initial_ramdisk alongside vfs.
const std = @import("std");
const runtime = @import("runtime");
pub fn main() void {
const u = runtime.unistd;
const u = @import("posix").unistd;
const payload = "hello-vfs";
// The VFS server may not have registered yet — retry open until it's up.
+4 -4
View File
@@ -1,4 +1,4 @@
//! /sbin/vfs — the user-space VFS server. Shipped in the initrd, spawned as a
//! /sbin/vfs — the user-space VFS server. Shipped in the initial_ramdisk, spawned as a
//! ring-3 process, and reached by every other process through IPC (the `runtime`
//! file API marshals open/read/write/stat/close into calls to this server's
//! endpoint, published under the well-known `vfs` service id).
@@ -101,10 +101,10 @@ fn handle(message: []const u8, out: []u8) usize {
if (off + n > nd.size) nd.size = off + n;
return writeReply(out, .{ .status = 0, .len = @intCast(n) }, &.{});
},
.stat => {
.status => {
const of = openAt(request.node) orelse return fail(out);
const st = protocol.Stat{ .size = nodes[of.node].size, .kind = 0 };
return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.Stat) }, std.mem.asBytes(&st));
const st = protocol.FileStatus{ .size = nodes[of.node].size, .kind = 0 };
return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&st));
},
.close => {
if (request.node < opens.len) opens[@intCast(request.node)].used = false;