Add runtime.time, drop demo drivers, harden TSC timekeeping

Time is a kernel concern in danos: the kernel owns the scheduling timer and
already exposes monotonic time via the clock/sleep/timer_bind syscalls, so a
userspace time service would be a redundant, slower path. This adds the generic
runtime.time module over those syscalls, retires the two demonstration drivers,
reorganizes the milestone docs, and makes the monotonic clock correct on Intel,
AMD, and inside any VM.

runtime.time (library/runtime/time.zig)
- Instant/Duration interface: now, sleep, spin, after, monotonicNanos, available
- a thin layer over system.clock/sleep/timerOnce; unit-tested arithmetic

Remove the demo drivers hpet and bus (a teaching example belongs in the docs,
not shipped in the tree)
- system/drivers/ now holds only real drivers: pci-bus, ps2-bus, usb-xhci-bus
- device-manager end-to-end test repointed to pci-bus (asserts on kernel state:
  the process table and the device tree, not a racy serial marker)
- device_register containment moved to a new in-kernel `containment` test
- the driver-model worked example moved inline into docs/drivers.md

Reorganize milestone docs into topic docs
- m17-m18 / m19-m20 / m21 plans dissolved into process-lifecycle, device-manager,
  discovery, and acpi docs; new docs/power.md and docs/timers.md; ~20 citations
  repointed; plan docs deleted

TSC reliability (apic.zig, smp.zig, cpu.zig, kernel.zig)
- check the invariant-TSC bit (CPUID 0x80000007 EDX[8]) on Intel and AMD
- cross-core "warp" check at SMP bring-up, pairwise BSP<->AP as each core comes up
- fall back to the HPET clocksource when the TSC is not invariant (a bare VM) or
  not synchronized (a warp), switched continuously so time never jumps
- boot log reports the outcome; new tsc-sync test exercises the TSC + warp path

Verified: zig build; zig build test; 60/60 QEMU cases (incl. new containment and
tsc-sync).
This commit is contained in:
Daniel Samson
2026-07-13 11:56:43 +01:00
parent 5b63a841ba
commit 452080e997
36 changed files with 1150 additions and 1267 deletions
+230 -13
View File
@@ -80,6 +80,35 @@ var timer_hz: u32 = 0;
var tsc_hz: u64 = 0;
var tsc_base: u64 = 0;
/// Whether the TSC is architecturally **invariant** — a constant rate regardless of
/// P/C-state transitions, and thus valid as a clocksource (CPUID leaf 0x80000007,
/// EDX bit 8). AMD and modern Intel set it; the bare qemu64 model does not. Measured
/// frequency alone is not enough: a non-invariant TSC speeds up and slows down with
/// the core clock, so reading it as wall time would drift.
var tsc_invariant: bool = false;
/// Cleared if the cross-core warp check (checkWarpSource) ever sees the TSC read
/// lower on one core than the max another core has already published — i.e. the
/// per-core TSCs are not synchronized, and a task migrating cores could see time go
/// backward. Starts true (assume synchronized until proven otherwise).
var tsc_synced: bool = true;
/// The worst backward skew the warp check observed, in TSC cycles (0 = none).
var tsc_warp_cycles: u64 = 0;
/// The monotonic clock's source. The TSC when it is invariant *and* synchronized —
/// the fast `rdtsc` path taken on real Intel/AMD and modern VMs. Otherwise the HPET
/// main counter: a single fixed-rate counter, immune to both per-core skew and
/// frequency scaling, so it stays accurate on a bare VM or a warped machine.
const ClockSource = enum { tsc, hpet };
var clock_source: ClockSource = .tsc;
/// HPET standby clocksource, set up in calibrate() whenever an HPET exists (whether
/// or not calibration itself measured against it): its frequency, the counter value
/// chosen as the zero point, and its width mask. Only a 64-bit HPET is used as a
/// clocksource — a 32-bit one wraps too fast to be monotonic without accumulation.
var hpet_clock_hz: u64 = 0;
var hpet_clock_base: u64 = 0;
var hpet_clock_mask: u64 = ~@as(u64, 0);
/// Read the 64-bit Time Stamp Counter.
fn rdtsc() u64 {
var low: u32 = undefined;
@@ -220,6 +249,31 @@ pub fn calibrate() void {
}
tsc_base = rdtsc(); // the clock's zero point (boot)
// Decide whether the TSC is trustworthy as a clocksource. Frequency (measured
// above, possibly against the HPET/PIT) is necessary but not sufficient: the TSC
// must also be *invariant* (CPUID 0x80000007 EDX[8]). AMD and modern Intel set
// this; the bare qemu64 model does not.
tsc_invariant = tscIsInvariant();
// Bring up the HPET as a standby clocksource whenever one exists — even on the
// CPUID-0x15 path where calibration never touched it — so a non-invariant TSC
// (here) or an unsynchronized one (checkWarpSource, during SMP bring-up) can fall
// back to a source that is immune to both. hpetHz() maps + enables the counter
// and is idempotent if calibration already used it.
if (configuration_hpet_base != 0) {
if (hpetHz()) |hz| {
hpet_clock_mask = hpetMask();
if (hpet_clock_mask == ~@as(u64, 0)) { // only a 64-bit HPET is monotonic enough
hpet_clock_hz = hz;
hpet_clock_base = readHpet();
}
}
}
// Select the source: the fast TSC when invariant, else the HPET if we have one.
// (checkWarpSource may still demote TSC -> HPET later if the cores' TSCs skew.)
if (!tsc_invariant and hpet_clock_hz != 0) clock_source = .hpet;
}
/// Run the LAPIC timer one-shot from its maximum count while a monotonic reference
@@ -275,7 +329,7 @@ fn calibratePit() void {
// --- reference clocks ------------------------------------------------------
/// TSC frequency from CPUID leaf 0x15 (crystal_hz * numerator / denominator), or
/// null if the CPU doesn't enumerate it (common under QEMU).
/// null if the CPU doesn't enumerate it (common under QEMU, and on AMD).
fn cpuidTscHz() ?u64 {
if (cpuid(0).eax < 0x15) return null;
const r = cpuid(0x15);
@@ -283,6 +337,15 @@ fn cpuidTscHz() ?u64 {
return @as(u64, r.ecx) * r.ebx / r.eax;
}
/// Whether the CPU advertises an **invariant** TSC (CPUID leaf 0x80000007, EDX
/// bit 8) — the architectural guarantee, on both Intel and AMD, that the TSC ticks
/// at a constant rate across P/C-states and never stops. Requires the extended-leaf
/// range to reach 0x80000007 first.
fn tscIsInvariant() bool {
if (cpuid(0x80000000).eax < 0x80000007) return false;
return (cpuid(0x80000007).edx & (1 << 8)) != 0;
}
const CpuidRegs = struct { eax: u32, ebx: u32, ecx: u32, edx: u32 };
fn cpuid(leaf: u32) CpuidRegs {
@@ -310,11 +373,20 @@ fn hpetWrite64(off: usize, value: u64) void {
@as(*volatile u64, @ptrFromInt(configuration_hpet_base + off)).* = value;
}
/// Whether the HPET has been mapped into the physmap yet, so `configuration_hpet_base`
/// already holds the virtual address. `hpetHz` is called more than once (calibration
/// may use the HPET, and the standby-clocksource setup asks for it again), and mapping
/// an already-mapped base a second time would double-offset it into an overflow.
var hpet_mapped: bool = false;
/// Map + enable the HPET and return its tick frequency, or null if unusable.
/// Maps the HPET into the physmap and switches configuration_hpet_base to that virtual
/// address, so the register accessors reach it without the identity map.
/// address, so the register accessors reach it without the identity map. Idempotent.
fn hpetHz() ?u64 {
configuration_hpet_base = paging.mapMmio(configuration_hpet_base, 0x400, true);
if (!hpet_mapped) {
configuration_hpet_base = paging.mapMmio(configuration_hpet_base, 0x400, true);
hpet_mapped = true;
}
const caps = hpetRead64(0x00);
const period_fs = caps >> 32; // femtoseconds per tick
if (period_fs == 0) return null;
@@ -364,24 +436,169 @@ pub fn tscHz() u64 {
return tsc_hz;
}
// Monotonic high-resolution clock, from the TSC. A function per resolution, each
// scaling the cycle delta directly at its unit (the 128-bit intermediate avoids
// overflow across a long uptime). nanos() resolves to a few ns; millis() is what
// the scheduler uses for sleep deadlines.
// Monotonic high-resolution clock. A function per resolution, each scaling the
// counter delta directly at its unit (the 128-bit intermediate avoids overflow
// across a long uptime). nanos() resolves to a few ns on the TSC; millis() is what
// the scheduler uses for sleep deadlines. The source is the TSC when it is invariant
// and synchronized, else the HPET counter (see clock_source) — the branch is one
// global load and the TSC path is unchanged from before.
/// The selected source's counter delta since its zero point.
fn clockCount() u64 {
return switch (clock_source) {
.tsc => rdtsc() -% tsc_base,
// A 64-bit HPET (the only kind we select) never wraps in any realistic
// uptime, so the wrapping subtraction is exact.
.hpet => readHpet() -% hpet_clock_base,
};
}
/// The selected source's frequency (0 if the clock is unavailable/uncalibrated).
fn clockHertz() u64 {
return switch (clock_source) {
.tsc => tsc_hz,
.hpet => hpet_clock_hz,
};
}
pub fn nanos() u64 {
if (tsc_hz == 0) return 0;
return @intCast(@as(u128, rdtsc() -% tsc_base) * 1_000_000_000 / tsc_hz);
const hz = clockHertz();
if (hz == 0) return 0;
return @intCast(@as(u128, clockCount()) * 1_000_000_000 / hz);
}
pub fn micros() u64 {
if (tsc_hz == 0) return 0;
return @intCast(@as(u128, rdtsc() -% tsc_base) * 1_000_000 / tsc_hz);
const hz = clockHertz();
if (hz == 0) return 0;
return @intCast(@as(u128, clockCount()) * 1_000_000 / hz);
}
pub fn millis() u64 {
if (tsc_hz == 0) return 0;
return @intCast(@as(u128, rdtsc() -% tsc_base) * 1_000 / tsc_hz);
const hz = clockHertz();
if (hz == 0) return 0;
return @intCast(@as(u128, clockCount()) * 1_000 / hz);
}
/// Whether the CPU advertises an invariant TSC (CPUID 0x80000007 EDX[8]).
pub fn tscInvariant() bool {
return tsc_invariant;
}
/// Test hook: force the TSC clocksource on, as if the CPU had advertised an invariant
/// TSC. QEMU's TCG accelerator (the only one for an x86 guest on an Apple-Silicon
/// host) does not expose the invariant-TSC bit — its emulated TSC isn't invariant — so
/// the tsc-sync test can't reach the real-Intel/AMD/KVM path through CPUID. This lets
/// that test exercise the TSC clocksource and the cross-core warp check anyway. tsc_base
/// is left as-is so the switch from the HPET is continuous.
pub fn forceTscClocksourceForTest() void {
tsc_invariant = true;
clock_source = .tsc;
}
/// How many per-AP warp checks actually ran (a rendezvous completed) — lets a test
/// confirm the cross-core check executed rather than being skipped.
pub fn warpChecksRun() u32 {
return warp_checks;
}
/// Whether the per-core TSCs are synchronized (no backward warp seen at bring-up).
pub fn tscSynced() bool {
return tsc_synced;
}
/// The active monotonic clocksource, for the boot log and tests.
pub fn clockSourceName() []const u8 {
return switch (clock_source) {
.tsc => "tsc",
.hpet => "hpet",
};
}
// --- cross-core TSC synchronization ("warp") check -------------------------
// Two cores hammer a shared "max seen" TSC value under a lock; if either reads a
// value below that max, its TSC lags the other's, and time would run backward for a
// task migrating between them (Linux calls this a warp). danos brings APs up one at a
// time, so this runs pairwise: the BSP (source) against each AP (target) as it comes
// online. It only matters — and only runs — while the TSC is the clocksource; on a
// machine already on the HPET (a bare VM) the whole rendezvous is skipped.
var warp_lock: u32 = 0;
var warp_last: u64 = 0;
var warp_bsp_ready: u32 = 0;
var warp_ap_ready: u32 = 0;
var warp_stop: u32 = 0;
var warp_checks: u32 = 0; // completed per-AP rendezvous count (for the tsc-sync test)
const warp_rounds: u32 = 1 << 20; // locked reads on the BSP: ~1 ms at GHz rates
const warp_spin_limit: u64 = 1 << 32; // bound every rendezvous wait so a lost core can't hang boot
fn warpTick() void {
while (@cmpxchgWeak(u32, &warp_lock, 0, 1, .acquire, .monotonic) != null) asm volatile ("pause");
const t = rdtsc();
if (t < warp_last) {
const delta = warp_last - t;
if (delta > tsc_warp_cycles) tsc_warp_cycles = delta;
tsc_synced = false;
} else {
warp_last = t;
}
@atomicStore(u32, &warp_lock, 0, .release);
}
/// Spin (bounded) until `flag` is nonzero; false on timeout.
fn warpAwait(flag: *u32) bool {
var spins: u64 = 0;
while (@atomicLoad(u32, flag, .acquire) == 0) : (spins += 1) {
if (spins >= warp_spin_limit) return false;
asm volatile ("pause");
}
return true;
}
/// BSP side of the pairwise TSC warp check, run once per AP as it reports in. No-op
/// unless the TSC is the active clocksource. If the AP's TSC proves to lag, demote
/// the monotonic clock to the HPET without a discontinuity.
pub fn checkWarpSource() void {
if (clock_source != .tsc) return;
warp_last = 0;
@atomicStore(u32, &warp_stop, 0, .release);
@atomicStore(u32, &warp_ap_ready, 0, .release);
@atomicStore(u32, &warp_bsp_ready, 1, .release);
if (!warpAwait(&warp_ap_ready)) { // AP never joined the rendezvous; skip, don't hang
@atomicStore(u32, &warp_bsp_ready, 0, .release);
return;
}
var i: u32 = 0;
while (i < warp_rounds) : (i += 1) warpTick();
@atomicStore(u32, &warp_stop, 1, .release);
@atomicStore(u32, &warp_bsp_ready, 0, .release);
warp_checks += 1;
if (!tsc_synced and hpet_clock_hz != 0) demoteToHpet();
}
/// AP side: join the BSP's warp check, then return so the core can enter the
/// scheduler. Bounded so a missing BSP can't strand the core.
pub fn checkWarpTarget() void {
if (clock_source != .tsc) return;
if (!warpAwait(&warp_bsp_ready)) return;
@atomicStore(u32, &warp_ap_ready, 1, .release);
var spins: u64 = 0;
while (@atomicLoad(u32, &warp_stop, .acquire) == 0) : (spins += 1) {
if (spins >= warp_spin_limit) return;
warpTick();
}
}
/// Switch the clocksource from the TSC to the HPET without a discontinuity: choose
/// the HPET zero point so it reads the same nanosecond value the TSC does right now,
/// so time neither jumps nor runs backward across the switch. Called when the warp
/// check proves the per-core TSCs unsynchronized.
fn demoteToHpet() void {
const now_ns = @as(u128, rdtsc() -% tsc_base) * 1_000_000_000 / tsc_hz;
const equivalent_ticks: u64 = @intCast(now_ns * hpet_clock_hz / 1_000_000_000);
hpet_clock_base = readHpet() -% equivalent_ticks;
clock_source = .hpet;
}
/// Acknowledge the current interrupt so the LAPIC will deliver the next one.
+29
View File
@@ -469,6 +469,35 @@ pub fn clockHz() u64 {
return apic.tscHz();
}
/// Whether the CPU guarantees an **invariant** TSC (CPUID 0x80000007 EDX[8] on
/// x86; the analogous architectural guarantee elsewhere). When false the TSC is not
/// used as the clocksource.
pub fn clockInvariant() bool {
return apic.tscInvariant();
}
/// Whether the per-core clock counters are synchronized (no backward warp observed
/// at SMP bring-up). When false the clock falls back off the TSC.
pub fn clockSynchronized() bool {
return apic.tscSynced();
}
/// The active monotonic clocksource, for the boot log ("tsc" or "hpet" on x86).
pub fn clockSourceName() []const u8 {
return apic.clockSourceName();
}
/// Test hook: force the TSC clocksource on, to exercise the TSC + warp-check path on
/// a hypervisor that won't advertise an invariant TSC (see apic.forceTscClocksourceForTest).
pub fn forceTscClocksourceForTest() void {
apic.forceTscClocksourceForTest();
}
/// How many per-AP TSC warp checks completed (for the tsc-sync test).
pub fn warpChecksRun() u32 {
return apic.warpChecksRun();
}
/// Unmask maskable interrupts (`sti`) so device interrupts get delivered.
pub fn enableInterrupts() void {
asm volatile ("sti");
+13 -1
View File
@@ -148,7 +148,14 @@ pub fn startAp(apic_id: u32, stack_top: usize, percpu: usize, index: usize, cr3:
// Wait up to 100 ms for the AP to reach apEntry and set the flag.
const deadline = apic.millis() + 100;
while (apic.millis() < deadline) {
if (@atomicLoad(u32, &ap_alive, .acquire) != 0) return true;
if (@atomicLoad(u32, &ap_alive, .acquire) != 0) {
// Cross-check this core's TSC against the BSP's before it joins the run
// loop: an unsynchronized TSC must be caught before any task can migrate
// onto this core and observe time going backward. No-op unless the TSC is
// the clocksource (apic.checkWarpSource).
apic.checkWarpSource();
return true;
}
asm volatile ("pause");
}
return false;
@@ -177,6 +184,11 @@ fn apEntry(percpu: usize) callconv(.c) noreturn {
@atomicStore(u32, &ap_alive, 1, .release); // "architecture state up" — BSP is polling this
// Rendezvous with the BSP for the TSC warp check (no-op unless the TSC is the
// clocksource) before joining the run loop, so this core's clock is vetted before
// it can run any task.
apic.checkWarpTarget();
if (secondary_entry) |enterScheduler| enterScheduler(); // joins the run loop
while (true) asm volatile ("hlt"); // (only if no entry was registered)
}
+1 -1
View File
@@ -189,7 +189,7 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
if (!ok) return error.NotContained;
}
// Idempotent on exact match (docs/m19-m20-plan.md decision 3): a restarted
// Idempotent on exact match (docs/device-manager.md): a restarted
// registering bus re-registers what it rediscovers, and the table has no
// unregister — an identical (class, identity, resources) child under the
// same parent returns the existing id instead of appending a duplicate.
+21 -1
View File
@@ -257,10 +257,30 @@ fn kmain(boot_information: *const BootInformation) noreturn {
log.checkpoint(cp_timer);
log.print("/system/kernel: timer online ({d} Hz tick; timer clock {d} MHz, clock {d} MHz; calibrated via {s})\n", .{ architecture.timer_hz, architecture.timerClockHz() / 1_000_000, architecture.clockHz() / 1_000_000, architecture.timerCalibrationSource() });
// The tsc-sync test forces the TSC clocksource on before the cores come up, so the
// TSC + warp-check path is exercised even under TCG (which won't advertise an
// invariant TSC). Inert in a normal build (docs/timers.md).
if (build_options.test_case) |tc| {
if (std.mem.eql(u8, tc, "tsc-sync")) architecture.forceTscClocksourceForTest();
}
// Wake the other cores (application processors). A no-op on a single-core
// machine; on SMP each AP climbs to long mode and reports in (docs/smp.md).
// machine; on SMP each AP climbs to long mode and reports in (docs/smp.md). The
// per-core TSC warp check rides this: each AP is vetted before it joins the run
// loop (docs/timers.md).
bringUpSecondaries();
// Report the monotonic clock's final reliability, now the warp check has run on
// every core. On real Intel/AMD this is the invariant, synchronized TSC; a bare
// VM (no invariant bit) or a machine whose cores' TSCs skew uses the HPET instead.
log.print("/system/kernel: clocksource {s} (TSC invariant: {s}, synchronized: {s})\n", .{
architecture.clockSourceName(),
if (architecture.clockInvariant()) "yes" else "no",
if (architecture.clockSynchronized()) "yes" else "no",
});
if (!architecture.clockSynchronized())
log.write("/system/kernel: WARNING: per-core TSCs are not synchronized; monotonic clock moved off the TSC\n");
// In a test build (`zig build -Dtest-case=<name>`), run that case and stop.
// Normal builds fall through to the idle halt.
if (build_options.test_case) |case| {
+143 -179
View File
@@ -102,6 +102,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
stressTest();
} else if (eql(case, "smp-retry")) {
smpRetryTest();
} else if (eql(case, "tsc-sync")) {
tscSyncTest();
} else if (eql(case, "fault-ud")) {
faultInvalidOpcode();
} else if (eql(case, "fault-pf")) {
@@ -162,14 +164,12 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
vfsTest(boot_information);
} else if (eql(case, "input")) {
inputTest(boot_information);
} else if (eql(case, "hpet")) {
hpetTest(boot_information);
} else if (eql(case, "iopass")) {
ioPassTest();
} else if (eql(case, "irqfree")) {
irqFreeTest();
} else if (eql(case, "bus")) {
busTest(boot_information);
} else if (eql(case, "containment")) {
containmentTest();
} else if (eql(case, "device-manager")) {
deviceManagerTest(boot_information);
} else if (eql(case, "poweroff")) {
@@ -213,7 +213,7 @@ fn eql(a: []const u8, b: []const u8) bool {
/// Whether the captured last-write buffer *contains* `needle`. Markers are
/// matched as substrings, not prefixes, so a service's source-path debug prefix
/// (`system/drivers/hpet: ok`) still satisfies a marker like `hpet: ok`.
/// (`system/drivers/pci-bus: ...`) still satisfies a marker like `pci-bus: `.
fn bufferHas(needle: []const u8) bool {
return std.mem.indexOf(u8, process.write_buffer[0..process.write_len], needle) != null;
}
@@ -1211,6 +1211,35 @@ fn clockTest() void {
result();
}
/// The TSC clocksource + cross-core warp check (`-smp 4`). This is the real
/// Intel/AMD / KVM path — an invariant, synchronized TSC. TCG (the only x86
/// accelerator on an Apple-Silicon host) won't advertise an invariant TSC, so the
/// boot forces the TSC clocksource on (kernel.zig, gated on this case) to exercise
/// the machinery: the kernel must run the per-AP warp check as each core came up,
/// find the cores' TSCs synchronized, and keep the clock on the TSC (no HPET
/// fallback). The default suite (no force) exercises the HPET fallback instead.
fn tscSyncTest() void {
log("DANOS-TEST-BEGIN: tsc-sync\n", .{});
check("clocksource is the TSC (forced-invariant path)", eql(architecture.clockSourceName(), "tsc"));
check("the cross-core warp check ran on the APs", architecture.warpChecksRun() >= 1);
check("per-core TSCs synchronized (no warp, no HPET fallback)", architecture.clockSynchronized());
// A warp that slipped past the bring-up check would surface as a backward reading.
var last = architecture.nanos();
var monotonic = true;
var advanced = false;
var i: u32 = 0;
while (i < 1_000_000) : (i += 1) {
const t = architecture.nanos();
if (t < last) monotonic = false;
if (t > last) advanced = true;
last = t;
}
check("monotonic clock advanced", advanced);
check("monotonic clock never ran backward", monotonic);
result();
}
var proc_worker_run: bool = true;
var proc_worker_ran: bool = false;
@@ -2258,68 +2287,7 @@ fn spawnNamed(rd: initial_ramdisk.Reader, name: []const u8) bool {
return false;
}
/// IO passthrough + IRQ-as-IPC: a user-space driver drives real hardware and is
/// *woken by it*. Spawn hpet, which claims the HPET, maps its registers into its
/// own ring-3 address space, arms a level-triggered comparator, binds the interrupt
/// to an IPC endpoint, and then blocks. It prints "hpet: ok" only after being woken
/// `target_ticks` times — it cannot reach that line by polling, because the loop's
/// only exit is through `replyWait` returning a notification badge.
///
/// The interesting assertion is the last one, which doesn't trust hpet at all: it
/// reads the I/O APIC's redirection entry back and checks the kernel really routed
/// the line (our vector, level-triggered) and really left it unmasked after the
/// driver's final `irq_ack`. hpet disables its comparator on the last interrupt, so
/// that state is quiescent and not a race.
fn hpetTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: hpet\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.write_count = 0;
process.write_from_user = false;
check("hpet spawned from the initial_ramdisk", spawnNamed(rd, "hpet"));
const prefix = "hpet: ok";
scheduler.setPriority(1);
const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (bufferHas(prefix) and process.write_count >= 2) break;
scheduler.yield();
}
scheduler.setPriority(4);
const ok = bufferHas(prefix);
check("user driver mapped HPET MMIO and was woken by its interrupt", ok);
check("driver syscalls came from user mode (CPL 3)", process.write_from_user);
check("kernel routed and re-armed the HPET's line at the I/O APIC", hpetRouteOk());
result();
}
/// Read back the I/O APIC redirection entry for the HPET's GSI and confirm the
/// kernel programmed it: a vector in the device window, level-triggered, unmasked.
/// Independent of anything the driver reported about itself.
fn hpetRouteOk() bool {
const gsi = hpetGsi() orelse return false;
if (gsi >= architecture.irqRouteCount()) return false;
const low = architecture.irqRouteRaw(gsi); // entry index == GSI (this I/O APIC's gsi_base is 0)
const vector: u8 = @truncate(low & 0xFF);
const masked = low & (1 << 16) != 0;
const level = low & (1 << 15) != 0;
return vector >= architecture.irq_vector_base and
vector < architecture.irq_vector_base + architecture.irq_vector_count and
level and !masked;
}
/// The GSI discovery recorded for the HPET, from the same device table the driver saw.
/// The GSI discovery recorded for the HPET, from the same device table drivers see.
fn hpetGsi() ?u32 {
var buffer: [16]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
@@ -2334,87 +2302,87 @@ fn hpetGsi() ?u32 {
return null;
}
/// Bus driver: a user process claims a device that contains other devices, enumerates
/// them from the hardware, and publishes each as a child via `device_register` — the
/// primitive a PCI bridge or USB hub driver is built from.
///
/// `bus` treats the HPET's register block as a bus and its comparators as children,
/// giving each a 0x20 sub-window. It checks its own work (children come back from the
/// table with the right parent and a strictly narrower window) and, importantly, that
/// the kernel **refuses** a child whose window escapes the parent's — without that,
/// `device_register` would be a system_call for mapping arbitrary physical memory. It prints
/// "bus: ok" only if all of that holds.
///
/// The kernel-side check here is the one bus can't make: that the children really did
/// land in the device table with the containment invariant intact.
fn busTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: bus\n", .{});
/// `device_register` containment — a kernel security property, tested directly against
/// the broker (no user-space demo driver). A bus driver publishes children of a device
/// it owns; the kernel must **refuse** any child whose resource escapes the parent's
/// grant, or `device_register` would become a system call for mapping arbitrary physical
/// memory. This is the property the old `bus` demo driver proved end-to-end; with the
/// demo gone, the property is asserted where it lives — in the kernel. Also checks the
/// idempotence rule (M19.0): re-registering an identical child returns the same id
/// instead of appending a duplicate.
fn containmentTest() void {
log("DANOS-TEST-BEGIN: containment\n", .{});
// M19.0: device_register is idempotent on exact match — a restarted
// registering bus must not duplicate its children. Driven directly against
// the broker: claim an unclaimed node, register the same (class, hid,
// resourceless) child twice, expect one id and one table entry.
{
const me = scheduler.currentId();
var probe: [1]device_abi.DeviceDescriptor = undefined;
const total = devices_broker.enumerate(&probe);
check("device tree is seeded for the idempotence check", total >= 1);
if (devices_broker.ownerOf(0) == null) {
check("claimed device 0 for the idempotence check", devices_broker.claim(0, me));
var child = std.mem.zeroes(device_abi.DeviceDescriptor);
child.class = @intFromEnum(device_abi.DeviceClass.unknown);
child.pci_class = device_abi.no_pci_class;
child.hid_len = 4;
child.hid[0..4].* = "idem".*;
const first = devices_broker.register(0, me, &child) catch 0;
check("first register succeeded", first != 0);
const before = devices_broker.enumerate(&probe);
const second = devices_broker.register(0, me, &child) catch 0;
check("re-register returned the same id", second == first);
check("re-register grew nothing", devices_broker.enumerate(&probe) == before);
devices_broker.releaseAllOwnedBy(me);
} else {
check("device 0 unexpectedly claimed before the idempotence check", false);
}
}
const me = scheduler.currentId();
var buffer: [64]device_abi.DeviceDescriptor = undefined;
check("device tree is seeded", devices_broker.enumerate(&buffer) >= 1);
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
// The kernel-seeded HPET timer block is a device with a memory resource — a natural
// parent to publish sub-window children under, as a PCI bridge or USB hub would.
const parent_id = hpetDeviceId() orelse {
check("found a device with a memory window to parent children under", false);
result();
return;
};
const parent = buffer[@intCast(parent_id)];
var window: ?device_abi.ResourceDescriptor = null;
for (0..parent.resource_count) |j| {
if (parent.resources[j].kind == @intFromEnum(device_abi.ResourceKind.memory)) window = parent.resources[j];
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
const parent_window = window orelse {
check("parent exposes a memory window", false);
result();
return;
};
process.write_count = 0;
process.write_from_user = false;
check("bus spawned from the initial_ramdisk", spawnNamed(rd, "bus"));
const prefix = "bus: ok";
scheduler.setPriority(1);
const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (bufferHas(prefix)) break;
scheduler.yield();
if (devices_broker.ownerOf(parent_id) != null) {
check("parent device was unclaimed at the start of the test", false);
result();
return;
}
scheduler.setPriority(4);
check("claimed the parent device", devices_broker.claim(parent_id, me));
defer devices_broker.releaseAllOwnedBy(me);
// A child whose window lies inside the parent's is accepted.
var fits = childDescriptor("cfit", parent_window.start, 0x20);
const before = devices_broker.enumerate(&buffer);
const good = devices_broker.register(parent_id, me, &fits) catch 0;
check("a contained child is registered", good != 0);
check("the contained child was appended to the table", devices_broker.enumerate(&buffer) == before + 1);
// A child whose window escapes the parent's is refused with NotContained.
var escapes = childDescriptor("cesc", parent_window.start, parent_window.len + 0x1000);
const refused = if (devices_broker.register(parent_id, me, &escapes)) |_| false else |err| err == error.NotContained;
check("an out-of-window child is refused (NotContained)", refused);
check("the refused child left the table unchanged", devices_broker.enumerate(&buffer) == before + 1);
// Idempotent on exact match: re-registering the accepted child returns its id and
// appends nothing (M19.0 — a restarted bus re-reports what it rediscovers).
const again = devices_broker.register(parent_id, me, &fits) catch 0;
check("re-registering an identical child returns the same id", again != 0 and again == good);
check("re-registering grew nothing", devices_broker.enumerate(&buffer) == before + 1);
const ok = bufferHas(prefix);
check("bus driver published children and the kernel refused an out-of-window one", ok);
check("driver syscalls came from user mode (CPL 3)", process.write_from_user);
check("every registered child is contained in its parent", childrenContained());
result();
}
/// A minimal child descriptor with one memory resource, for the containment test.
fn childDescriptor(hid: []const u8, start: u64, len: u64) device_abi.DeviceDescriptor {
var child = std.mem.zeroes(device_abi.DeviceDescriptor);
child.class = @intFromEnum(device_abi.DeviceClass.unknown);
child.pci_class = device_abi.no_pci_class;
child.hid_len = @intCast(hid.len);
@memcpy(child.hid[0..hid.len], hid);
child.resource_count = 1;
child.resources[0] = .{ .kind = @intFromEnum(device_abi.ResourceKind.memory), .start = start, .len = len };
return child;
}
/// The device manager (a ring-3 service) enumerates /system/devices, matches each
/// device to a driver, and — eventually — spawns it. This increment only checks the
/// discovery+matching half: it must find the HPET (a timer) and decide `hpet` serves
/// it, printing "device-manager: ok". It uses no special privilege — the same
/// `device_enumerate` any process could call. (Spawning is the next increment.)
/// device to a driver, and spawns it. Proof of the whole discover -> match -> spawn ->
/// driver-up chain: boot only the device-manager; it must discover the PCI host bridge,
/// match `pci-bus`, and spawn it (with the bridge id as its argument) — and the spawned
/// pci-bus must reach its own live marker. It uses no special privilege — the same
/// `device_enumerate` any process could call.
fn deviceManagerTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: device-manager\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
@@ -2430,70 +2398,65 @@ fn deviceManagerTest(boot_information: *const BootInformation) void {
};
// Let `system_spawn` find bundled binaries by name (the normal boot path does
// this too). Only the device-manager is spawned here — so if `hpet` runs at all,
// it's because the manager discovered the timer, matched, and spawned it.
// this too). Only the device-manager is spawned here — so if `pci-bus` runs at
// all, it's because the manager discovered the PCI host bridge, matched, and
// spawned it.
process.setInitialRamdisk(image);
process.write_count = 0;
process.write_from_user = false;
check("device-manager spawned from the initial_ramdisk", spawnNamed(rd, "device-manager"));
// End-to-end proof: the driver the manager spawned reaches its own live marker.
// `hpet: ok` is hpet's final, stable message (it claims the timer, maps its MMIO,
// binds its IRQ, services one, then sleeps) — nothing overwrites the buffer after,
// so it's race-free to poll for. Its arrival means the whole
// discover -> match -> system_spawn -> driver-up chain worked.
const prefix = "hpet: ok";
// End-to-end proof, read from kernel state — not the racy last-write serial buffer,
// since many services keep logging after pci-bus. The manager must discover the PCI
// host bridge, match pci-bus, and spawn it, and pci-bus must come up: claim the
// bridge, map its ECAM, and register the functions it enumerates as children in the
// device tree.
scheduler.setPriority(1);
const deadline = architecture.millis() + 10000;
var spawned = false;
while (architecture.millis() < deadline) {
if (bufferHas(prefix)) break;
if (processRunning("pci-bus")) spawned = true;
if (spawned and pciFunctionsRegistered()) break;
scheduler.yield();
}
scheduler.setPriority(4);
const ok = bufferHas(prefix);
check("device manager matched the timer and system_spawn'd hpet, which came up", ok);
check("device manager discovered the PCI host bridge and spawned pci-bus", spawned);
check("pci-bus came up and registered the functions it enumerated", pciFunctionsRegistered());
check("its syscalls came from user mode (CPL 3)", process.write_from_user);
result();
}
/// Every child `bus` registered must have each of its resources inside a parent
/// resource of the same kind — the invariant `device_register` exists to maintain,
/// checked from the kernel's own table rather than the driver's word for it.
///
/// Only *registered* children are checked, not the whole tree. Firmware topology is
/// trusted and doesn't obey containment: a PCI function's BAR is not inside its host
/// bridge's `bus_range`, because a bus-number range isn't an address window.
fn childrenContained() bool {
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
const bus_id = hpetDeviceId() orelse return false;
const p = buffer[@intCast(bus_id)];
var children: usize = 0;
for (buffer[0..n]) |d| {
if (d.parent != bus_id) continue;
children += 1;
for (0..d.resource_count) |i| {
const r = d.resources[i];
var ok = false;
for (0..p.resource_count) |j| {
const pr = p.resources[j];
if (pr.kind != r.kind) continue;
if (r.kind == @intFromEnum(device_abi.ResourceKind.irq)) {
if (pr.start == r.start) ok = true;
} else if (r.len != 0 and r.start >= pr.start and
r.start + r.len <= pr.start + pr.len) ok = true;
}
if (!ok) return false;
}
/// Whether a live task was spawned under `name` (its argv[0]) — read from the kernel
/// task table, the same snapshot `process_enumerate` exposes.
fn processRunning(name: []const u8) bool {
var table: [64]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table);
for (table[0..@min(total, table.len)]) |d| {
if (std.mem.eql(u8, d.name[0..d.name_length], name)) return true;
}
return children > 0; // bus must have published at least one
return false;
}
/// Device id of the HPET (the bus bus claims), from the same table drivers see.
/// Whether pci-bus registered at least one function under the PCI host bridge — proof
/// it came up, claimed the bridge, mapped its ECAM, and walked configuration space.
fn pciFunctionsRegistered() bool {
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var bridge_id: ?u64 = null;
for (buffer[0..n]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_host_bridge)) bridge_id = d.id;
}
const bid = bridge_id orelse return false;
for (buffer[0..n]) |d| {
if (d.parent == bid) return true;
}
return false;
}
/// Device id of the kernel-seeded HPET timer block (the node with a memory resource),
/// from the same device table drivers see. Used as a containment-test parent.
fn hpetDeviceId() ?u64 {
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
@@ -2511,8 +2474,9 @@ fn hpetDeviceId() ?u64 {
/// (so a dead driver's device goes quiet instead of storming) and the slot cleared
/// (so an ISR never posts a notification into the endpoint that is about to be freed).
///
/// This is the path `hpet` never takes — it runs forever — so it gets its own test.
/// Two properties, both read back from the hardware rather than from our own state:
/// A long-running driver that never exits wouldn't reach this teardown path, so it
/// gets its own test that binds and releases directly. Two properties, both read back
/// from the hardware rather than from our own state:
///
/// 1. A bound GSI is routed and unmasked.
/// 2. After `releaseOwner` for the binding's owner, that same entry is masked again.