Files
danos/docs/os-development/timers.md
T
Daniel Samson bf0595763e docs: flip stale status markers across the tracks (audit found 23)
An all-tracks docs-vs-code audit (the same method that caught the storage
drift) found 23 confirmed inaccuracies where a doc's build-status claim no
longer matches the source — status markers that were never flipped after a
track landed, and a few paths left over from completed flag-days. All verified
against the code before editing; docs only, no behavior change.

The systemic ones:
- IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was
  not built and "device_claim = ring 0" / "memory-safe is not true yet". It is
  built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE
  rollback, dma_alloc buffers bound and torn down at death; fail-open only with
  no IOMMU). Restated; M16 marker flipped to done.
- The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv
  (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log ->
  /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb
  (process-management.md). Following the old paths silently breaks driver match.
- protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate
  fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built;
  SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer
  trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static
  hole -> dynamic per-registrar quota; init spawns fat -> volume-manager;
  config "hardcoded, move to /etc" -> already CSV data files; vdso.md three-
  value call; zig-self-hosting library/ layout; python argv "new" -> built.

Found and fixed by a multi-agent audit across 12 doc clusters, each finding
adversarially verified against the source.
2026-08-09 22:04:17 +01:00

6.7 KiB

Timers and time

Two different needs hide under the word "timer", and danos keeps them apart:

  • Reading the clock — what time is it? A read of a free-running counter.
  • Waiting — wake me in N milliseconds, or notify me when a deadline passes.

Both are answered by the kernel, because the kernel already owns a timer: it has to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this are built in device-interrupts.md; the scheduler's blocking and wait queues are in scheduling.md. This page is about the surface a ring-3 program actually uses, and one deliberate absence: there is no user-space time service.

Why time is a syscall, not a service

The tempting microkernel move is to put a timer driver in user space and have applications ask it for the time over IPC. For a monotonic clock that is wrong — reading now() should never cost an IPC round trip. The kernel is already holding the answer: it computes the current time every time it schedules, from the TSC, in a couple of instructions. Surfacing that as a system call is pure mechanism; routing it through a message to another process would be slower and redundant, and a device like the HPET (uncacheable MMIO reads) is a particularly bad thing to read on every now().

This is the same conclusion every serious system reaches: Linux and Zircon read the counter in the vDSO, L4 exposes a clock field in a shared kernel page, seL4 reads the cycle counter directly. None of them make a clock read an IPC. danos makes it a syscall.

That "from the TSC" hides a portability question, because the TSC is only a valid clock when the CPU guarantees it is invariant and when every core's TSC is synchronized. danos checks both — the invariant-TSC CPUID bit (0x80000007 EDX[8], set on Intel and AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET counter when either fails. So now() stays accurate on a real Intel box, a real AMD box, and inside a VM alike; only the source behind it differs. The mechanism is in device-interrupts.md.

So the timer hardware lives in the kernel, and there is no hpet driver and no time server to consume. (An earlier HPET driver existed only to demonstrate the driver model; that role now lives in drivers.md, as documentation.) The one part left to user space — calendar policy over wall-clock time (time zones, formatting) — is discussed at the end; the wall-clock seconds it builds on are a kernel syscall (wall_clock), like the monotonic clock.

The three system calls

Time and waiting are three entries in the small syscall table (syscall.md):

  • clock (#23) → monotonic nanoseconds since boot. It only moves forward. Not wall-clock: no date, no timezone. Backed by architecture.nanos() (TSC, scaled with a 128-bit intermediate so a long uptime can't overflow) — a few nanoseconds of resolution, and just an rdtsc plus a multiply.
  • sleep (#3) → block the caller for N milliseconds. The scheduler records a wake deadline and the tick sweep wakes it (scheduler.sleep).
  • timer_bind (#31) → arm a one-shot timer that, after N milliseconds, posts a timer notification to an IPC endpoint. Unlike sleep it does not block: a service can keep answering messages on the same endpoint while a deadline is pending. This is the timed wait that stop-sequence escalation, hello deadlines, and restart backoff are built from (process-lifecycle.md, device-manager.md).

The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space; programs read the TSC through clock and get timed wakeups through sleep/timer_bind, both riding the scheduler tick.

time — the generic interface

Applications don't call the syscalls directly; they use time (library/kernel/time.zig), a thin Instant/Duration layer over them — an ergonomic front door, not new mechanism.

const time = @import("time");

const start = time.now();                 // Instant — monotonic
doWork();
const took = start.elapsed();             // Duration
time.sleep(time.Duration.fromMillis(5));  // block ~5 ms

// A deadline delivered as a notification, so a service keeps serving meanwhile:
_ = time.after(endpoint, time.Duration.fromMillis(200));
  • Duration is nanoseconds under the hood, with fromNanos/fromMicros/fromMillis/ fromSeconds and asNanos/asMillis. ceilMillis rounds up to the kernel's millisecond granularity, so a sub-millisecond sleep never rounds down to zero and returns early. All arithmetic saturates rather than wraps.
  • Instant is a point on the monotonic clock: since, elapsed, plus, reached — built for deadline loops (while (!deadline.reached()) …).
  • now() / monotonicNanos() wrap clock. available() reports whether the clock is calibrated at all (the kernel returns 0 until the TSC frequency is known, so a caller that needs real time can treat 0 as "unavailable" rather than assume it advances).
  • sleep(d) wraps sleep; spin(d) busy-polls now() for the sub-millisecond delays the millisecond tick can't express; after(endpoint, d) wraps timer_bind.

The raw wrappers (clock, sleepMillis, timerOnce) and the ergonomic Instant/Duration layer both live in the time module (library/kernel/time.zig); the latter is what everyday code uses.

Wall-clock time

Everything above is monotonic: elapsed time since boot, perfect for timeouts and measurement, useless for "what is the date?" Calendar time needs a real-time clock. The kernel owns wall-clock seconds as mechanism, exactly like the monotonic clock: the wall_clock syscall (#33) returns Unix epoch seconds (UTC). The CMOS RTC is read once at boot and anchored to the monotonic clock (system/kernel/wall-clock.zig), so a query is a cheap arithmetic offset rather than a per-call CMOS poll; time's wallClock() (library/kernel/time.zig) wraps it. Reading the hardware's value is not policy — time zones, leap seconds, calendars, and formatting layer on top in user space. It exists because the filesystem needs real timestamps (mtime).

Verifying it

time's Instant/Duration arithmetic has unit tests that run on the host:

$ zig build test        # includes library/kernel/time.zig

End to end, the proof the clock is real is that it advances: read now(), sleep a Duration, read now() again, and the second reading is later — the kernel's timer driving a ring-3 program with no service in between.