An all-tracks docs-vs-code audit (the same method that caught the storage drift) found 23 confirmed inaccuracies where a doc's build-status claim no longer matches the source — status markers that were never flipped after a track landed, and a few paths left over from completed flag-days. All verified against the code before editing; docs only, no behavior change. The systemic ones: - IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was not built and "device_claim = ring 0" / "memory-safe is not true yet". It is built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE rollback, dma_alloc buffers bound and torn down at death; fail-open only with no IOMMU). Restated; M16 marker flipped to done. - The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log -> /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb (process-management.md). Following the old paths silently breaks driver match. - protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built; SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static hole -> dynamic per-registrar quota; init spawns fat -> volume-manager; config "hardcoded, move to /etc" -> already CSV data files; vdso.md three- value call; zig-self-hosting library/ layout; python argv "new" -> built. Found and fixed by a multi-agent audit across 12 doc clusters, each finding adversarially verified against the source.
6.7 KiB
Timers and time
Two different needs hide under the word "timer", and danos keeps them apart:
- Reading the clock — what time is it? A read of a free-running counter.
- Waiting — wake me in N milliseconds, or notify me when a deadline passes.
Both are answered by the kernel, because the kernel already owns a timer: it has to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this are built in device-interrupts.md; the scheduler's blocking and wait queues are in scheduling.md. This page is about the surface a ring-3 program actually uses, and one deliberate absence: there is no user-space time service.
Why time is a syscall, not a service
The tempting microkernel move is to put a timer driver in user space and have
applications ask it for the time over IPC. For a monotonic clock that is wrong —
reading now() should never cost an IPC round trip. The kernel is already holding the
answer: it computes the current time every time it schedules, from the TSC, in a couple
of instructions. Surfacing that as a system call is pure mechanism; routing it through a
message to another process would be slower and redundant, and a device like the HPET
(uncacheable MMIO reads) is a particularly bad thing to read on every now().
This is the same conclusion every serious system reaches: Linux and Zircon read the counter in the vDSO, L4 exposes a clock field in a shared kernel page, seL4 reads the cycle counter directly. None of them make a clock read an IPC. danos makes it a syscall.
That "from the TSC" hides a portability question, because the TSC is only a valid clock
when the CPU guarantees it is invariant and when every core's TSC is synchronized.
danos checks both — the invariant-TSC CPUID bit (0x80000007 EDX[8], set on Intel and
AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET
counter when either fails. So now() stays accurate on a real Intel box, a real AMD box,
and inside a VM alike; only the source behind it differs. The mechanism is in
device-interrupts.md.
So the timer hardware lives in the kernel, and there is no hpet driver and no time
server to consume. (An earlier HPET driver existed only to demonstrate the driver
model; that role now lives in drivers.md, as documentation.) The one part
left to user space — calendar policy over wall-clock time (time zones, formatting) —
is discussed at the end; the wall-clock seconds it builds on are a kernel syscall
(wall_clock), like the monotonic clock.
The three system calls
Time and waiting are three entries in the small syscall table (syscall.md):
clock(#23) → monotonic nanoseconds since boot. It only moves forward. Not wall-clock: no date, no timezone. Backed byarchitecture.nanos()(TSC, scaled with a 128-bit intermediate so a long uptime can't overflow) — a few nanoseconds of resolution, and just anrdtscplus a multiply.sleep(#3) → block the caller for N milliseconds. The scheduler records a wake deadline and the tick sweep wakes it (scheduler.sleep).timer_bind(#31) → arm a one-shot timer that, after N milliseconds, posts a timer notification to an IPC endpoint. Unlikesleepit does not block: a service can keep answering messages on the same endpoint while a deadline is pending. This is the timed wait that stop-sequence escalation, hello deadlines, and restart backoff are built from (process-lifecycle.md, device-manager.md).
The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space;
programs read the TSC through clock and get timed wakeups through sleep/timer_bind,
both riding the scheduler tick.
time — the generic interface
Applications don't call the syscalls directly; they use time
(library/kernel/time.zig), a thin Instant/Duration layer over them — an ergonomic
front door, not new mechanism.
const time = @import("time");
const start = time.now(); // Instant — monotonic
doWork();
const took = start.elapsed(); // Duration
time.sleep(time.Duration.fromMillis(5)); // block ~5 ms
// A deadline delivered as a notification, so a service keeps serving meanwhile:
_ = time.after(endpoint, time.Duration.fromMillis(200));
Durationis nanoseconds under the hood, withfromNanos/fromMicros/fromMillis/ fromSecondsandasNanos/asMillis.ceilMillisrounds up to the kernel's millisecond granularity, so a sub-millisecondsleepnever rounds down to zero and returns early. All arithmetic saturates rather than wraps.Instantis a point on the monotonic clock:since,elapsed,plus,reached— built for deadline loops (while (!deadline.reached()) …).now()/monotonicNanos()wrapclock.available()reports whether the clock is calibrated at all (the kernel returns 0 until the TSC frequency is known, so a caller that needs real time can treat 0 as "unavailable" rather than assume it advances).sleep(d)wrapssleep;spin(d)busy-pollsnow()for the sub-millisecond delays the millisecond tick can't express;after(endpoint, d)wrapstimer_bind.
The raw wrappers (clock, sleepMillis, timerOnce) and the ergonomic
Instant/Duration layer both live in the time module
(library/kernel/time.zig); the latter is what everyday code uses.
Wall-clock time
Everything above is monotonic: elapsed time since boot, perfect for timeouts and
measurement, useless for "what is the date?" Calendar time needs a real-time clock.
The kernel owns wall-clock seconds as mechanism, exactly like the monotonic clock: the
wall_clock syscall (#33) returns Unix epoch seconds (UTC). The CMOS RTC is read
once at boot and anchored to the monotonic clock (system/kernel/wall-clock.zig), so a
query is a cheap arithmetic offset rather than a per-call CMOS poll; time's
wallClock() (library/kernel/time.zig) wraps it. Reading the hardware's value is not
policy — time zones, leap seconds, calendars, and formatting layer on top in user space.
It exists because the filesystem needs real timestamps (mtime).
Verifying it
time's Instant/Duration arithmetic has unit tests that run on the host:
$ zig build test # includes library/kernel/time.zig
End to end, the proof the clock is real is that it advances: read now(), sleep a
Duration, read now() again, and the second reading is later — the kernel's timer
driving a ring-3 program with no service in between.