docs: flip stale status markers across the tracks (audit found 23)
An all-tracks docs-vs-code audit (the same method that caught the storage drift) found 23 confirmed inaccuracies where a doc's build-status claim no longer matches the source — status markers that were never flipped after a track landed, and a few paths left over from completed flag-days. All verified against the code before editing; docs only, no behavior change. The systemic ones: - IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was not built and "device_claim = ring 0" / "memory-safe is not true yet". It is built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE rollback, dma_alloc buffers bound and torn down at death; fail-open only with no IOMMU). Restated; M16 marker flipped to done. - The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log -> /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb (process-management.md). Following the old paths silently breaks driver match. - protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built; SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static hole -> dynamic per-registrar quota; init spawns fat -> volume-manager; config "hardcoded, move to /etc" -> already CSV data files; vdso.md three- value call; zig-self-hosting library/ layout; python argv "new" -> built. Found and fixed by a multi-agent audit across 12 doc clusters, each finding adversarially verified against the source.
This commit is contained in:
@@ -89,14 +89,19 @@ comptime {
|
||||
}
|
||||
```
|
||||
|
||||
`maximum_domains = 64` and `maximum_devices = 64` agree today only by a sentence in a
|
||||
comment, and the agreement fails open. This is the clause with a live hole behind it,
|
||||
and the reason raising `maximum_devices` alone would be a privilege escalation rather
|
||||
than a fix.
|
||||
`maximum_domains = 64` and `maximum_devices = 64` once agreed only by a sentence in a
|
||||
comment, and that agreement failed open — the clause with the live hole behind it, where
|
||||
raising `maximum_devices` alone left every device id past the end of `iommu.confined`
|
||||
unconfined while `confineDevice` still reported success, a privilege escalation rather
|
||||
than a fix. The assert closed that: it held the two together while both stayed fixed, and
|
||||
when the device table was later made dynamic — no `maximum_devices` any more, only a
|
||||
per-registrar quota — that forced them apart, the assert having done its job. `confined`
|
||||
now grows to cover every id the broker mints, and `confineDevice` refuses when it cannot
|
||||
record a confinement rather than failing open.
|
||||
|
||||
## The worked bad case
|
||||
|
||||
`devices_broker.maximum_devices`, which had no comment at all:
|
||||
`devices_broker.maximum_devices`, which had no comment at all, before it was made dynamic:
|
||||
|
||||
```zig
|
||||
/// bound: device nodes for the whole machine — firmware-discovered plus registered
|
||||
|
||||
@@ -118,17 +118,22 @@ regions `init` frees. Keeping that boot-protocol knowledge on the loader side is
|
||||
deliberate — the kernel has no notion of "reclaimable" or of UEFI at all.
|
||||
|
||||
The one live piece in that memory is the boot stack the kernel starts on; the loader
|
||||
leaves the single region containing it `reserved`, so `init` won't hand it out. A
|
||||
later step will move task 0 onto a kernel-owned stack, freeing that last ~1 MiB
|
||||
region too (and giving user mode the clean stack it wants).
|
||||
leaves the single region containing it `reserved`, so `init` won't hand it out. The
|
||||
kernel is only on it for an instant, though — `_start`'s first instruction switches
|
||||
to a kernel-owned 64 KiB stack in `.bss` (that context becomes task 0). The region
|
||||
stays `reserved` because the loader's own `convertMemoryMap` was executing on that
|
||||
stack when it reclassified the RAM, and, like the map buffers below, nothing frees it
|
||||
yet.
|
||||
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Contiguous allocation** — done: `allocContiguous` scans for a run of clear
|
||||
bits, with an optional physical ceiling for DMA (`dma_alloc` is its user), and
|
||||
`allocBelow` serves the SMP trampoline.
|
||||
- **A kernel stack for task 0** — still open: the boot processor's idle task runs
|
||||
on the boot stack to this day, so that region can't be freed.
|
||||
- **A kernel stack for task 0** — done: `_start`'s first instruction switches `rsp`
|
||||
to a kernel-owned 64 KiB stack in `.bss` (`bootstrap_stack`), and `scheduler.init`
|
||||
registers that running context as task 0 — the kernel is on the loader's boot
|
||||
stack for that one instruction and never again.
|
||||
- **Freeing the `reserved` `loader_data`** (the boot-time map buffers) — still
|
||||
open: the bitmap deliberately tracks those frames so they *can* be freed, but
|
||||
nothing frees them yet.
|
||||
|
||||
@@ -14,7 +14,7 @@ are mirrored to it explicitly (`system/kernel/kernel.zig`).
|
||||
## The pipeline
|
||||
|
||||
```
|
||||
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /var/log/<boot-stamp>/<binary-path>.log
|
||||
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /system/logs/<boot-stamp>/<binary-path>.log
|
||||
kernel log.print ─┘ │
|
||||
└▶ serial / 0xE9 sinks (QEMU, -Dserial)
|
||||
```
|
||||
@@ -49,12 +49,12 @@ kernel log.print ─┘ │
|
||||
|
||||
5. **Persist.** The **logger service** (`system/services/logger`) drains the
|
||||
ring every 250 ms and demultiplexes records into one file per source under
|
||||
`/var/log/<boot-stamp>/`, e.g.
|
||||
`/system/logs/<boot-stamp>/`, e.g.
|
||||
|
||||
```
|
||||
/var/log/2026-07-21T150434Z/kernel.log
|
||||
/var/log/2026-07-21T150434Z/system/services/fat.log
|
||||
/var/log/2026-07-21T150434Z/system/drivers/usb-storage.log
|
||||
/system/logs/2026-07-21T150434Z/kernel.log
|
||||
/system/logs/2026-07-21T150434Z/system/services/fat.log
|
||||
/system/logs/2026-07-21T150434Z/system/drivers/usb-storage.log
|
||||
```
|
||||
|
||||
The boot stamp is the RTC anchor from `klog_status` (FAT-safe: no colons; a
|
||||
@@ -95,6 +95,6 @@ written last so a reader only trusts a complete record).
|
||||
- A write-spamming process can evict other processes' unread records from the
|
||||
ring (a per-process quota is future work); the loss is at least visible via
|
||||
sequence gaps in every affected file.
|
||||
- `/var/log` files have no privacy until the VFS grows permissions.
|
||||
- `/system/logs` files have no privacy until the VFS grows permissions.
|
||||
- Records emitted after the logger's final shutdown drain reach serial and the
|
||||
ring but not the files.
|
||||
|
||||
@@ -14,7 +14,7 @@ process-manager server, and Fuchsia/seL4 control processes only through handles.
|
||||
|
||||
danos rules out `/proc` **as the primitive**: the path router lives in the
|
||||
kernel (`fs_resolve`), but what is mounted under a path is served by a
|
||||
user-process filesystem server (the way FAT serves `/mnt/usb`) — a `/proc`
|
||||
user-process filesystem server (the way FAT serves `/volumes/usb`) — a `/proc`
|
||||
would be one more such server, which would put a user process in the path of
|
||||
process control. If that server (or anything under it) hangs, nothing could be
|
||||
listed or killed, *including the hung server*. The control plane for processes
|
||||
@@ -117,9 +117,11 @@ the architecture layer calls up into `tick`.
|
||||
`process_exit_reason` (`process.exitReason`). This is the input to
|
||||
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
|
||||
for the clean case can still ride alongside later.
|
||||
- Enumerate writes through the caller's raw pointer under the bring-up trust
|
||||
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
|
||||
isolation break).
|
||||
- ~~Enumerate writes through the caller's raw pointer under the bring-up trust
|
||||
model, like `device_enumerate`~~ Closed (8d4a7cf): both `process_enumerate`
|
||||
and `device_enumerate` describe a chunk into a kernel buffer and place it with
|
||||
`copyToUser`, which validates the range and resolves each page — an unmapped
|
||||
page returns `-EFAULT`, and the kernel never stores through the user pointer.
|
||||
|
||||
## Tests
|
||||
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
# The protocol namespace
|
||||
|
||||
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P3 of the
|
||||
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P4 of the
|
||||
migration plan at the end have landed (the envelope, the registry and the
|
||||
`ServiceId` flag-day, and restriction stage one); P4 and P5 are the remaining
|
||||
work list.*
|
||||
`ServiceId` flag-day, restriction stage one, and the protocol rebase); P5
|
||||
(restriction stage two) is the remaining work.*
|
||||
|
||||
How a program finds, connects to, and is restricted from the things it talks to.
|
||||
Three ideas, kept deliberately separate:
|
||||
|
||||
@@ -23,8 +23,8 @@ INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor
|
||||
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
|
||||
the `smp` self-test confirms worker tasks executing on all four cores at once under
|
||||
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
|
||||
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
|
||||
and thread-to-core affinity (see [Implementation status](#implementation-status)).
|
||||
kernel lock. What's left is refinement, not first-light: per-core run queues and IPIs
|
||||
(see [Implementation status](#implementation-status)).
|
||||
|
||||
## The common microkernel instinct: don't share kernel state
|
||||
|
||||
@@ -238,10 +238,12 @@ next lands.
|
||||
- **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock
|
||||
contention ever bites. (Thread *affinity* already exists — see above; this is the
|
||||
further step of giving each core its own primary run queue for load distribution.)
|
||||
- **Fault recovery** — today a fault halts (only) the faulting core. Turning that into
|
||||
"kill the task, keep the core running" is the [resilience](resilience.md) track (it
|
||||
needs the task's lock/resource state handled), and for taking a core fully offline,
|
||||
its tasks migrated first.
|
||||
- **Fault recovery** — a ring-3 fault already **kills the faulting process and keeps the
|
||||
core (and the rest of the system) running**: the kernel trapped it on the task's own
|
||||
kernel stack, reclaims what the process held, and reschedules
|
||||
(`process.killCurrentProcess`; the `fault-recovery` test proves init keeps heartbeating
|
||||
through the kill) — the [resilience](resilience.md) track. What's still open here is
|
||||
taking a core fully **offline**, which additionally needs its tasks migrated off first.
|
||||
|
||||
## Further reading
|
||||
|
||||
|
||||
@@ -84,7 +84,7 @@ address space. Threads deliberately remove that boundary *within* a process:
|
||||
there is no isolation **between** threads.
|
||||
- Threads share fate — by contract: a fault in any thread, or a "kill the process"
|
||||
decision, takes down **all** of them, so restartability lives at the process level,
|
||||
not the thread level. (The kernel does not yet enforce this fan-out — see the
|
||||
not the thread level. (The kernel enforces this fan-out — see the
|
||||
Lifecycle note under
|
||||
[Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).)
|
||||
- Shared mutable state reintroduces data races — the failure class the
|
||||
|
||||
@@ -36,9 +36,10 @@ and inside a VM alike; only the source behind it differs. The mechanism is in
|
||||
|
||||
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
|
||||
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
|
||||
model; that role now lives in [drivers.md](../device-driver-development/drivers.md), as documentation.) The one place
|
||||
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
|
||||
at the end; it is deliberately not built yet.
|
||||
model; that role now lives in [drivers.md](../device-driver-development/drivers.md), as documentation.) The one part
|
||||
left to user space — **calendar policy** over wall-clock time (time zones, formatting) —
|
||||
is discussed at the end; the wall-clock *seconds* it builds on are a kernel syscall
|
||||
(`wall_clock`), like the monotonic clock.
|
||||
|
||||
## The three system calls
|
||||
|
||||
@@ -95,15 +96,17 @@ The raw wrappers (`clock`, `sleepMillis`, `timerOnce`) and the ergonomic
|
||||
`Instant`/`Duration` layer both live in the `time` module
|
||||
(`library/kernel/time.zig`); the latter is what everyday code uses.
|
||||
|
||||
## Wall-clock time (not built)
|
||||
## Wall-clock time
|
||||
|
||||
Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and
|
||||
measurement, useless for "what is the date?" Calendar time — a real-time clock, time
|
||||
zones, leap seconds — is genuinely a **user-space** concern, and it *is* the case a time
|
||||
service is for. It would be backed by an **RTC** driver (the CMOS real-time clock), not
|
||||
the HPET, and exposed as a `CLOCK_REALTIME`-style service alongside the monotonic
|
||||
syscall. It is deferred until something needs it; the monotonic clock the kernel already
|
||||
owns covers every current use.
|
||||
measurement, useless for "what is the date?" Calendar time needs a **real-time clock**.
|
||||
The kernel owns wall-clock *seconds* as mechanism, exactly like the monotonic clock: the
|
||||
`wall_clock` syscall (#33) returns Unix epoch seconds (UTC). The CMOS **RTC** is read
|
||||
once at boot and anchored to the monotonic clock (`system/kernel/wall-clock.zig`), so a
|
||||
query is a cheap arithmetic offset rather than a per-call CMOS poll; `time`'s
|
||||
`wallClock()` (`library/kernel/time.zig`) wraps it. Reading the hardware's value is not
|
||||
policy — time zones, leap seconds, calendars, and formatting layer on top in user space.
|
||||
It exists because the filesystem needs real timestamps (mtime).
|
||||
|
||||
## Verifying it
|
||||
|
||||
|
||||
@@ -123,12 +123,15 @@ The calls that return two values in `rax:rdx` today — `dma_alloc`
|
||||
`fs_resolve` (route tag + node token / backend handle) —
|
||||
become functions returning a two-`u64` struct. The System V ABI returns a
|
||||
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
|
||||
C-ABI spelling of the existing convention, at zero cost. The one call that
|
||||
returns *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
|
||||
`rdx`, received capability in `r8`) — exceeds the two-register return: its
|
||||
function returns a three-`u64` struct, which the ABI passes via a hidden
|
||||
result pointer, so that one stub stores `rax`/`rdx`/`r8` through the pointer
|
||||
after the `syscall` — a few instructions rather than one.
|
||||
C-ABI spelling of the existing convention, at zero cost. The calls that
|
||||
return *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
|
||||
`rdx`, received capability in `r8`) and `dma_alloc` when the region is
|
||||
`shareable` (virtual in `rax`, physical in `rdx`, handle in `r8`) — exceed the
|
||||
two-register return: their functions return a three-`u64` struct, which the
|
||||
ABI passes via a hidden result pointer, so each stub stores `rax`/`rdx`/`r8`
|
||||
through the pointer after the `syscall` — a few instructions rather than one.
|
||||
(`ipc_call` likewise carries a received capability in `r8` alongside its `rax`
|
||||
result.)
|
||||
|
||||
Grouped as `abi.zig` groups them:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user