docs: flip stale status markers across the tracks (audit found 23)

An all-tracks docs-vs-code audit (the same method that caught the storage
drift) found 23 confirmed inaccuracies where a doc's build-status claim no
longer matches the source — status markers that were never flipped after a
track landed, and a few paths left over from completed flag-days. All verified
against the code before editing; docs only, no behavior change.

The systemic ones:
- IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was
  not built and "device_claim = ring 0" / "memory-safe is not true yet". It is
  built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE
  rollback, dma_alloc buffers bound and torn down at death; fail-open only with
  no IOMMU). Restated; M16 marker flipped to done.
- The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv
  (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log ->
  /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb
  (process-management.md). Following the old paths silently breaks driver match.
- protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate
  fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built;
  SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer
  trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static
  hole -> dynamic per-registrar quota; init spawns fat -> volume-manager;
  config "hardcoded, move to /etc" -> already CSV data files; vdso.md three-
  value call; zig-self-hosting library/ layout; python argv "new" -> built.

Found and fixed by a multi-agent audit across 12 doc clusters, each finding
adversarially verified against the source.
This commit is contained in:
Daniel Samson
2026-08-09 22:04:17 +01:00
parent 8216be991d
commit bf0595763e
17 changed files with 134 additions and 105 deletions
@@ -136,7 +136,7 @@ xHCI match already uses); USB children carry the (class, subclass, protocol) tri
from usb-ids.zig — each bus's native language, decoded by the shared ids modules. from usb-ids.zig — each bus's native language, decoded by the shared ids modules.
(Since the registry landed, `child_added` also carries a `bus` discriminator and (Since the registry landed, `child_added` also carries a `bus` discriminator and
the numeric `vendor`/`device`/`subsystem` ids the finer match levels need — the numeric `vendor`/`device`/`subsystem` ids the finer match levels need —
see [/etc/devices.csv](devices-csv.md).) see [/system/configuration/devices.csv](devices-csv.md).)
## Supervision and restart ## Supervision and restart
@@ -227,7 +227,7 @@ published exit events, signals + `process`). On top of those:
class triple alone, so a virtio-gpu could only be matched as a generic display class triple alone, so a virtio-gpu could only be matched as a generic display
function and the driver had to re-confirm its `1AF4:1050` identity from config function and the driver had to re-confirm its `1AF4:1050` identity from config
space after being spawned. The manifest the earlier note anticipated landed as a space after being spawned. The manifest the earlier note anticipated landed as a
human-readable registry: **[/etc/devices.csv](devices-csv.md)**, parsed by the human-readable registry: **[/system/configuration/devices.csv](devices-csv.md)**, parsed by the
pure `device-registry` module and read by the manager at boot. A row binds a pure `device-registry` module and read by the manager at boot. A row binds a
driver to a device by any of base / subclass / prog-IF / vendor / device / driver to a device by any of base / subclass / prog-IF / vendor / device /
subsystem / `_HID`, most-specific match winning; it is authoritative (no subsystem / `_HID`, most-specific match winning; it is authoritative (no
@@ -1,6 +1,6 @@
# /etc/devices.csv — the device registry # /system/configuration/devices.csv — the device registry
**Status: built (2026-07-26).** The device manager reads `/etc/devices.csv` at **Status: built (2026-07-26).** The device manager reads `/system/configuration/devices.csv` at
boot and binds every device a bus driver reports to the driver the registry boot and binds every device a bus driver reports to the driver the registry
names. It replaces the three hand-written `switch` tables that used to live in names. It replaces the three hand-written `switch` tables that used to live in
the manager (`pciDriverForIdentity`, `hidDriverFor`, `usbDriverForIdentity`) — the manager (`pciDriverForIdentity`, `hidDriverFor`, `usbDriverForIdentity`) —
@@ -73,15 +73,15 @@ the first, so the shadowed rule is visible rather than silently dropped.
There is no compiled-in default table behind the registry. A device that no row There is no compiled-in default table behind the registry. A device that no row
matches goes **unbound** and is logged; the manager never guesses. A missing or matches goes **unbound** and is logged; the manager never guesses. A missing or
empty `/etc/devices.csv` therefore means nothing matches — which is loud at boot, empty `/system/configuration/devices.csv` therefore means nothing matches — which is loud at boot,
not a silent half-working system. not a silent half-working system.
## How the manager reads it ## How the manager reads it
`/etc/devices.csv` is bundled into the initial ramdisk (`build.zig`'s `bundled` `/system/configuration/devices.csv` is bundled into the initial ramdisk (`build.zig`'s `bundled`
list). The kernel serves the initrd's `/etc` tree directly — the `fat` service is list). The kernel serves the initrd's `/system/configuration` tree directly — the `fat` service is
spawned *after* the device manager and is irrelevant to `/etc` — so the manager spawned *after* the device manager and is irrelevant to `/system/configuration` — so the manager
reads the file with a plain `fs.open("/etc/devices.csv")` + `read`, with no reads the file with a plain `fs.open("/system/configuration/devices.csv")` + `read`, with no
filesystem service running and no boot-ordering dependency. It parses the bytes filesystem service running and no boot-ordering dependency. It parses the bytes
once in `initialise`, before any bus driver can report a device to match. once in `initialise`, before any bus driver can report a device to match.
@@ -105,6 +105,6 @@ in different namespaces — against the right `bus` column.
`build_support.programModule`). Then bundle it at `/system/drivers/<name>`: `build_support.programModule`). Then bundle it at `/system/drivers/<name>`:
one dependency + one bundled entry in the root `build.zig`, one line in the one dependency + one bundled entry in the root `build.zig`, one line in the
root `build.zig.zon`. root `build.zig.zon`.
2. Add a row to `etc/devices.csv` naming the identity it binds and its full path. 2. Add a row to `system/configuration/devices.csv` naming the identity it binds and its full path.
No device-manager change is required — the registry is the seam. No device-manager change is required — the registry is the seam.
+27 -21
View File
@@ -193,12 +193,14 @@ class driver, the device manager, or the kernel may share them freely.
a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with
a syscall per access. `io_port` resources were recorded by discovery and ignored — now a syscall per access. `io_port` resources were recorded by discovery and ignored — now
they're used. they're used.
- **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table, - **M16 (enforcement)** — the IOMMU is now *found and used*: discovery parses the ACPI
maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the DMAR table, maps the first VT-d unit, and reads its version + capabilities
platform info). This is detection only — **no translation domains are programmed, so (`iommu_present` in the platform info), and **`device_claim` programs a private
DMA is still unprotected** (the caveat below). Enforcement lands with the first DMA per-device translation domain** for the claimed function — rolling the claim back with
driver, which is what there is to protect and test against. Proven in the `iommu` test, `-ECONFINE` if it cannot confine it — so `dma_alloc` buffers are bound into that domain
booted with an emulated `intel-iommu`. and torn down at process death. DMA is protected on any IOMMU-equipped machine; the
system fails open only when there is no IOMMU at all. Proven in the `iommu` test, booted
with an emulated `intel-iommu`.
- **`system_spawn`** — a user-space supervisor starts a driver: - **`system_spawn`** — a user-space supervisor starts a driver:
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a `system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
fresh ring-3 process; `name` becomes the child's argv[0] and the optional fresh ring-3 process; `name` becomes the child's argv[0] and the optional
@@ -369,25 +371,29 @@ capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so an HPET timer could never Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so an HPET timer could never
exercise this path. The first MSI driver will be the first PCI driver. exercise this path. The first MSI driver will be the first PCI driver.
## M16 — the IOMMU, and the honest caveat ◑ detection done, enforcement pending ## M16 — the IOMMU, and the honest caveat ✅ done
*The IOMMU is now detected (DMAR parsed, VT-d unit mapped and read — see the `iommu` *The IOMMU is now detected (DMAR parsed, VT-d unit mapped and read — see the `iommu`
test), but **enforcement is not built**: no translation domains are programmed, so the test) **and enforced**: `device_claim` programs a private per-device VT-d/AMD-Vi
caveat below still holds in full. Detection can't be taken further usefully until there translation domain for the claimed function and rolls the claim back with `-ECONFINE` if
is a DMA driver to protect and QEMU's `intel-iommu` to test the protection against — it cannot confine it, `dma_alloc` buffers are bound into that domain and torn down at
building the per-device domains alongside that first driver is both the natural order process death, and the machine fails open only when it has no IOMMU at all. So the caveat
and the only way to verify them. The rest of this section is the original caveat.* below no longer holds except on IOMMU-less hardware. The rest of this section is the
original caveat, kept for the reasoning.*
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*. Everything above is capability-gated at the *CPU*. CPU page tables alone do not gate the
A driver that can program a bus-mastering engine can make that device write to any *device*: a driver that can program a bus-mastering engine could make that device write to
physical address, because page tables sit between the CPU and RAM, not between a device any physical address, because those page tables sit between the CPU and RAM, not between a
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, **`device_claim` device and RAM. That is exactly what the IOMMU closes. Now that VT-d/DMAR (and AMD-Vi; SMMU
on any DMA-capable device is equivalent to granting ring 0.** on ARM) is programmed, **`device_claim` confines the function into a private translation
domain** and rolls the claim back with `-ECONFINE` if it cannot — so a claimed DMA-capable
device is no longer equivalent to granting ring 0.
This does not make the model useless — it's the same position Linux is in with the This puts the model ahead of Linux-with-the-IOMMU-off: with an IOMMU present,
IOMMU off, and every other guarantee (crash isolation, restart, no shared address "user-space drivers are memory-safe" now holds, alongside every other guarantee (crash
space) still holds. But "user-space drivers are memory-safe" is not true yet, and the isolation, restart, no shared address space). The one remaining gap — a machine with no
gap should be named rather than implied. IOMMU at all, where the system deliberately fails open — should be named rather than
implied.
## Ordering ## Ordering
+16 -15
View File
@@ -32,7 +32,7 @@ kernel ──spawns──► init (PID 1) ──spawns──► device-manag
spawns only init, the service supervisor: the driver supervisor: enumerates spawns only init, the service supervisor: the driver supervisor: enumerates
publishes the starts the system /system/devices, matches each device publishes the starts the system /system/devices, matches each device
initial-ramdisk services (device-manager, to a driver, and system_spawn's it initial-ramdisk services (device-manager, to a driver, and system_spawn's it
so user space can fat, logger, ...). Its so user space can volume-manager, logger, ...). Its
system_spawn from it list is init policy. system_spawn from it list is init policy.
``` ```
@@ -43,7 +43,7 @@ the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Every
- **init** ([system/services/init](system/services/init/init.zig)) is the **service - **init** ([system/services/init](system/services/init/init.zig)) is the **service
supervisor**. It spawns the system services danos brings up at boot — today `input`, supervisor**. It spawns the system services danos brings up at boot — today `input`,
the `device-manager`, `fat`, `display`, `display-demo`, and the `logger` — from a the `device-manager`, `volume-manager`, `display`, `display-demo`, and the `logger` — from a
small list. Drivers are deliberately *not* its job. (An earlier draft listed a `vfs` small list. Drivers are deliberately *not* its job. (An earlier draft listed a `vfs`
service here; that service is retired — the router moved into the kernel as service here; that service is retired — the router moved into the kernel as
`fs_resolve`.) `fs_resolve`.)
@@ -66,9 +66,10 @@ the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Every
So "how is a driver discovered and configured" has two halves: **discovery** is the So "how is a driver discovered and configured" has two halves: **discovery** is the
kernel's device table, read by anyone; **configuration** is two user-space policies — kernel's device table, read by anyone; **configuration** is two user-space policies —
init's service list and the device-manager's match table. Both are hardcoded in their init's service list and the device-manager's match table. Both are now data, not code —
respective programs today; the natural next step is to move them into `/etc` (see the init reads `/system/configuration/init.csv` and the device manager reads
milestone notes in [driver-model.md](driver-model.md)). `system_spawn` is currently `/system/configuration/devices.csv`, each parsed at startup (the compiled-in switch
tables are gone). `system_spawn` is currently
ungated — any process may spawn any bundled binary — because there is no spawn ungated — any process may spawn any bundled binary — because there is no spawn
capability yet. capability yet.
@@ -301,12 +302,13 @@ process releases its claims and IRQ/MSI bindings — `releaseAllOwnedBy`,
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means - **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
granting one grants the other. A `device_register`ed child's *resource* can be narrower granting one grants the other. A `device_register`ed child's *resource* can be narrower
than a page, but its *mapping* can't. than a page, but its *mapping* can't.
- **DMA is not contained.** A driver that can program a bus-mastering device can make - **DMA is contained — except on a machine with no IOMMU.** A driver that can program a
that device write to *any* physical address — page tables don't sit between a device bus-mastering device could make that device write to *any* physical address — page
and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains tables don't sit between a device and RAM; an IOMMU does. `device_claim` now confines
are programmed, so `device_claim` on a DMA-capable device is still effectively each claimed PCI function into its own VT-d/AMD-Vi translation domain and rolls the
equivalent to granting ring 0. This is the largest gap between the design's promise and claim back with `-ECONFINE` if it can't (`system/kernel/process.zig`); DMA buffers are
what it delivers; enforcement lands with the first DMA driver. bound into that domain and torn down at process death. The residual gap is fail-open:
where the machine exposes **no IOMMU at all**, a DMA-capable claim still reaches RAM.
- **No voluntary `dev_release`.** A *live* driver can't drop a claim — only exit - **No voluntary `dev_release`.** A *live* driver can't drop a claim — only exit
releases it (any path out of a process runs `releaseAllOwnedBy`) — so handing a releases it (any path out of a process runs `releaseAllOwnedBy`) — so handing a
device between running drivers still means exiting. device between running drivers still means exiting.
@@ -365,10 +367,9 @@ $ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree
## What's next (not done here) ## What's next (not done here)
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI, The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as and per-device IOMMU confinement — are **now done** ([driver-model.md](driver-model.md),
is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550 M13–M16), as is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on PS/2 or 16550 driver possible). What's left is a handful of smaller items:
the first DMA driver to protect and test against) and these smaller items:
- **Releasing a claim** — half done. The kernel now drops *all* of a dead driver's - **Releasing a claim** — half done. The kernel now drops *all* of a dead driver's
claims on every path out of a process (`releaseAllOwnedBy`, called from process claims on every path out of a process (`releaseAllOwnedBy`, called from process
@@ -162,7 +162,7 @@ Without the boot-tree row the binary never reaches the image and the
device-manager has nothing to spawn. (The package also builds standalone: device-manager has nothing to spawn. (The package also builds standalone:
`cd system/drivers/intel-uhd-graphics-750 && zig build`.) `cd system/drivers/intel-uhd-graphics-750 && zig build`.)
## 3. Add the match rule to `etc/devices.csv` ## 3. Add the match rule to `system/configuration/devices.csv`
One row: bus, class triplet, vendor/device, driver path. **Copy the class One row: bus, class triplet, vendor/device, driver path. **Copy the class
triplet from the pci-bus boot log line, not from another row** — for the iGPU triplet from the pci-bus boot log line, not from another row** — for the iGPU
@@ -195,9 +195,9 @@ mapping — with zero risk to the hardware.
## 5. Verify the plumbing ## 5. Verify the plumbing
- `zig build test` still passes. - `zig build test` still passes.
- On the image: `/var/log/<boot-stamp>/system/services/device-manager.log` - On the image: `/system/logs/<boot-stamp>/system/services/device-manager.log`
shows `spawned <name> for device <N>`, and shows `spawned <name> for device <N>`, and
`/var/log/<boot-stamp>/system/drivers/<name>.log` holds the resource list and `/system/logs/<boot-stamp>/system/drivers/<name>.log` holds the resource list and
your first read. your first read.
- If the driver did not spawn, diagnose in this order: binary on the image - If the driver did not spawn, diagnose in this order: binary on the image
(step 2) → CSV row matches the log line exactly (step 3) → path identical in (step 2) → CSV row matches the log line exactly (step 3) → path identical in
@@ -32,8 +32,9 @@ Read-only and writable halves of `/system`: the program subtrees (`kernel`,
initrd-backed or synthetic — while `configuration` and `logs` are mutable initrd-backed or synthetic — while `configuration` and `logs` are mutable
machine state served by the boot-volume FAT backend. The kernel's machine state served by the boot-volume FAT backend. The kernel's
reserved-prefix rule (no mount may shadow `/system`, `/test`, or `/protocol`) reserved-prefix rule (no mount may shadow `/system`, `/test`, or `/protocol`)
needs a carve-out for exactly these two writable subtrees; that lands with the carves out exactly these two writable subtrees (`initrd_carve_outs` in
path migration below. `system/kernel/vfs.zig`), so the boot volume mounts them while every other
`/system` path stays initrd-served.
Deliberately not defined yet: a temporary-files location and per-application Deliberately not defined yet: a temporary-files location and per-application
mutable storage. Both belong to the `/applications` design and will be mutable storage. Both belong to the `/applications` design and will be
@@ -74,10 +75,10 @@ expect, mapped onto the real tree; the tree itself stays danos-native.
## Migration ## Migration
The tree above is the specification; some code still writes the unix paths it The tree above is the specification, and the code writes it. A flag-day already
replaced. The flag-day converting them: converted the unix paths it replaced:
| Today (in code) | Becomes | Where | | Was | Now | Where |
|------------------------------------------|-------------------------------------------|-----------------------------------------------------------------| |------------------------------------------|-------------------------------------------|-----------------------------------------------------------------|
| `/etc/init.csv` | `/system/configuration/init.csv` | `system/services/init/init.zig` | | `/etc/init.csv` | `/system/configuration/init.csv` | `system/services/init/init.zig` |
| `/etc/devices.csv` | `/system/configuration/devices.csv` | `system/services/device-manager/device-manager.zig` | | `/etc/devices.csv` | `/system/configuration/devices.csv` | `system/services/device-manager/device-manager.zig` |
@@ -85,6 +86,6 @@ replaced. The flag-day converting them:
| `/mnt/usb` | `/volumes/usb` | `system/services/fat/fat.zig`, the fat/vfs tests | | `/mnt/usb` | `/volumes/usb` | `system/services/fat/fat.zig`, the fat/vfs tests |
| `ServiceId` lookup | resolve + open under `/protocol` | every service and client; [protocol-namespace.md](../os-development/protocol-namespace.md) | | `ServiceId` lookup | resolve + open under `/protocol` | every service and client; [protocol-namespace.md](../os-development/protocol-namespace.md) |
The boot-image builder and the on-volume directory layout move in the same The boot-image builder and the on-volume directory layout moved in the same
change, so a freshly written image and the paths the services expect never change, so a freshly written image and the paths the services expect never
disagree. disagree.
+10 -5
View File
@@ -89,14 +89,19 @@ comptime {
} }
``` ```
`maximum_domains = 64` and `maximum_devices = 64` agree today only by a sentence in a `maximum_domains = 64` and `maximum_devices = 64` once agreed only by a sentence in a
comment, and the agreement fails open. This is the clause with a live hole behind it, comment, and that agreement failed open — the clause with the live hole behind it, where
and the reason raising `maximum_devices` alone would be a privilege escalation rather raising `maximum_devices` alone left every device id past the end of `iommu.confined`
than a fix. unconfined while `confineDevice` still reported success, a privilege escalation rather
than a fix. The assert closed that: it held the two together while both stayed fixed, and
when the device table was later made dynamic — no `maximum_devices` any more, only a
per-registrar quota — that forced them apart, the assert having done its job. `confined`
now grows to cover every id the broker mints, and `confineDevice` refuses when it cannot
record a confinement rather than failing open.
## The worked bad case ## The worked bad case
`devices_broker.maximum_devices`, which had no comment at all: `devices_broker.maximum_devices`, which had no comment at all, before it was made dynamic:
```zig ```zig
/// bound: device nodes for the whole machine — firmware-discovered plus registered /// bound: device nodes for the whole machine — firmware-discovered plus registered
+10 -5
View File
@@ -118,17 +118,22 @@ regions `init` frees. Keeping that boot-protocol knowledge on the loader side is
deliberate — the kernel has no notion of "reclaimable" or of UEFI at all. deliberate — the kernel has no notion of "reclaimable" or of UEFI at all.
The one live piece in that memory is the boot stack the kernel starts on; the loader The one live piece in that memory is the boot stack the kernel starts on; the loader
leaves the single region containing it `reserved`, so `init` won't hand it out. A leaves the single region containing it `reserved`, so `init` won't hand it out. The
later step will move task 0 onto a kernel-owned stack, freeing that last ~1 MiB kernel is only on it for an instant, though — `_start`'s first instruction switches
region too (and giving user mode the clean stack it wants). to a kernel-owned 64 KiB stack in `.bss` (that context becomes task 0). The region
stays `reserved` because the loader's own `convertMemoryMap` was executing on that
stack when it reclassified the RAM, and, like the map buffers below, nothing frees it
yet.
## What's next (partly done since) ## What's next (partly done since)
- **Contiguous allocation** — done: `allocContiguous` scans for a run of clear - **Contiguous allocation** — done: `allocContiguous` scans for a run of clear
bits, with an optional physical ceiling for DMA (`dma_alloc` is its user), and bits, with an optional physical ceiling for DMA (`dma_alloc` is its user), and
`allocBelow` serves the SMP trampoline. `allocBelow` serves the SMP trampoline.
- **A kernel stack for task 0** — still open: the boot processor's idle task runs - **A kernel stack for task 0** — done: `_start`'s first instruction switches `rsp`
on the boot stack to this day, so that region can't be freed. to a kernel-owned 64 KiB stack in `.bss` (`bootstrap_stack`), and `scheduler.init`
registers that running context as task 0 — the kernel is on the loader's boot
stack for that one instruction and never again.
- **Freeing the `reserved` `loader_data`** (the boot-time map buffers) — still - **Freeing the `reserved` `loader_data`** (the boot-time map buffers) — still
open: the bitmap deliberately tracks those frames so they *can* be freed, but open: the bitmap deliberately tracks those frames so they *can* be freed, but
nothing frees them yet. nothing frees them yet.
+6 -6
View File
@@ -14,7 +14,7 @@ are mirrored to it explicitly (`system/kernel/kernel.zig`).
## The pipeline ## The pipeline
``` ```
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /var/log/<boot-stamp>/<binary-path>.log process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /system/logs/<boot-stamp>/<binary-path>.log
kernel log.print ─┘ │ kernel log.print ─┘ │
└▶ serial / 0xE9 sinks (QEMU, -Dserial) └▶ serial / 0xE9 sinks (QEMU, -Dserial)
``` ```
@@ -49,12 +49,12 @@ kernel log.print ─┘ │
5. **Persist.** The **logger service** (`system/services/logger`) drains the 5. **Persist.** The **logger service** (`system/services/logger`) drains the
ring every 250 ms and demultiplexes records into one file per source under ring every 250 ms and demultiplexes records into one file per source under
`/var/log/<boot-stamp>/`, e.g. `/system/logs/<boot-stamp>/`, e.g.
``` ```
/var/log/2026-07-21T150434Z/kernel.log /system/logs/2026-07-21T150434Z/kernel.log
/var/log/2026-07-21T150434Z/system/services/fat.log /system/logs/2026-07-21T150434Z/system/services/fat.log
/var/log/2026-07-21T150434Z/system/drivers/usb-storage.log /system/logs/2026-07-21T150434Z/system/drivers/usb-storage.log
``` ```
The boot stamp is the RTC anchor from `klog_status` (FAT-safe: no colons; a The boot stamp is the RTC anchor from `klog_status` (FAT-safe: no colons; a
@@ -95,6 +95,6 @@ written last so a reader only trusts a complete record).
- A write-spamming process can evict other processes' unread records from the - A write-spamming process can evict other processes' unread records from the
ring (a per-process quota is future work); the loss is at least visible via ring (a per-process quota is future work); the loss is at least visible via
sequence gaps in every affected file. sequence gaps in every affected file.
- `/var/log` files have no privacy until the VFS grows permissions. - `/system/logs` files have no privacy until the VFS grows permissions.
- Records emitted after the logger's final shutdown drain reach serial and the - Records emitted after the logger's final shutdown drain reach serial and the
ring but not the files. ring but not the files.
+6 -4
View File
@@ -14,7 +14,7 @@ process-manager server, and Fuchsia/seL4 control processes only through handles.
danos rules out `/proc` **as the primitive**: the path router lives in the danos rules out `/proc` **as the primitive**: the path router lives in the
kernel (`fs_resolve`), but what is mounted under a path is served by a kernel (`fs_resolve`), but what is mounted under a path is served by a
user-process filesystem server (the way FAT serves `/mnt/usb`) — a `/proc` user-process filesystem server (the way FAT serves `/volumes/usb`) — a `/proc`
would be one more such server, which would put a user process in the path of would be one more such server, which would put a user process in the path of
process control. If that server (or anything under it) hangs, nothing could be process control. If that server (or anything under it) hangs, nothing could be
listed or killed, *including the hung server*. The control plane for processes listed or killed, *including the hung server*. The control plane for processes
@@ -117,9 +117,11 @@ the architecture layer calls up into `tick`.
`process_exit_reason` (`process.exitReason`). This is the input to `process_exit_reason` (`process.exitReason`). This is the input to
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code* restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
for the clean case can still ride alongside later. for the clean case can still ride alongside later.
- Enumerate writes through the caller's raw pointer under the bring-up trust - ~~Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an model, like `device_enumerate`~~ Closed (8d4a7cf): both `process_enumerate`
isolation break). and `device_enumerate` describe a chunk into a kernel buffer and place it with
`copyToUser`, which validates the range and resolves each page — an unmapped
page returns `-EFAULT`, and the kernel never stores through the user pointer.
## Tests ## Tests
+3 -3
View File
@@ -1,9 +1,9 @@
# The protocol namespace # The protocol namespace
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P3 of the *Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P4 of the
migration plan at the end have landed (the envelope, the registry and the migration plan at the end have landed (the envelope, the registry and the
`ServiceId` flag-day, and restriction stage one); P4 and P5 are the remaining `ServiceId` flag-day, restriction stage one, and the protocol rebase); P5
work list.* (restriction stage two) is the remaining work.*
How a program finds, connects to, and is restricted from the things it talks to. How a program finds, connects to, and is restricted from the things it talks to.
Three ideas, kept deliberately separate: Three ideas, kept deliberately separate:
+8 -6
View File
@@ -23,8 +23,8 @@ INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** — LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
the `smp` self-test confirms worker tasks executing on all four cores at once under the `smp` self-test confirms worker tasks executing on all four cores at once under
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs, kernel lock. What's left is refinement, not first-light: per-core run queues and IPIs
and thread-to-core affinity (see [Implementation status](#implementation-status)). (see [Implementation status](#implementation-status)).
## The common microkernel instinct: don't share kernel state ## The common microkernel instinct: don't share kernel state
@@ -238,10 +238,12 @@ next lands.
- **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock - **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock
contention ever bites. (Thread *affinity* already exists — see above; this is the contention ever bites. (Thread *affinity* already exists — see above; this is the
further step of giving each core its own primary run queue for load distribution.) further step of giving each core its own primary run queue for load distribution.)
- **Fault recovery** — today a fault halts (only) the faulting core. Turning that into - **Fault recovery** — a ring-3 fault already **kills the faulting process and keeps the
"kill the task, keep the core running" is the [resilience](resilience.md) track (it core (and the rest of the system) running**: the kernel trapped it on the task's own
needs the task's lock/resource state handled), and for taking a core fully offline, kernel stack, reclaims what the process held, and reschedules
its tasks migrated first. (`process.killCurrentProcess`; the `fault-recovery` test proves init keeps heartbeating
through the kill) — the [resilience](resilience.md) track. What's still open here is
taking a core fully **offline**, which additionally needs its tasks migrated off first.
## Further reading ## Further reading
+1 -1
View File
@@ -84,7 +84,7 @@ address space. Threads deliberately remove that boundary *within* a process:
there is no isolation **between** threads. there is no isolation **between** threads.
- Threads share fate — by contract: a fault in any thread, or a "kill the process" - Threads share fate — by contract: a fault in any thread, or a "kill the process"
decision, takes down **all** of them, so restartability lives at the process level, decision, takes down **all** of them, so restartability lives at the process level,
not the thread level. (The kernel does not yet enforce this fan-out — see the not the thread level. (The kernel enforces this fan-out — see the
Lifecycle note under Lifecycle note under
[Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).) [Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).)
- Shared mutable state reintroduces data races — the failure class the - Shared mutable state reintroduces data races — the failure class the
+13 -10
View File
@@ -36,9 +36,10 @@ and inside a VM alike; only the source behind it differs. The mechanism is in
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
model; that role now lives in [drivers.md](../device-driver-development/drivers.md), as documentation.) The one place model; that role now lives in [drivers.md](../device-driver-development/drivers.md), as documentation.) The one part
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed left to user space — **calendar policy** over wall-clock time (time zones, formatting) —
at the end; it is deliberately not built yet. is discussed at the end; the wall-clock *seconds* it builds on are a kernel syscall
(`wall_clock`), like the monotonic clock.
## The three system calls ## The three system calls
@@ -95,15 +96,17 @@ The raw wrappers (`clock`, `sleepMillis`, `timerOnce`) and the ergonomic
`Instant`/`Duration` layer both live in the `time` module `Instant`/`Duration` layer both live in the `time` module
(`library/kernel/time.zig`); the latter is what everyday code uses. (`library/kernel/time.zig`); the latter is what everyday code uses.
## Wall-clock time (not built) ## Wall-clock time
Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and
measurement, useless for "what is the date?" Calendar time — a real-time clock, time measurement, useless for "what is the date?" Calendar time needs a **real-time clock**.
zones, leap seconds — is genuinely a **user-space** concern, and it *is* the case a time The kernel owns wall-clock *seconds* as mechanism, exactly like the monotonic clock: the
service is for. It would be backed by an **RTC** driver (the CMOS real-time clock), not `wall_clock` syscall (#33) returns Unix epoch seconds (UTC). The CMOS **RTC** is read
the HPET, and exposed as a `CLOCK_REALTIME`-style service alongside the monotonic once at boot and anchored to the monotonic clock (`system/kernel/wall-clock.zig`), so a
syscall. It is deferred until something needs it; the monotonic clock the kernel already query is a cheap arithmetic offset rather than a per-call CMOS poll; `time`'s
owns covers every current use. `wallClock()` (`library/kernel/time.zig`) wraps it. Reading the hardware's value is not
policy — time zones, leap seconds, calendars, and formatting layer on top in user space.
It exists because the filesystem needs real timestamps (mtime).
## Verifying it ## Verifying it
+9 -6
View File
@@ -123,12 +123,15 @@ The calls that return two values in `rax:rdx` today — `dma_alloc`
`fs_resolve` (route tag + node token / backend handle) — `fs_resolve` (route tag + node token / backend handle) —
become functions returning a two-`u64` struct. The System V ABI returns a become functions returning a two-`u64` struct. The System V ABI returns a
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the 16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
C-ABI spelling of the existing convention, at zero cost. The one call that C-ABI spelling of the existing convention, at zero cost. The calls that
returns *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in return *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
`rdx`, received capability in `r8`) — exceeds the two-register return: its `rdx`, received capability in `r8`) and `dma_alloc` when the region is
function returns a three-`u64` struct, which the ABI passes via a hidden `shareable` (virtual in `rax`, physical in `rdx`, handle in `r8`) — exceed the
result pointer, so that one stub stores `rax`/`rdx`/`r8` through the pointer two-register return: their functions return a three-`u64` struct, which the
after the `syscall` — a few instructions rather than one. ABI passes via a hidden result pointer, so each stub stores `rax`/`rdx`/`r8`
through the pointer after the `syscall` — a few instructions rather than one.
(`ipc_call` likewise carries a received capability in `r8` alongside its `rax`
result.)
Grouped as `abi.zig` groups them: Grouped as `abi.zig` groups them:
+2 -2
View File
@@ -143,8 +143,8 @@ Python shell uses it.
The kernel/VFS cluster a shell forces (any shell, any language): The kernel/VFS cluster a shell forces (any shell, any language):
- **exec-of-path** — spawn an arbitrary VFS path, not a named ramdisk binary; - **exec-of-path** — spawn an arbitrary VFS path, not a named ramdisk binary;
- **argv/envp** — carried through spawn onto the child's entry stack (env from - **argv/envp** — carried through spawn onto the child's entry stack (argv
P1, argv new); already built and tested via `spawnWithArguments`; envp new, from P1);
- **numeric exit status** — extend the exit record beyond the categorical - **numeric exit status** — extend the exit record beyond the categorical
`ExitReason` (the gotcha the Zig roadmap flagged: `WEXITSTATUS` must be real); `ExitReason` (the gotcha the Zig roadmap flagged: `WEXITSTATUS` must be real);
- **fd inheritance + pipes** — a kernel or service pipe (a character device by - **fd inheritance + pipes** — a kernel or service pipe (a character device by
+3 -2
View File
@@ -151,8 +151,9 @@ Its footprint was tiny: **five** call sites, all `unistd` file operations —
(from the boot-log work) `init.zig` and `log-flush.zig`. `stdio.zig` was dead — nothing (from the boot-log work) `init.zig` and `log-flush.zig`. `stdio.zig` was dead — nothing
imported it. The plan — build `runtime.fs`, migrate those five to it, delete imported it. The plan — build `runtime.fs`, migrate those five to it, delete
`library/posix/`, and drop the `posix` module from `build.zig`'s `addUserBinary` — has `library/posix/`, and drop the `posix` module from `build.zig`'s `addUserBinary` — has
since been carried out: `library/` today holds only `mmio`, `runtime`, and since been carried out: `library/` today holds `client`, `csv`, `device`,
`xkeyboard-config`. `kernel`, `protocol`, and `xkeyboard-config` (the `runtime` namespace was later
reorganized into `library/kernel`, and `mmio` moved under `library/device`).
## Where danos stands: coverage vs. the gaps ## Where danos stands: coverage vs. the gaps