docs: flip stale status markers across the tracks (audit found 23)
An all-tracks docs-vs-code audit (the same method that caught the storage drift) found 23 confirmed inaccuracies where a doc's build-status claim no longer matches the source — status markers that were never flipped after a track landed, and a few paths left over from completed flag-days. All verified against the code before editing; docs only, no behavior change. The systemic ones: - IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was not built and "device_claim = ring 0" / "memory-safe is not true yet". It is built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE rollback, dma_alloc buffers bound and torn down at death; fail-open only with no IOMMU). Restated; M16 marker flipped to done. - The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log -> /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb (process-management.md). Following the old paths silently breaks driver match. - protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built; SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static hole -> dynamic per-registrar quota; init spawns fat -> volume-manager; config "hardcoded, move to /etc" -> already CSV data files; vdso.md three- value call; zig-self-hosting library/ layout; python argv "new" -> built. Found and fixed by a multi-agent audit across 12 doc clusters, each finding adversarially verified against the source.
This commit is contained in:
@@ -136,7 +136,7 @@ xHCI match already uses); USB children carry the (class, subclass, protocol) tri
|
||||
from usb-ids.zig — each bus's native language, decoded by the shared ids modules.
|
||||
(Since the registry landed, `child_added` also carries a `bus` discriminator and
|
||||
the numeric `vendor`/`device`/`subsystem` ids the finer match levels need —
|
||||
see [/etc/devices.csv](devices-csv.md).)
|
||||
see [/system/configuration/devices.csv](devices-csv.md).)
|
||||
|
||||
## Supervision and restart
|
||||
|
||||
@@ -227,7 +227,7 @@ published exit events, signals + `process`). On top of those:
|
||||
class triple alone, so a virtio-gpu could only be matched as a generic display
|
||||
function and the driver had to re-confirm its `1AF4:1050` identity from config
|
||||
space after being spawned. The manifest the earlier note anticipated landed as a
|
||||
human-readable registry: **[/etc/devices.csv](devices-csv.md)**, parsed by the
|
||||
human-readable registry: **[/system/configuration/devices.csv](devices-csv.md)**, parsed by the
|
||||
pure `device-registry` module and read by the manager at boot. A row binds a
|
||||
driver to a device by any of base / subclass / prog-IF / vendor / device /
|
||||
subsystem / `_HID`, most-specific match winning; it is authoritative (no
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# /etc/devices.csv — the device registry
|
||||
# /system/configuration/devices.csv — the device registry
|
||||
|
||||
**Status: built (2026-07-26).** The device manager reads `/etc/devices.csv` at
|
||||
**Status: built (2026-07-26).** The device manager reads `/system/configuration/devices.csv` at
|
||||
boot and binds every device a bus driver reports to the driver the registry
|
||||
names. It replaces the three hand-written `switch` tables that used to live in
|
||||
the manager (`pciDriverForIdentity`, `hidDriverFor`, `usbDriverForIdentity`) —
|
||||
@@ -73,15 +73,15 @@ the first, so the shadowed rule is visible rather than silently dropped.
|
||||
|
||||
There is no compiled-in default table behind the registry. A device that no row
|
||||
matches goes **unbound** and is logged; the manager never guesses. A missing or
|
||||
empty `/etc/devices.csv` therefore means nothing matches — which is loud at boot,
|
||||
empty `/system/configuration/devices.csv` therefore means nothing matches — which is loud at boot,
|
||||
not a silent half-working system.
|
||||
|
||||
## How the manager reads it
|
||||
|
||||
`/etc/devices.csv` is bundled into the initial ramdisk (`build.zig`'s `bundled`
|
||||
list). The kernel serves the initrd's `/etc` tree directly — the `fat` service is
|
||||
spawned *after* the device manager and is irrelevant to `/etc` — so the manager
|
||||
reads the file with a plain `fs.open("/etc/devices.csv")` + `read`, with no
|
||||
`/system/configuration/devices.csv` is bundled into the initial ramdisk (`build.zig`'s `bundled`
|
||||
list). The kernel serves the initrd's `/system/configuration` tree directly — the `fat` service is
|
||||
spawned *after* the device manager and is irrelevant to `/system/configuration` — so the manager
|
||||
reads the file with a plain `fs.open("/system/configuration/devices.csv")` + `read`, with no
|
||||
filesystem service running and no boot-ordering dependency. It parses the bytes
|
||||
once in `initialise`, before any bus driver can report a device to match.
|
||||
|
||||
@@ -105,6 +105,6 @@ in different namespaces — against the right `bus` column.
|
||||
`build_support.programModule`). Then bundle it at `/system/drivers/<name>`:
|
||||
one dependency + one bundled entry in the root `build.zig`, one line in the
|
||||
root `build.zig.zon`.
|
||||
2. Add a row to `etc/devices.csv` naming the identity it binds and its full path.
|
||||
2. Add a row to `system/configuration/devices.csv` naming the identity it binds and its full path.
|
||||
|
||||
No device-manager change is required — the registry is the seam.
|
||||
|
||||
@@ -193,12 +193,14 @@ class driver, the device manager, or the kernel may share them freely.
|
||||
a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with
|
||||
a syscall per access. `io_port` resources were recorded by discovery and ignored — now
|
||||
they're used.
|
||||
- **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table,
|
||||
maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the
|
||||
platform info). This is detection only — **no translation domains are programmed, so
|
||||
DMA is still unprotected** (the caveat below). Enforcement lands with the first DMA
|
||||
driver, which is what there is to protect and test against. Proven in the `iommu` test,
|
||||
booted with an emulated `intel-iommu`.
|
||||
- **M16 (enforcement)** — the IOMMU is now *found and used*: discovery parses the ACPI
|
||||
DMAR table, maps the first VT-d unit, and reads its version + capabilities
|
||||
(`iommu_present` in the platform info), and **`device_claim` programs a private
|
||||
per-device translation domain** for the claimed function — rolling the claim back with
|
||||
`-ECONFINE` if it cannot confine it — so `dma_alloc` buffers are bound into that domain
|
||||
and torn down at process death. DMA is protected on any IOMMU-equipped machine; the
|
||||
system fails open only when there is no IOMMU at all. Proven in the `iommu` test, booted
|
||||
with an emulated `intel-iommu`.
|
||||
- **`system_spawn`** — a user-space supervisor starts a driver:
|
||||
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
|
||||
fresh ring-3 process; `name` becomes the child's argv[0] and the optional
|
||||
@@ -369,25 +371,29 @@ capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
|
||||
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so an HPET timer could never
|
||||
exercise this path. The first MSI driver will be the first PCI driver.
|
||||
|
||||
## M16 — the IOMMU, and the honest caveat ◑ detection done, enforcement pending
|
||||
## M16 — the IOMMU, and the honest caveat ✅ done
|
||||
|
||||
*The IOMMU is now detected (DMAR parsed, VT-d unit mapped and read — see the `iommu`
|
||||
test), but **enforcement is not built**: no translation domains are programmed, so the
|
||||
caveat below still holds in full. Detection can't be taken further usefully until there
|
||||
is a DMA driver to protect and QEMU's `intel-iommu` to test the protection against —
|
||||
building the per-device domains alongside that first driver is both the natural order
|
||||
and the only way to verify them. The rest of this section is the original caveat.*
|
||||
test) **and enforced**: `device_claim` programs a private per-device VT-d/AMD-Vi
|
||||
translation domain for the claimed function and rolls the claim back with `-ECONFINE` if
|
||||
it cannot confine it, `dma_alloc` buffers are bound into that domain and torn down at
|
||||
process death, and the machine fails open only when it has no IOMMU at all. So the caveat
|
||||
below no longer holds except on IOMMU-less hardware. The rest of this section is the
|
||||
original caveat, kept for the reasoning.*
|
||||
|
||||
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*.
|
||||
A driver that can program a bus-mastering engine can make that device write to any
|
||||
physical address, because page tables sit between the CPU and RAM, not between a device
|
||||
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, **`device_claim`
|
||||
on any DMA-capable device is equivalent to granting ring 0.**
|
||||
Everything above is capability-gated at the *CPU*. CPU page tables alone do not gate the
|
||||
*device*: a driver that can program a bus-mastering engine could make that device write to
|
||||
any physical address, because those page tables sit between the CPU and RAM, not between a
|
||||
device and RAM. That is exactly what the IOMMU closes. Now that VT-d/DMAR (and AMD-Vi; SMMU
|
||||
on ARM) is programmed, **`device_claim` confines the function into a private translation
|
||||
domain** and rolls the claim back with `-ECONFINE` if it cannot — so a claimed DMA-capable
|
||||
device is no longer equivalent to granting ring 0.
|
||||
|
||||
This does not make the model useless — it's the same position Linux is in with the
|
||||
IOMMU off, and every other guarantee (crash isolation, restart, no shared address
|
||||
space) still holds. But "user-space drivers are memory-safe" is not true yet, and the
|
||||
gap should be named rather than implied.
|
||||
This puts the model ahead of Linux-with-the-IOMMU-off: with an IOMMU present,
|
||||
"user-space drivers are memory-safe" now holds, alongside every other guarantee (crash
|
||||
isolation, restart, no shared address space). The one remaining gap — a machine with no
|
||||
IOMMU at all, where the system deliberately fails open — should be named rather than
|
||||
implied.
|
||||
|
||||
## Ordering
|
||||
|
||||
|
||||
@@ -32,7 +32,7 @@ kernel ──spawns──► init (PID 1) ──spawns──► device-manag
|
||||
spawns only init, the service supervisor: the driver supervisor: enumerates
|
||||
publishes the starts the system /system/devices, matches each device
|
||||
initial-ramdisk services (device-manager, to a driver, and system_spawn's it
|
||||
so user space can fat, logger, ...). Its
|
||||
so user space can volume-manager, logger, ...). Its
|
||||
system_spawn from it list is init policy.
|
||||
```
|
||||
|
||||
@@ -43,7 +43,7 @@ the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Every
|
||||
|
||||
- **init** ([system/services/init](system/services/init/init.zig)) is the **service
|
||||
supervisor**. It spawns the system services danos brings up at boot — today `input`,
|
||||
the `device-manager`, `fat`, `display`, `display-demo`, and the `logger` — from a
|
||||
the `device-manager`, `volume-manager`, `display`, `display-demo`, and the `logger` — from a
|
||||
small list. Drivers are deliberately *not* its job. (An earlier draft listed a `vfs`
|
||||
service here; that service is retired — the router moved into the kernel as
|
||||
`fs_resolve`.)
|
||||
@@ -66,9 +66,10 @@ the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Every
|
||||
|
||||
So "how is a driver discovered and configured" has two halves: **discovery** is the
|
||||
kernel's device table, read by anyone; **configuration** is two user-space policies —
|
||||
init's service list and the device-manager's match table. Both are hardcoded in their
|
||||
respective programs today; the natural next step is to move them into `/etc` (see the
|
||||
milestone notes in [driver-model.md](driver-model.md)). `system_spawn` is currently
|
||||
init's service list and the device-manager's match table. Both are now data, not code —
|
||||
init reads `/system/configuration/init.csv` and the device manager reads
|
||||
`/system/configuration/devices.csv`, each parsed at startup (the compiled-in switch
|
||||
tables are gone). `system_spawn` is currently
|
||||
ungated — any process may spawn any bundled binary — because there is no spawn
|
||||
capability yet.
|
||||
|
||||
@@ -301,12 +302,13 @@ process releases its claims and IRQ/MSI bindings — `releaseAllOwnedBy`,
|
||||
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
|
||||
granting one grants the other. A `device_register`ed child's *resource* can be narrower
|
||||
than a page, but its *mapping* can't.
|
||||
- **DMA is not contained.** A driver that can program a bus-mastering device can make
|
||||
that device write to *any* physical address — page tables don't sit between a device
|
||||
and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
|
||||
are programmed, so `device_claim` on a DMA-capable device is still effectively
|
||||
equivalent to granting ring 0. This is the largest gap between the design's promise and
|
||||
what it delivers; enforcement lands with the first DMA driver.
|
||||
- **DMA is contained — except on a machine with no IOMMU.** A driver that can program a
|
||||
bus-mastering device could make that device write to *any* physical address — page
|
||||
tables don't sit between a device and RAM; an IOMMU does. `device_claim` now confines
|
||||
each claimed PCI function into its own VT-d/AMD-Vi translation domain and rolls the
|
||||
claim back with `-ECONFINE` if it can't (`system/kernel/process.zig`); DMA buffers are
|
||||
bound into that domain and torn down at process death. The residual gap is fail-open:
|
||||
where the machine exposes **no IOMMU at all**, a DMA-capable claim still reaches RAM.
|
||||
- **No voluntary `dev_release`.** A *live* driver can't drop a claim — only exit
|
||||
releases it (any path out of a process runs `releaseAllOwnedBy`) — so handing a
|
||||
device between running drivers still means exiting.
|
||||
@@ -365,10 +367,9 @@ $ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree
|
||||
## What's next (not done here)
|
||||
|
||||
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
|
||||
and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as
|
||||
is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550
|
||||
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on
|
||||
the first DMA driver to protect and test against) and these smaller items:
|
||||
and per-device IOMMU confinement — are **now done** ([driver-model.md](driver-model.md),
|
||||
M13–M16), as is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a
|
||||
PS/2 or 16550 driver possible). What's left is a handful of smaller items:
|
||||
|
||||
- **Releasing a claim** — half done. The kernel now drops *all* of a dead driver's
|
||||
claims on every path out of a process (`releaseAllOwnedBy`, called from process
|
||||
|
||||
@@ -162,7 +162,7 @@ Without the boot-tree row the binary never reaches the image and the
|
||||
device-manager has nothing to spawn. (The package also builds standalone:
|
||||
`cd system/drivers/intel-uhd-graphics-750 && zig build`.)
|
||||
|
||||
## 3. Add the match rule to `etc/devices.csv`
|
||||
## 3. Add the match rule to `system/configuration/devices.csv`
|
||||
|
||||
One row: bus, class triplet, vendor/device, driver path. **Copy the class
|
||||
triplet from the pci-bus boot log line, not from another row** — for the iGPU
|
||||
@@ -195,9 +195,9 @@ mapping — with zero risk to the hardware.
|
||||
## 5. Verify the plumbing
|
||||
|
||||
- `zig build test` still passes.
|
||||
- On the image: `/var/log/<boot-stamp>/system/services/device-manager.log`
|
||||
- On the image: `/system/logs/<boot-stamp>/system/services/device-manager.log`
|
||||
shows `spawned <name> for device <N>`, and
|
||||
`/var/log/<boot-stamp>/system/drivers/<name>.log` holds the resource list and
|
||||
`/system/logs/<boot-stamp>/system/drivers/<name>.log` holds the resource list and
|
||||
your first read.
|
||||
- If the driver did not spawn, diagnose in this order: binary on the image
|
||||
(step 2) → CSV row matches the log line exactly (step 3) → path identical in
|
||||
|
||||
@@ -32,8 +32,9 @@ Read-only and writable halves of `/system`: the program subtrees (`kernel`,
|
||||
initrd-backed or synthetic — while `configuration` and `logs` are mutable
|
||||
machine state served by the boot-volume FAT backend. The kernel's
|
||||
reserved-prefix rule (no mount may shadow `/system`, `/test`, or `/protocol`)
|
||||
needs a carve-out for exactly these two writable subtrees; that lands with the
|
||||
path migration below.
|
||||
carves out exactly these two writable subtrees (`initrd_carve_outs` in
|
||||
`system/kernel/vfs.zig`), so the boot volume mounts them while every other
|
||||
`/system` path stays initrd-served.
|
||||
|
||||
Deliberately not defined yet: a temporary-files location and per-application
|
||||
mutable storage. Both belong to the `/applications` design and will be
|
||||
@@ -74,10 +75,10 @@ expect, mapped onto the real tree; the tree itself stays danos-native.
|
||||
|
||||
## Migration
|
||||
|
||||
The tree above is the specification; some code still writes the unix paths it
|
||||
replaced. The flag-day converting them:
|
||||
The tree above is the specification, and the code writes it. A flag-day already
|
||||
converted the unix paths it replaced:
|
||||
|
||||
| Today (in code) | Becomes | Where |
|
||||
| Was | Now | Where |
|
||||
|------------------------------------------|-------------------------------------------|-----------------------------------------------------------------|
|
||||
| `/etc/init.csv` | `/system/configuration/init.csv` | `system/services/init/init.zig` |
|
||||
| `/etc/devices.csv` | `/system/configuration/devices.csv` | `system/services/device-manager/device-manager.zig` |
|
||||
@@ -85,6 +86,6 @@ replaced. The flag-day converting them:
|
||||
| `/mnt/usb` | `/volumes/usb` | `system/services/fat/fat.zig`, the fat/vfs tests |
|
||||
| `ServiceId` lookup | resolve + open under `/protocol` | every service and client; [protocol-namespace.md](../os-development/protocol-namespace.md) |
|
||||
|
||||
The boot-image builder and the on-volume directory layout move in the same
|
||||
The boot-image builder and the on-volume directory layout moved in the same
|
||||
change, so a freshly written image and the paths the services expect never
|
||||
disagree.
|
||||
|
||||
@@ -89,14 +89,19 @@ comptime {
|
||||
}
|
||||
```
|
||||
|
||||
`maximum_domains = 64` and `maximum_devices = 64` agree today only by a sentence in a
|
||||
comment, and the agreement fails open. This is the clause with a live hole behind it,
|
||||
and the reason raising `maximum_devices` alone would be a privilege escalation rather
|
||||
than a fix.
|
||||
`maximum_domains = 64` and `maximum_devices = 64` once agreed only by a sentence in a
|
||||
comment, and that agreement failed open — the clause with the live hole behind it, where
|
||||
raising `maximum_devices` alone left every device id past the end of `iommu.confined`
|
||||
unconfined while `confineDevice` still reported success, a privilege escalation rather
|
||||
than a fix. The assert closed that: it held the two together while both stayed fixed, and
|
||||
when the device table was later made dynamic — no `maximum_devices` any more, only a
|
||||
per-registrar quota — that forced them apart, the assert having done its job. `confined`
|
||||
now grows to cover every id the broker mints, and `confineDevice` refuses when it cannot
|
||||
record a confinement rather than failing open.
|
||||
|
||||
## The worked bad case
|
||||
|
||||
`devices_broker.maximum_devices`, which had no comment at all:
|
||||
`devices_broker.maximum_devices`, which had no comment at all, before it was made dynamic:
|
||||
|
||||
```zig
|
||||
/// bound: device nodes for the whole machine — firmware-discovered plus registered
|
||||
|
||||
@@ -118,17 +118,22 @@ regions `init` frees. Keeping that boot-protocol knowledge on the loader side is
|
||||
deliberate — the kernel has no notion of "reclaimable" or of UEFI at all.
|
||||
|
||||
The one live piece in that memory is the boot stack the kernel starts on; the loader
|
||||
leaves the single region containing it `reserved`, so `init` won't hand it out. A
|
||||
later step will move task 0 onto a kernel-owned stack, freeing that last ~1 MiB
|
||||
region too (and giving user mode the clean stack it wants).
|
||||
leaves the single region containing it `reserved`, so `init` won't hand it out. The
|
||||
kernel is only on it for an instant, though — `_start`'s first instruction switches
|
||||
to a kernel-owned 64 KiB stack in `.bss` (that context becomes task 0). The region
|
||||
stays `reserved` because the loader's own `convertMemoryMap` was executing on that
|
||||
stack when it reclassified the RAM, and, like the map buffers below, nothing frees it
|
||||
yet.
|
||||
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Contiguous allocation** — done: `allocContiguous` scans for a run of clear
|
||||
bits, with an optional physical ceiling for DMA (`dma_alloc` is its user), and
|
||||
`allocBelow` serves the SMP trampoline.
|
||||
- **A kernel stack for task 0** — still open: the boot processor's idle task runs
|
||||
on the boot stack to this day, so that region can't be freed.
|
||||
- **A kernel stack for task 0** — done: `_start`'s first instruction switches `rsp`
|
||||
to a kernel-owned 64 KiB stack in `.bss` (`bootstrap_stack`), and `scheduler.init`
|
||||
registers that running context as task 0 — the kernel is on the loader's boot
|
||||
stack for that one instruction and never again.
|
||||
- **Freeing the `reserved` `loader_data`** (the boot-time map buffers) — still
|
||||
open: the bitmap deliberately tracks those frames so they *can* be freed, but
|
||||
nothing frees them yet.
|
||||
|
||||
@@ -14,7 +14,7 @@ are mirrored to it explicitly (`system/kernel/kernel.zig`).
|
||||
## The pipeline
|
||||
|
||||
```
|
||||
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /var/log/<boot-stamp>/<binary-path>.log
|
||||
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /system/logs/<boot-stamp>/<binary-path>.log
|
||||
kernel log.print ─┘ │
|
||||
└▶ serial / 0xE9 sinks (QEMU, -Dserial)
|
||||
```
|
||||
@@ -49,12 +49,12 @@ kernel log.print ─┘ │
|
||||
|
||||
5. **Persist.** The **logger service** (`system/services/logger`) drains the
|
||||
ring every 250 ms and demultiplexes records into one file per source under
|
||||
`/var/log/<boot-stamp>/`, e.g.
|
||||
`/system/logs/<boot-stamp>/`, e.g.
|
||||
|
||||
```
|
||||
/var/log/2026-07-21T150434Z/kernel.log
|
||||
/var/log/2026-07-21T150434Z/system/services/fat.log
|
||||
/var/log/2026-07-21T150434Z/system/drivers/usb-storage.log
|
||||
/system/logs/2026-07-21T150434Z/kernel.log
|
||||
/system/logs/2026-07-21T150434Z/system/services/fat.log
|
||||
/system/logs/2026-07-21T150434Z/system/drivers/usb-storage.log
|
||||
```
|
||||
|
||||
The boot stamp is the RTC anchor from `klog_status` (FAT-safe: no colons; a
|
||||
@@ -95,6 +95,6 @@ written last so a reader only trusts a complete record).
|
||||
- A write-spamming process can evict other processes' unread records from the
|
||||
ring (a per-process quota is future work); the loss is at least visible via
|
||||
sequence gaps in every affected file.
|
||||
- `/var/log` files have no privacy until the VFS grows permissions.
|
||||
- `/system/logs` files have no privacy until the VFS grows permissions.
|
||||
- Records emitted after the logger's final shutdown drain reach serial and the
|
||||
ring but not the files.
|
||||
|
||||
@@ -14,7 +14,7 @@ process-manager server, and Fuchsia/seL4 control processes only through handles.
|
||||
|
||||
danos rules out `/proc` **as the primitive**: the path router lives in the
|
||||
kernel (`fs_resolve`), but what is mounted under a path is served by a
|
||||
user-process filesystem server (the way FAT serves `/mnt/usb`) — a `/proc`
|
||||
user-process filesystem server (the way FAT serves `/volumes/usb`) — a `/proc`
|
||||
would be one more such server, which would put a user process in the path of
|
||||
process control. If that server (or anything under it) hangs, nothing could be
|
||||
listed or killed, *including the hung server*. The control plane for processes
|
||||
@@ -117,9 +117,11 @@ the architecture layer calls up into `tick`.
|
||||
`process_exit_reason` (`process.exitReason`). This is the input to
|
||||
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
|
||||
for the clean case can still ride alongside later.
|
||||
- Enumerate writes through the caller's raw pointer under the bring-up trust
|
||||
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
|
||||
isolation break).
|
||||
- ~~Enumerate writes through the caller's raw pointer under the bring-up trust
|
||||
model, like `device_enumerate`~~ Closed (8d4a7cf): both `process_enumerate`
|
||||
and `device_enumerate` describe a chunk into a kernel buffer and place it with
|
||||
`copyToUser`, which validates the range and resolves each page — an unmapped
|
||||
page returns `-EFAULT`, and the kernel never stores through the user pointer.
|
||||
|
||||
## Tests
|
||||
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
# The protocol namespace
|
||||
|
||||
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P3 of the
|
||||
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P4 of the
|
||||
migration plan at the end have landed (the envelope, the registry and the
|
||||
`ServiceId` flag-day, and restriction stage one); P4 and P5 are the remaining
|
||||
work list.*
|
||||
`ServiceId` flag-day, restriction stage one, and the protocol rebase); P5
|
||||
(restriction stage two) is the remaining work.*
|
||||
|
||||
How a program finds, connects to, and is restricted from the things it talks to.
|
||||
Three ideas, kept deliberately separate:
|
||||
|
||||
@@ -23,8 +23,8 @@ INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor
|
||||
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
|
||||
the `smp` self-test confirms worker tasks executing on all four cores at once under
|
||||
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
|
||||
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
|
||||
and thread-to-core affinity (see [Implementation status](#implementation-status)).
|
||||
kernel lock. What's left is refinement, not first-light: per-core run queues and IPIs
|
||||
(see [Implementation status](#implementation-status)).
|
||||
|
||||
## The common microkernel instinct: don't share kernel state
|
||||
|
||||
@@ -238,10 +238,12 @@ next lands.
|
||||
- **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock
|
||||
contention ever bites. (Thread *affinity* already exists — see above; this is the
|
||||
further step of giving each core its own primary run queue for load distribution.)
|
||||
- **Fault recovery** — today a fault halts (only) the faulting core. Turning that into
|
||||
"kill the task, keep the core running" is the [resilience](resilience.md) track (it
|
||||
needs the task's lock/resource state handled), and for taking a core fully offline,
|
||||
its tasks migrated first.
|
||||
- **Fault recovery** — a ring-3 fault already **kills the faulting process and keeps the
|
||||
core (and the rest of the system) running**: the kernel trapped it on the task's own
|
||||
kernel stack, reclaims what the process held, and reschedules
|
||||
(`process.killCurrentProcess`; the `fault-recovery` test proves init keeps heartbeating
|
||||
through the kill) — the [resilience](resilience.md) track. What's still open here is
|
||||
taking a core fully **offline**, which additionally needs its tasks migrated off first.
|
||||
|
||||
## Further reading
|
||||
|
||||
|
||||
@@ -84,7 +84,7 @@ address space. Threads deliberately remove that boundary *within* a process:
|
||||
there is no isolation **between** threads.
|
||||
- Threads share fate — by contract: a fault in any thread, or a "kill the process"
|
||||
decision, takes down **all** of them, so restartability lives at the process level,
|
||||
not the thread level. (The kernel does not yet enforce this fan-out — see the
|
||||
not the thread level. (The kernel enforces this fan-out — see the
|
||||
Lifecycle note under
|
||||
[Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).)
|
||||
- Shared mutable state reintroduces data races — the failure class the
|
||||
|
||||
@@ -36,9 +36,10 @@ and inside a VM alike; only the source behind it differs. The mechanism is in
|
||||
|
||||
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
|
||||
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
|
||||
model; that role now lives in [drivers.md](../device-driver-development/drivers.md), as documentation.) The one place
|
||||
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
|
||||
at the end; it is deliberately not built yet.
|
||||
model; that role now lives in [drivers.md](../device-driver-development/drivers.md), as documentation.) The one part
|
||||
left to user space — **calendar policy** over wall-clock time (time zones, formatting) —
|
||||
is discussed at the end; the wall-clock *seconds* it builds on are a kernel syscall
|
||||
(`wall_clock`), like the monotonic clock.
|
||||
|
||||
## The three system calls
|
||||
|
||||
@@ -95,15 +96,17 @@ The raw wrappers (`clock`, `sleepMillis`, `timerOnce`) and the ergonomic
|
||||
`Instant`/`Duration` layer both live in the `time` module
|
||||
(`library/kernel/time.zig`); the latter is what everyday code uses.
|
||||
|
||||
## Wall-clock time (not built)
|
||||
## Wall-clock time
|
||||
|
||||
Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and
|
||||
measurement, useless for "what is the date?" Calendar time — a real-time clock, time
|
||||
zones, leap seconds — is genuinely a **user-space** concern, and it *is* the case a time
|
||||
service is for. It would be backed by an **RTC** driver (the CMOS real-time clock), not
|
||||
the HPET, and exposed as a `CLOCK_REALTIME`-style service alongside the monotonic
|
||||
syscall. It is deferred until something needs it; the monotonic clock the kernel already
|
||||
owns covers every current use.
|
||||
measurement, useless for "what is the date?" Calendar time needs a **real-time clock**.
|
||||
The kernel owns wall-clock *seconds* as mechanism, exactly like the monotonic clock: the
|
||||
`wall_clock` syscall (#33) returns Unix epoch seconds (UTC). The CMOS **RTC** is read
|
||||
once at boot and anchored to the monotonic clock (`system/kernel/wall-clock.zig`), so a
|
||||
query is a cheap arithmetic offset rather than a per-call CMOS poll; `time`'s
|
||||
`wallClock()` (`library/kernel/time.zig`) wraps it. Reading the hardware's value is not
|
||||
policy — time zones, leap seconds, calendars, and formatting layer on top in user space.
|
||||
It exists because the filesystem needs real timestamps (mtime).
|
||||
|
||||
## Verifying it
|
||||
|
||||
|
||||
@@ -123,12 +123,15 @@ The calls that return two values in `rax:rdx` today — `dma_alloc`
|
||||
`fs_resolve` (route tag + node token / backend handle) —
|
||||
become functions returning a two-`u64` struct. The System V ABI returns a
|
||||
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
|
||||
C-ABI spelling of the existing convention, at zero cost. The one call that
|
||||
returns *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
|
||||
`rdx`, received capability in `r8`) — exceeds the two-register return: its
|
||||
function returns a three-`u64` struct, which the ABI passes via a hidden
|
||||
result pointer, so that one stub stores `rax`/`rdx`/`r8` through the pointer
|
||||
after the `syscall` — a few instructions rather than one.
|
||||
C-ABI spelling of the existing convention, at zero cost. The calls that
|
||||
return *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
|
||||
`rdx`, received capability in `r8`) and `dma_alloc` when the region is
|
||||
`shareable` (virtual in `rax`, physical in `rdx`, handle in `r8`) — exceed the
|
||||
two-register return: their functions return a three-`u64` struct, which the
|
||||
ABI passes via a hidden result pointer, so each stub stores `rax`/`rdx`/`r8`
|
||||
through the pointer after the `syscall` — a few instructions rather than one.
|
||||
(`ipc_call` likewise carries a received capability in `r8` alongside its `rax`
|
||||
result.)
|
||||
|
||||
Grouped as `abi.zig` groups them:
|
||||
|
||||
|
||||
@@ -143,8 +143,8 @@ Python shell uses it.
|
||||
The kernel/VFS cluster a shell forces (any shell, any language):
|
||||
|
||||
- **exec-of-path** — spawn an arbitrary VFS path, not a named ramdisk binary;
|
||||
- **argv/envp** — carried through spawn onto the child's entry stack (env from
|
||||
P1, argv new);
|
||||
- **argv/envp** — carried through spawn onto the child's entry stack (argv
|
||||
already built and tested via `spawnWithArguments`; envp new, from P1);
|
||||
- **numeric exit status** — extend the exit record beyond the categorical
|
||||
`ExitReason` (the gotcha the Zig roadmap flagged: `WEXITSTATUS` must be real);
|
||||
- **fd inheritance + pipes** — a kernel or service pipe (a character device by
|
||||
|
||||
@@ -151,8 +151,9 @@ Its footprint was tiny: **five** call sites, all `unistd` file operations —
|
||||
(from the boot-log work) `init.zig` and `log-flush.zig`. `stdio.zig` was dead — nothing
|
||||
imported it. The plan — build `runtime.fs`, migrate those five to it, delete
|
||||
`library/posix/`, and drop the `posix` module from `build.zig`'s `addUserBinary` — has
|
||||
since been carried out: `library/` today holds only `mmio`, `runtime`, and
|
||||
`xkeyboard-config`.
|
||||
since been carried out: `library/` today holds `client`, `csv`, `device`,
|
||||
`kernel`, `protocol`, and `xkeyboard-config` (the `runtime` namespace was later
|
||||
reorganized into `library/kernel`, and `mmio` moved under `library/device`).
|
||||
|
||||
## Where danos stands: coverage vs. the gaps
|
||||
|
||||
|
||||
Reference in New Issue
Block a user