C6: update docs for the library/kernel + client split

Rewrote the repository-layout and driver-model docs to describe the new tree —
library/kernel (the kernel32-style system library: syscall surface split by
concern), library/device (mmio/model/pci/usb/acpi/driver/block), library/client
(display, input service clients), library/protocol (wire contracts) — and swept
the reference docs off the retired runtime shim:

  runtime.system.*  -> logging.* / time.* / process.* / memory.*
  runtime.dma.* / runtime.shared_memory.* / runtime.allocator -> memory.*
  runtime.ipc/process/time/service/Thread/block/display/input -> the module name
  runtime.device / runtime.device_manager -> driver
  library/runtime/<file>.zig -> its new home (kernel/ client/ device/)

docs/README.md (repository layout + source map), docs/driver-model.md (the module
graph + import lists), and the concern/reference docs (ipc, threading, timers,
logging, power, process-lifecycle, device-manager, display, vdso, sysv, drivers,
input, vfs-protocol, coding-standards, ...) now reflect the split. "runtime" that
remains is the userspace-library *concept*, which is still accurate.

The historical plan docs (display-plan, display-v2-plan, threading-plan) and the
zig-self-hosting design note are left as point-in-time snapshots.
This commit is contained in:
Daniel Samson
2026-07-22 23:42:25 +01:00
parent 37326c7664
commit 6271278d4d
17 changed files with 110 additions and 95 deletions
+25 -18
View File
@@ -73,7 +73,7 @@ rather than restate it. Roughly in the order things happen at runtime:
19. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
(M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
`runtime.process` interface, exit reasons, published exit events any stateful
`process` module interface, exit reasons, published exit events any stateful
service can subscribe to (the VFS releasing dead clients' handles), and the two
iron rules (cleanup is the kernel's job; kill is not a signal).
20. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
@@ -114,11 +114,11 @@ Start with the north star:
- **[zig-self-hosting.md](zig-self-hosting.md) — running Zig on danos.** A design note
(not built yet) on making danos a real Zig target (`-target x86_64-danos`) and
eventually running the compiler on it. The key realisation: Zig 0.16 reduces an OS
port to **one seam** (`std.os.danos`), so we build `runtime.os` (→ that seam) plus a
thin `runtime.fs`, retire the `posix` shim, and follow a phased path to
port to **one seam** (`std.os.danos`), so we build an `os` seam module (→ that seam) plus
the thin `file-system` module, retire the `posix` shim, and follow a phased path to
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
- **[threading.md](threading.md) — threads, the std-shaped way.** **Built** (M1–M6):
`runtime.Thread` mirrors `std.Thread`'s API (spawn/join/detach, Mutex/Condition/
the `thread` module's `Thread` mirrors `std.Thread`'s API (spawn/join/detach, Mutex/Condition/
Semaphore) over a **private** thread ABI — several tasks sharing one address space via
a `thread_spawn` syscall, futex-backed blocking, address-space refcounting. Why it's the
native type and not literal `std.Thread` (the [private ABI](syscall.md)), and why
@@ -220,7 +220,7 @@ addressed as **`system/services/init`** — the repeated leaf resolves away:
|----------------------------------------|--------------------------------------------|
| `system/services/init/init.zig` | `system/services/init` → `/system/services/init` |
| `system/drivers/ps2-bus/ps2-bus.zig` | `system/drivers/ps2-bus` → `/system/drivers/ps2-bus` |
| `library/runtime/runtime.zig` | `library/runtime` (the `runtime` module) |
| `library/device/pci/pci.zig` | `library/device/pci` (the `pci` module) |
In **source**, a sub-project is a directory so it can hold many files — the entry is
`fat/fat.zig`, beside it `fat/engine.zig`, `fat/on-disk.zig`, and so on. When
@@ -234,7 +234,7 @@ A sub-project's extra files are reached through the module, never as separate pa
```
system/ → /system danos's own internals (the self-representation)
boot-handoff.zig the loader↔kernel contract (the `boot-handoff` module)
abi.zig the private kernel↔runtime syscall ABI (the `abi` module)
abi.zig the private kernel↔userspace syscall ABI (the `abi` module)
parameters.zig initial-ramdisk.zig shared contracts
kernel/ IPC, memory, scheduling, the VFS root, the private syscall dispatch
architecture/x86_64/ the `architecture` module (never named by generic code)
@@ -245,10 +245,16 @@ system/ → /system danos's own internals (the self-representation)
services/ init/ fat/ device-manager/ system servers → /system/services (fat/ holds
fat.zig, engine.zig, on-disk.zig)
library/ → /lib libraries, one sub-directory each
runtime/ the danos-native runtime + file API (fs) — the stable application ABI
device/ device code by domain — mmio/ model/ pci/ usb/ acpi/ — each a
shareable data module (device-abi, pci-class, usb-abi/ids,
acpi-ids) plus a logic module (mmio, pci, usb, aml)
kernel/ the danos-native system library (kernel32-style): the syscall
surface split by concern — ipc, memory (heap/dma/shared-memory),
process, time, logging, file-system, thread, service, plus the
system-call stubs and the start/root entry shim
device/ device code by domain — mmio/ model/ pci/ usb/ acpi/ driver/
block/ — each a shareable data module (device-abi, pci-class,
usb-abi/ids, acpi-ids) plus a logic module (mmio, pci, usb, aml,
driver — the device-access + device-manager-hello client)
client/ userspace service clients (display, input) — a program's view of
a service, layered over that service's protocol
protocol/ driver↔service wire contracts (vfs block display scanout input
power device-manager usb-transfer), one module per directory
boot/ → /boot the loaders
@@ -260,11 +266,11 @@ tools/ test/ host-side build + QEMU test harness
name. A protocol is the seam between a low-level driver and the higher-level service it
serves — block ↔ the filesystem, a scanout driver ↔ the compositor — so both sides depend
on the contract, not on each other, and the contract belongs to neither sub-project. A
`runtime` client may *wrap* one for application convenience (`runtime.fs` over
`vfs-protocol`, `runtime.block`, `runtime.display`, `runtime.input`), but the module is the
boundary and `runtime` re-exports no protocol. A driver's private wire to its *hardware*
(virtio-gpu's command set) is not a service seam and stays a driver-private file, beside
the transport that reaches the same device.
client module may *wrap* one for application convenience (the `file-system` module over
`vfs-protocol`, and the `block`, `display`, `input` clients over theirs), but the protocol
module is the boundary — a client re-exports no protocol, it imports it by name. A driver's
private wire to its *hardware* (virtio-gpu's command set) is not a service seam and stays a
driver-private file, beside the transport that reaches the same device.
**Device code lives in `library/device/<domain>/`**, grouped by what it is about (pci, usb,
acpi, and the cross-cutting device model) and split by dependency weight: a data module of
@@ -275,7 +281,7 @@ marshals across the syscall boundary), and nothing with logic or a taxonomy in i
lone pure-data import is the only edge from `system/kernel/` into `library/`.
There is **no POSIX/C compatibility layer today**: danos programs do file I/O through the
danos-native `runtime.fs` (open/read/write/list over the VFS). A hand-rolled POSIX shim
danos-native `file-system` module (open/read/write/list over the VFS). A hand-rolled POSIX shim
(`library/posix/`) was retired as premature — the real POSIX/C surface will come later
from the `std.os.danos` seam (and, eventually, musl) when danos becomes a Zig target (see
[zig-self-hosting.md](zig-self-hosting.md)). When it does, the foreign-ABI naming
@@ -288,7 +294,7 @@ exception in [coding-standards.md](coding-standards.md) applies to that seam.
| Boot methods (one per way of booting the kernel) | `boot/` — `efi.zig` (UEFI) → `BOOTX64.efi` |
| Kernel entry, panic, bring-up | `system/kernel/kernel.zig` |
| Loader↔kernel handoff (`BootInformation`, `Framebuffer`, `MemoryMap`, VM layout) | `system/boot-handoff.zig` |
| Private kernel↔runtime syscall ABI (`SystemCall`, mmap prot flags, `page_size`) — the runtime speaks it, not apps | `system/abi.zig` |
| Private kernel↔userspace syscall ABI (`SystemCall`, mmap prot flags, `page_size`) — the system library speaks it, not apps | `system/abi.zig` |
| Device wire types (`DeviceDescriptor`, `DeviceClass`, …) | `library/device/model/device-abi.zig` |
| Physical frame allocator | `system/kernel/pmm.zig` |
| Kernel heap (`std.mem.Allocator`) | `system/kernel/heap.zig` |
@@ -304,7 +310,8 @@ exception in [coding-standards.md](coding-standards.md) applies to that seam.
| Framebuffer text console (mirrors to serial) | `system/kernel/console.zig` |
| In-kernel test cases | `system/kernel/tests.zig` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/IO-APIC/timer, serial, linker script) | `system/kernel/architecture/x86_64/` |
| danos-native runtime (`runtime`): syscall wrappers, heap, IPC, device access, the file API (`fs`) — the stable application ABI | `library/runtime/` |
| danos-native system library (kernel32-style): the syscall surface by concern — `ipc`, `memory`, `process`, `time`, `logging`, `file-system`, `thread`, `service` — the stable application ABI | `library/kernel/` |
| Service clients (a program's view of a service) and device clients | `library/client/` (display, input), `library/device/driver` |
| System services (init, the `fat` filesystem, the device-manager) | `system/services/` |
| Device drivers, one sub-project each (`pci-bus`, `ps2-bus`, `usb-xhci-bus` bus drivers) | `system/drivers/` |
| Build + `run-x86-64` (QEMU/OVMF) + `release-x86-64` (the flashable ISO) | `build.zig` |
+3 -3
View File
@@ -67,7 +67,7 @@ Three, and only three.
**This exception is scoped to a file that *is* a foreign ABI, and nothing else.**
danos has no such file today: the old `library/posix/` compatibility shim was retired
once its callers moved to the danos-native `runtime.fs`, since a hand-rolled POSIX
once its callers moved to the danos-native `file_system`, since a hand-rolled POSIX
layer is premature until danos actually needs it (see
[zig-self-hosting.md](zig-self-hosting.md)). The exception will apply again to the
`std.os.danos` seam when danos becomes a real Zig target — that module *is* the C-ABI
@@ -141,8 +141,8 @@ single word or acronym needs no hyphen: `scheduler.zig`, `paging.zig`, `apic.zig
conventions above — `snake_case` — because it's an identifier, not a filename.)
**A sub-project's entry point repeats its directory's name** — `init/init.zig`,
`runtime/runtime.zig`, `ps2-bus/ps2-bus.zig` — and the sub-project is addressed by the
*directory* (`system/services/init`, `library/runtime`), with the repeated leaf
`pci/pci.zig`, `ps2-bus/ps2-bus.zig` — and the sub-project is addressed by the
*directory* (`system/services/init`, `library/device/pci`), with the repeated leaf
resolving away. See the repository-layout section of [README.md](README.md).
## Named values, not magic numbers
+2 -2
View File
@@ -22,7 +22,7 @@ for drivers.
How processes stop, reload, and report their deaths is deliberately **not** in this
document: that is the universal lifecycle every danos process speaks —
[process-lifecycle.md](process-lifecycle.md), signals over IPC and the stable
`runtime.process` interface. The device manager is that design's first serious
`process` interface. The device manager is that design's first serious
customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a
driver is stopped, health-checked, and buried exactly like any other process.
@@ -137,7 +137,7 @@ way.
Increments 1–4 are the lifecycle prerequisites and live in
[process-lifecycle.md](process-lifecycle.md) (claim cleanup on death, exit reasons,
published exit events, signals + `runtime.process`). On top of those:
published exit events, signals + `process`). On top of those:
5. **device-manager-protocol**: `hello`, supervised spawn with restart policy;
usb-xhci-bus becomes the first conforming driver.
+7 -7
View File
@@ -96,7 +96,7 @@ rest of the system hasn't had to face:
▼ reached by name (ipc_lookup); clients drive it over the display protocol
┌────────────────────────────────────┬──────────────────────────────────────┐
drawing clients (v1) surface clients (deferred)
runtime.display commands: runtime.display surfaces:
display commands: display surfaces:
create_layer / configure_layer shared_memory_create → pass as a capability →
fill_rect / blit_tile / damage the compositor maps & composites the
present client-rendered bitmap directly
@@ -106,7 +106,7 @@ The bring-up sequence mirrors a hardware driver's — it is the
[`usb-xhci-bus` `initialise`](../system/drivers/usb-xhci-bus/usb-xhci-bus.zig) shape
(claim → `mmio_map` → run loop) — and the request/reply service shell is the
[FAT](../system/services/fat/fat.zig) / [input](../system/services/input/input.zig) shape
([`runtime.service.run`](../library/runtime/service.zig) with a `protocol.zig` of
([`service.run`](../library/kernel/service.zig) with a `protocol.zig` of
`extern struct` messages and an `Operation` tag).
**One process, for now.** v1 is a *single* service that both owns the framebuffer and
@@ -222,13 +222,13 @@ pacing on backends that have none (all of them today; see
[display-v2.md](display-v2.md), "Fenced is not vsync"). Bring-up paths that must put
pixels on screen synchronously (initialisation, the self-checks) bypass the clock.
## `runtime.display`
## `display`
Clients speak the protocol through a new [`library/runtime/display.zig`](../library/runtime/runtime.zig),
the [`runtime.block`](../library/runtime/block.zig) shape (a cached `.display` lookup
Clients speak the protocol through a new [`library/client/display/display.zig`](../library/client/display/display.zig),
the [`block`](../library/device/block/block.zig) shape (a cached `.display` lookup
with a boot-race retry): `display.info()`, a `Layer` handle with `fill` / `blitTile` /
`damage`, and `present()`. Application code never issues the raw syscalls — it calls the
runtime, as with every other danos service.
client module, as with every other danos service.
## The cursor: a mouse-listener thread feeding the compositor
@@ -244,7 +244,7 @@ thread** beside the compositor loop.
the compositor — so no lock guards the framebuffer. A parked `next()` leaves its core
free to halt ([halting.md](halting.md)).
- **The channel.** A single-slot *latest-value* cell (`CursorChannel`) guarded by a
`runtime.Thread.Mutex`: the renderer wants where the cursor *is now*, not a replay of
`Thread.Mutex`: the renderer wants where the cursor *is now*, not a replay of
every delta, so a new position overwrites the old. The listener also **pokes** the
compositor awake — the main loop is parked in `replyWait`, so the listener posts a
zero-payload `ipc.send` to the compositor's endpoint, which arrives as a
+16 -9
View File
@@ -105,7 +105,9 @@ module outlived it, which is rather the point.) The pattern generalises directly
```
library/
runtime/ module "runtime" — syscalls, ipc, lifecycle, memory, threads, log, fs
kernel/ the system library (kernel32-style): the syscall surface split by concern
— ipc, memory (heap/dma/shared-memory), process, time, logging,
file-system, thread, service, plus system-call stubs + start/root
device/ device code grouped by domain; each domain splits into a shareable
data module (enums/wire types, std-only) and a logic module (mmio/IPC)
mmio/ module "mmio" — typed volatile register access + barriers [M14]
@@ -113,13 +115,16 @@ library/
pci/ "pci-class" (data) + "pci" — config/BAR/capability walk (Function)
usb/ "usb-abi" + "usb-ids" (data) + "usb" — descriptors, control/interrupt/bulk client
acpi/ "acpi-ids" (data) + "aml" — _HID names, the AML interpreter
driver/ module "driver" — device-access syscalls + device-manager hello
block/ module "block" — the block-device client (a device type)
client/ userspace service clients — display, input (a program's view of a service)
protocol/ driver <-> service wire contracts, one module per directory
vfs/ block/ display/ scanout/ input/ power/ device-manager/ usb-transfer/
system/drivers/ one sub-project each → /system/drivers (no `d` suffix)
usb-xhci-bus/ HCD + bus driver imports runtime, usb, mmio, usb-transfer-protocol
usb-hid/ class driver imports runtime, usb, input-protocol
virtio-gpu/ scanout driver imports runtime, pci, mmio, display-/scanout-protocol
usb-xhci-bus/ HCD + bus driver imports usb, mmio, usb-transfer-protocol (+ kernel modules)
usb-hid/ class driver imports usb, input-protocol (+ kernel modules)
virtio-gpu/ scanout driver imports pci, mmio, display-/scanout-protocol (+ kernel modules)
```
The split by *dependency weight* is what lets the microkernel stay out of device
@@ -134,7 +139,9 @@ private wire to its *hardware* — virtio-gpu's command set — is not that; it
driver-private file, like the virtio-pci transport beside it.
The build side of this has since landed: [`addUserBinary`](build.zig) injects the
default modules (`runtime`, `mmio`, `xkeyboard-config`, `acpi-ids`) into every user
default modules — the library/kernel concern modules (`ipc`, `memory`, `process`, `time`,
`logging`, `file-system`, `thread`, `service`), the device/service clients (`driver`,
`block`, `display`, `input`), plus `mmio`, `xkeyboard-config`, `acpi-ids` — into every user
binary, and per-binary extras — protocol modules, bus logic — are added with
`programModule(exe).addImport(...)`. That's the *entire* mechanism — Zig modules
already give you everything else.
@@ -156,10 +163,10 @@ class driver, the device manager, or the kernel may share them freely.
and a `received_cap` return (r8): an endpoint travels with a message, installed into
the receiver's handle table (shared, refcount-bumped — a copy, not a move). A full
table fails `-ENOSPC` and does not half-deliver. This is the "open" primitive — a bus
driver mints a per-device endpoint and hands it to a class driver. The runtime exposes
`callCap` and `replyWait(..., send_cap)`, and class drivers consume them now: the
PS/2 keyboard and mouse drivers attach to ps2-bus this way, and `runtime.usb` /
`runtime.input` open their per-device and subscription channels with `callCap`.
driver mints a per-device endpoint and hands it to a class driver. The `ipc` module
exposes `callCap` and `replyWait(..., send_cap)`, and class drivers consume them now: the
PS/2 keyboard and mouse drivers attach to ps2-bus this way, and the `usb` / `input`
client modules open their per-device and subscription channels with `callCap`.
- **M14** — DMA memory + the memory-ordering layer. `/lib/device/mmio` gives drivers typed
volatile access and `memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier` (per-arch); `dma_alloc`/`dma_free` grant
physically-contiguous, pinned, uncacheable, reclaim-on-teardown buffers with the
+1 -1
View File
@@ -391,7 +391,7 @@ the first DMA driver to protect and test against) and these smaller items:
Claiming and mapping is half of being a danos driver; the other half is the
**lifecycle and protocol contract**, and the runtime makes it nearly free:
- Build on `runtime.service.run` — one replyWait loop folding protocol
- Build on `service.run` — one replyWait loop folding protocol
requests, signals, and notifications into callbacks. The harness answers the
universal zero-length ping and turns `terminate` into a clean exit for you
([process-lifecycle.md](process-lifecycle.md)).
+1 -1
View File
@@ -89,7 +89,7 @@ This is the async counterpart of `ipc_call`, and the input service is its first
- A **subscriber** calls `input.subscribe(mask)` — or a typed helper: `subscribeKeyboard()`,
`subscribeMouse()`, `subscribeJoystick()` (one class, `next()` returns the decoded event),
or `subscribeAll()` (every class, `next()` returns a tagged `InputEvent`)
([library/runtime/input.zig](../library/runtime/input.zig)). It creates its own endpoint
([library/client/input/input.zig](../library/client/input/input.zig)). It creates its own endpoint
and hands it to the service as a **capability** (M13 capability passing — the input
service is that feature's first real user), along with its `device_mask`. Then it loops on
`next()`, a `replyWait` on that endpoint returning each pushed event.
+4 -4
View File
@@ -118,15 +118,15 @@ Three conventions from [process-lifecycle.md](process-lifecycle.md) ride the
notification mechanism:
- **Signals** arrive as notifications on the endpoint a process nominated with
`signal_bind` (`runtime.process.bindSignals`): badge = the signal bit plus the
coalesced pending mask (`runtime.process.signalsFrom` decodes). Statements,
`signal_bind` (`process.bindSignals`): badge = the signal bit plus the
coalesced pending mask (`process.signalsFrom` decodes). Statements,
never questions; no payload, no reply.
- **One-shot timers** (`timer_bind`, `runtime.system.timerOnce`) land as a
- **One-shot timers** (`timer_bind`, `time.timerOnce`) land as a
timer-bit notification — the timed wait: a service arms a deadline and keeps
serving, instead of blocking in sleep.
- **The universal ping**: a **zero-length request is the liveness probe**,
answered with a zero-length reply by the service harness itself
(`runtime.service.run`). No protocol's requests start at length zero, so the
(`service.run`). No protocol's requests start at length zero, so the
encoding cannot collide, and a wedged service simply fails to answer — which
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
protocol message.
+2 -2
View File
@@ -21,9 +21,9 @@ kernel log.print ─┘ │
1. **Emit.** A program calls `std.log.info("mounted {s}", .{path})` — the
runtime's `logFn` (installed for every binary by the root shim,
`library/runtime/log.zig`) formats one line and issues one `debug_write`
`library/kernel/logging.zig`) formats one line and issues one `debug_write`
carrying the level. The payload does NOT contain the process's name.
`runtime.system.write` remains as the raw/bring-up path (panics, test
`logging.write` remains as the raw/bring-up path (panics, test
fixtures); raw bytes ride the same ring, attributed all the same.
2. **Stamp.** The kernel wraps every payload LINE in a record stamped with the
+2 -2
View File
@@ -17,7 +17,7 @@ What subscribers want is not: *the lid closed* means the same thing regardless o
who noticed. So the surface is **domain-named**. There is a `power-protocol`
module and a well-known `ServiceId.power = 5`; on x86 the **acpi service**
registers it, and on ARM a PSCI/mailbox service will register the *same* id.
Subscribers call `runtime.ipc.lookup(.power)` and never learn which firmware they
Subscribers call `ipc.lookup(.power)` and never learn which firmware they
are on — the neutrality the whole [discovery](discovery.md) migration exists to
preserve, carried one layer up into a running-system surface.
@@ -75,7 +75,7 @@ notifications, the lifecycle **signals** it can receive (`terminate`), and the
On a `power_button` event or a `terminate` signal, init:
1. logs that it is shutting down,
2. runs the standard stop sequence — `runtime.process.stop(child, deadline,
2. runs the standard stop sequence — `process.stop(child, deadline,
endpoint)` — over its children **in reverse spawn order**, so the VFS stops
last (other services may flush through it), each child getting the
*terminate → deadline → kill* escalation from
+7 -7
View File
@@ -6,7 +6,7 @@ harness are all in — the interface below is as-built. The primitives underneat
predate this design ([process-management.md](process-management.md):
spawn, the supervision link, kill, child-exit notifications); this document designs
the layer above them — the standard vocabulary a danos process speaks about its own
life, and the stable `runtime.process` interface that carries it. Nothing here is
life, and the stable `process` interface that carries it. Nothing here is
device- or driver-specific: a driver, the VFS, and a user application all stop,
reload, and die the same way. The device manager is simply this design's first
serious customer ([device-manager.md](device-manager.md)).
@@ -184,7 +184,7 @@ zombie state or privileged snooping:
same encoding supervisors already decode, the IRQ-as-IPC pattern once more). The
subscriber filters for ids it holds state for and releases what the dead client
held. Correlating is free of bookkeeping: an IPC sender's badge already *is* its
task id (`runtime.ipc.Received`), so the id a service has been keying client
task id (`ipc.Received`), so the id a service has been keying client
state by all along is the id the exit event carries.
Subscription, not broadcast-to-everyone: only processes that asked receive
@@ -200,9 +200,9 @@ its clients cleaning up after themselves.** Handle release on client death is th
service's job, triggered by the published exit event — never by a courtesy
"closing now" message that a crashed client will never send.
## The stable interface: `runtime.process`
## The stable interface: `process`
`runtime.process` already owns what a process receives at birth (`Init`, the
`process` already owns what a process receives at birth (`Init`, the
argv contract). It grows to own the other end of life.
**The runtime is the stable interface; the numbers are not.** danos applications do
@@ -283,7 +283,7 @@ callbacks (`on_terminate`, `on_reload`) for programs that want defaults.
### The service harness
`runtime.service` owns the `replyWait` loop and folds every event source — signals,
`service` owns the `replyWait` loop and folds every event source — signals,
child exits, protocol messages — into callbacks, with the vocabulary's defaults:
`terminate` returns from the loop (clean exit), the common `ping` is answered automatically,
`reload` is ignored unless overridden. One loop, no locking, nothing reentrant. A
@@ -312,12 +312,12 @@ get POSIX; danos-native programs never pay for it.
kill a claiming driver, spawn it again, the claim succeeds.
2. **Exit reason in the death notification** (`ExitReason` above).
3. **Exit events**: `process_subscribe` in the kernel (bounded subscriber table,
publishes on every death), `runtime.process.subscribeExits`; the userspace VFS
publishes on every death), `process.subscribeExits`; the userspace VFS
router was the first subscriber — releasing a dead client's handles was its
proof test — and the FAT server inherited the role when the router moved into
the kernel (clients now hold the filesystem server's node ids directly).
4. **Signals**: `signal_bind` + `process_signal` + the pending mask in the kernel;
`runtime.process` grows the interface above; the service harness handles
`process` grows the interface above; the service harness handles
`terminate` and answers the common `ping`; `stop()` for supervisors.
[device-manager.md](device-manager.md) builds directly on all four.
+1 -1
View File
@@ -114,7 +114,7 @@ the architecture layer calls up into `tick`.
- ~~There is no exit status in the notification~~ Closed (M17.2): the kernel
records how every process ends — exited, a fault class, or killed — before it
posts the exit notification, and the supervisor reads it with
`process_exit_reason` (`runtime.process.exitReason`). This is the input to
`process_exit_reason` (`process.exitReason`). This is the input to
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
for the clean case can still ride alongside later.
- Enumerate writes through the caller's raw pointer under the bring-up trust
+3 -3
View File
@@ -90,9 +90,9 @@ process) instead of silently corrupting the image
(`buildEntryStack` in `system/kernel/process.zig`); `argv[0]` is always the path
or initial-ramdisk name the process was spawned as, and `system_spawn`'s optional
argument blob becomes `argv[1..]`. The runtime's `_start`
(`library/runtime/start.zig`) hands the block to `rt_start`, which builds a
`runtime.process.Init` from it and passes that to the program's `main`
(`pub fn main(init: runtime.process.Init)`; a parameterless `main()` is also
(`library/kernel/start.zig`) hands the block to `rt_start`, which builds a
`process.Init` from it and passes that to the program's `main`
(`pub fn main(init: process.Init)`; a parameterless `main()` is also
accepted). A C runtime's `crt0` would walk
the identical layout unmodified — that's the compatibility being bought. The
`args` test proves the round trip.
+17 -17
View File
@@ -1,8 +1,8 @@
# Threading: `runtime.Thread`, a std-shaped API over a private thread ABI
# Threading: `Thread`, a std-shaped API over a private thread ABI
A note on danos **threads** — several tasks sharing one address space — provided by a
`runtime.Thread` type that mirrors the shape of Zig's `std.Thread` while keeping every
kernel entry behind the [runtime](../library/runtime). **Built** (M1–M11, see
`Thread` type that mirrors the shape of Zig's `std.Thread` while keeping every
kernel entry behind the [runtime](../library/kernel). **Built** (M1–M11, see
[threading-plan.md](threading-plan.md)): `spawn`/`join`/`detach`, cross-core parallelism,
a futex, `Mutex`/`Condition`/`Semaphore`/`RwLock`/`WaitGroup`, `getCurrentId`/`currentCore`,
per-thread thread-pointer TLS, thread-safe allocation, and a task reaper that reclaims dead
@@ -17,13 +17,13 @@ treat upstream shapes as "0.16.x."
A danos service can write
```zig
const t = try runtime.Thread.spawn(.{}, worker, .{ctx});
const t = try Thread.spawn(.{}, worker, .{ctx});
// ... do other work concurrently ...
t.join();
```
and get real parallelism across cores — with `runtime.Thread.Mutex`,
`runtime.Thread.Condition`, and `runtime.Thread.Semaphore` available for
and get real parallelism across cores — with `Thread.Mutex`,
`Thread.Condition`, and `Thread.Semaphore` available for
coordination — **without any code path reaching the kernel except through the
runtime**. The call sites read exactly like `std.Thread`, so the day danos becomes a
real Zig target (see [self-hosting](#the-self-hosting-endgame)) we swap the
@@ -31,7 +31,7 @@ implementation underneath, not the API above.
## Locked decisions (do not relitigate)
- **We build `runtime.Thread`, not literal `std.Thread`.** It mirrors std's *API and
- **We build `Thread`, not literal `std.Thread`.** It mirrors std's *API and
features*; the implementation underneath is danos-native. See
[Why not literal std.Thread](#why-not-literal-stdthread).
- **Threads are a narrow, opt-in capability — not the default concurrency tool.** The
@@ -42,7 +42,7 @@ implementation underneath, not the API above.
- **Per-binary opt-in to multi-threaded codegen.** Only a service that asks for
threads is built `single_threaded = false`; the rest stay lean and single-threaded.
- **The thread ABI is private.** New syscalls extend [abi.zig](../system/abi.zig)
`SystemCall` and are reached only through `library/runtime` wrappers, exactly like
`SystemCall` and are reached only through `library/kernel` wrappers, exactly like
every other danos syscall ([syscall.md](syscall.md)) — numbers stay renumberable.
## Why not literal `std.Thread`
@@ -99,7 +99,7 @@ processes. The isolation boundary stays at process granularity.
## The API surface (mirrors `std.Thread`)
Lives in `library/runtime/thread.zig`, re-exported as `runtime.Thread`.
Lives in `library/kernel/thread.zig`, re-exported as `Thread`.
```zig
pub const Thread = struct {
@@ -129,14 +129,14 @@ Deviations from `std.Thread`, called out honestly:
`void`). Return data through shared state or a `Semaphore`/`Condition`, not the
return.
- No `getCpuCount()` (a service rarely needs it) and no `Thread.yield()` — `yield`
lives in `runtime.system`. Instead `currentCore()` exposes the calling core's dense
lives in the `process` module. Instead `currentCore()` exposes the calling core's dense
index ([smp.md](smp.md)), used to observe genuine cross-core parallelism.
## Kernel primitives (new private syscalls)
Five core entries extend [abi.zig](../system/abi.zig) `SystemCall` after
`shared_memory_physical = 36` (plus small helpers `current_core`, `thread_self`, and
`set_thread_pointer`), each with a `library/runtime` wrapper:
`set_thread_pointer`), each with a `library/kernel` wrapper:
| Syscall | Signature | Purpose |
|---|---|---|
@@ -239,7 +239,7 @@ see the intro). Two scoped pieces, as built:
A binary opts in by being added with `addThreadedUserBinary` — as `addUserBinary`,
but the shared implementation builds it `single_threaded = false` — so atomics and
(later) TLS are real. Threads and atomics are unsound in a `single_threaded` image,
so a binary must opt in **before** it may call `runtime.Thread.spawn`. Everyone else
so a binary must opt in **before** it may call `Thread.spawn`. Everyone else
stays single-threaded and lean.
## Interaction with the rest of the kernel
@@ -313,9 +313,9 @@ a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
Follow [coding-standards.md](coding-standards.md): spell out non-acronym
abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New syscalls
extend [abi.zig](../system/abi.zig) `SystemCall` + a `library/runtime` wrapper
([syscall.md](syscall.md)). `runtime.Thread` is a first-class runtime module, the same
way `runtime.process` ([process-lifecycle.md](process-lifecycle.md)) and `runtime.ipc`
extend [abi.zig](../system/abi.zig) `SystemCall` + a `library/kernel` wrapper
([syscall.md](syscall.md)). `Thread` is a first-class runtime module, the same
way `process` ([process-lifecycle.md](process-lifecycle.md)) and `ipc`
are — user code never names a syscall.
## Non-goals
@@ -333,8 +333,8 @@ are — user code never names a syscall.
When danos becomes a real Zig target and we (eventually) add a danos backend to std
([zig-self-hosting.md](zig-self-hosting.md)), `std.Thread` can sit *on top of* these
same kernel primitives — the danos `std.Thread.Impl` would call the very
`thread_spawn`/`futex_*` wrappers `runtime.Thread` already uses. Because
`runtime.Thread` was built API-compatible from day one, that transition swaps the
`thread_spawn`/`futex_*` wrappers `Thread` already uses. Because
`Thread` was built API-compatible from day one, that transition swaps the
implementation, not a single call site. Designing to the std shape now is what makes
the later self-hosting lift cheap.
+8 -7
View File
@@ -61,10 +61,10 @@ The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to use
programs read the TSC through `clock` and get timed wakeups through `sleep`/`timer_bind`,
both riding the scheduler tick.
## `runtime.time` — the generic interface
## `time` — the generic interface
Applications don't call the syscalls directly; they use `runtime.time`
(`library/runtime/time.zig`), a thin `Instant`/`Duration` layer over them — an ergonomic
Applications don't call the syscalls directly; they use `time`
(`library/kernel/time.zig`), a thin `Instant`/`Duration` layer over them — an ergonomic
front door, not new mechanism.
```zig
@@ -91,8 +91,9 @@ _ = time.after(endpoint, time.Duration.fromMillis(200));
- `sleep(d)` wraps `sleep`; `spin(d)` busy-polls `now()` for the sub-millisecond delays
the millisecond tick can't express; `after(endpoint, d)` wraps `timer_bind`.
The raw wrappers (`system.clock`, `system.sleep`, `system.timerOnce`) stay in
`library/runtime/system.zig`; `runtime.time` is the layer meant for everyday use.
The raw wrappers (`clock`, `sleepMillis`, `timerOnce`) and the ergonomic
`Instant`/`Duration` layer both live in the `time` module
(`library/kernel/time.zig`); the latter is what everyday code uses.
## Wall-clock time (not built)
@@ -106,10 +107,10 @@ owns covers every current use.
## Verifying it
`runtime.time`'s `Instant`/`Duration` arithmetic has unit tests that run on the host:
`time`'s `Instant`/`Duration` arithmetic has unit tests that run on the host:
```
$ zig build test # includes library/runtime/time.zig
$ zig build test # includes library/kernel/time.zig
```
End to end, the proof the clock is real is that it *advances*: read `now()`, `sleep` a
+9 -9
View File
@@ -1,7 +1,7 @@
# The vDSO — the public system-call boundary
> **Status:** design note, not built. The runtime today issues raw `syscall`
> instructions from `library/runtime/system-call.zig` using the numbers in
> instructions from `library/kernel/system-call.zig` using the numbers in
> `system/abi.zig`. This note designs the layer that replaces that arrangement:
> a **kernel-supplied, C-ABI entry library** mapped into every process — the
> only supported way into the kernel — so the raw numbers can stay private,
@@ -25,8 +25,8 @@ ourselves:
this mistake: it issued XNU syscalls directly instead of going through
libSystem, and macOS updates repeatedly broke every Go binary until Go
switched to the library like everyone else.
2. **Not everything is Zig.** A Rust or C program can't import the `runtime`
module. The public boundary has to be expressible in the one calling
2. **Not everything is Zig.** A Rust or C program can't import the danos Zig
modules. The public boundary has to be expressible in the one calling
convention every language speaks: the C ABI.
3. **Randomised syscall numbers** — a hardening option we want open — only
work if no user binary anywhere knows a number at build time. The binding
@@ -49,10 +49,10 @@ The public danos ABI then has exactly two layers, neither of which is
| Layer | Contract | Spoken by |
|-------|----------|-----------|
| **vDSO** | C-ABI functions, this note | every language's thin shim (`runtime.system` for Zig, a `-sys` crate for Rust, a header for C) |
| **vDSO** | C-ABI functions, this note | every language's thin shim (the `system-call` module for Zig, a `-sys` crate for Rust, a header for C) |
| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](vfs-protocol.md) is the first one documented) | any client that can lay out bytes |
Everything above those — the heap, `runtime.fs`, the service harness — is
Everything above those — the heap, `file_system`, the service harness — is
per-language convenience, compiled into each binary from source, exactly as
today. Nothing about the Zig runtime's shape changes; it just stops being the
*only* door.
@@ -107,7 +107,7 @@ convenience, not a requirement.)
The kernel already builds a System V entry block — argc, argv, envp
terminator, **auxiliary vector** — on every new process's stack
(`buildEntryStack`, read by `runtime.start`). The vDSO base rides in a new
(`buildEntryStack`, read by the `start` module). The vDSO base rides in a new
auxv entry, exactly Linux's `AT_SYSINFO_EHDR` move. No new syscall, no magic
address, and a language shim finds it the same portable way on every
architecture.
@@ -183,11 +183,11 @@ second — but the design should never be sold as more than that.
Phased so every step ships alone (the M-milestone discipline):
1. **The blob + the table.** Build the vDSO, map it at spawn, deliver the
base via auxv. `runtime.system-call.zig` binds through the table when the
base via auxv. `library/kernel/system-call.zig` binds through the table when the
auxv entry is present, falls back to raw `syscall` when absent — the whole
tree keeps booting during the transition.
2. **Cut the runtime over.** Delete the raw stubs; `runtime` no longer
imports the `SystemCall` numbers at all (`abi.zig`'s enum becomes
2. **Cut the system library over.** Delete the raw stubs; the `system-call`
module no longer imports the `SystemCall` numbers at all (`abi.zig`'s enum becomes
kernel-internal). The QEMU suite passing proves the table carries the
whole system.
3. **Enforce + randomise.** Add the `rip`-range check, then per-boot number
+2 -2
View File
@@ -1,6 +1,6 @@
# The VFS wire protocol
> **Status:** built and spoken today between `runtime.fs` (the client) and the
> **Status:** built and spoken today between `file_system` (the client) and the
> filesystem BACKENDS (the FAT server). The mount router lives in the
> **kernel** (`system/kernel/vfs.zig`): `fs_resolve` routes a path and either
> serves it directly (the read-only /system initrd mount, via `fs_node`) or
@@ -113,7 +113,7 @@ Notes per operation:
prefix maps a mount into the backend's namespace (fat serves `/mnt/usb`
from its volume root and `/var` from its `/var` subtree).
- **rename** — same-directory rename only: the backend compares the old and
new parent paths and refuses a mismatch. The client (`runtime.fs`) refuses
new parent paths and refuses a mismatch. The client (`file_system`) refuses
earlier when the two paths resolve to different backend endpoints, but that
check is coarser than "one mount" — one endpoint can serve several mounts
(fat serves `/mnt/usb` and `/var`), so a cross-mount rename reaches the