An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115.
15 KiB
The vDSO — the public system-call boundary
Status: design note, not built. The runtime today issues raw
syscallinstructions fromlibrary/kernel/system-call.zigusing the numbers insystem/abi.zig. This note designs the layer that replaces that arrangement: a kernel-supplied, C-ABI entry library mapped into every process — the only supported way into the kernel — so the raw numbers can stay private, be renumbered at will, and eventually be randomised per boot.
Why: the ABI danos promises, and the one it doesn't
system/abi.zig is the private kernel ↔ runtime contract. Its header says
so: the numbers are an implementation detail the runtime hides and may
renumber, the same split as libSystem over the XNU syscalls on macOS or win32
over the NT syscalls on Windows. Linux — with its world-visible, frozen
syscall table — is the outlier, not the norm.
That stance has consequences the moment binaries exist that we don't rebuild ourselves:
- Third-party binaries (docs/zig-self-hosting.md) must keep working across
kernel updates. If they contain raw
syscallinstructions with today's numbers baked in, every renumbering breaks the world — the ABI would be de facto public no matter what the header says. Go on macOS made exactly this mistake: it issued XNU syscalls directly instead of going through libSystem, and macOS updates repeatedly broke every Go binary until Go switched to the library like everyone else. - Not everything is Zig. A Rust or C program can't import the danos Zig modules. The public boundary has to be expressible in the one calling convention every language speaks: the C ABI.
- Randomised syscall numbers — a hardening option we want open — only work if no user binary anywhere knows a number at build time. The binding must happen at load time, from something the kernel controls.
All three point at the same well-known shape: a vDSO (virtual dynamic
shared object). The kernel carries a small blob of user-mode code, maps it
into every process at spawn, and that blob — not the application — contains
the syscall instructions. Fuchsia works exactly this way: its vDSO is the
only kernel entry, version-matched by construction because the kernel itself
injects it. Because the kernel and the blob ship as one artifact, there is
no version skew, no loader, no search path, and no shared file on disk —
which is what makes this the resilient way to have a private ABI
(docs/resilience.md), where a conventional ld.so + /lib/libdanos.so
arrangement would add a loader to every spawn and a single shared point of
failure.
The public danos ABI then has exactly two layers, neither of which is
abi.zig:
| Layer | Contract | Spoken by |
|---|---|---|
| vDSO | C-ABI functions, this note | every language's thin shim (the system-call module for Zig, a -sys crate for Rust, a header for C) |
| IPC wire protocols | byte layouts over ipc_call (vfs-protocol.md is the first one documented) |
any client that can lay out bytes |
Everything above those — the heap, file_system, the service harness — is
per-language convenience, compiled into each binary from source, exactly as
today. Nothing about the Zig runtime's shape changes; it just stops being the
only door.
The blob
A single copy of the vDSO code lives in the kernel image (built by
build.zig as a tiny freestanding object, embedded like the AP trampoline).
At boot the kernel finalises it once — this is where randomised numbers would
be patched in — and thereafter maps the same physical pages read-execute
into every process's address space. The blob is:
- Position-independent. It is mapped at a per-process randomised base, so it must be PIC (rip-relative addressing only — no relocations to process).
- Stateless and re-entrant. No writable data. Anything stateful belongs to the process, not the vDSO.
- Architecture-specific. The x86-64 blob wraps
syscall; an aarch64 blob wrapssvc #0. It lives beside the other per-architecture kernel sources (system/kernel/architecture/<arch>/), selected the same way thearchitecturemodule is (docs/architecture.md).
Shape: a function table, not an ELF
A real .so with a dynamic symbol table is the conventional vDSO shape, but
linking against one at load time needs a dynamic linker in every binary —
machinery danos deliberately doesn't have. Instead the v1 shape is the
simplest thing that is still a stable contract — a function-pointer table
at the vDSO base:
offset 0 u64 magic 'danosVDS' — a mapped-the-wrong-thing guard
offset 8 u64 api_level incremented when the table grows
offset 16 u64 count number of table entries that follow
offset 24 u64 table[count] function pointers into the vDSO's own code
Table indices are the public constants (published in a C header,
danos.h), assigned once and append-only — the same discipline the IPC
protocols use for operation values. The pointers point at stubs inside the
blob; what those stubs put in rax is nobody's business but the kernel's.
A language shim binds in one step: read the base from the init block, check
the magic, keep the table pointer. Feature detection for a binary built
against older headers is count/api_level — a kernel never removes or
reorders entries.
(If danos ever grows a real dynamic linker, the same blob can additionally
present an ELF dynsym without breaking the table — Fuchsia's vDSO is
likewise both a mappable blob and a linkable .so. That is a later
convenience, not a requirement.)
Delivery: the auxiliary vector
The kernel already builds a System V entry block — argc, argv, envp
terminator, auxiliary vector — on every new process's stack
(buildEntryStack, read by the start module). The vDSO base rides in a new
auxv entry, exactly Linux's AT_SYSINFO_EHDR move. No new syscall, no magic
address, and a language shim finds it the same portable way on every
architecture.
The function surface
One table entry per kernel call, C ABI (System V AMD64), names prefixed
danos_. The current SystemCall set maps directly; integer arguments and
returns are u64, errors return as negative values exactly as today.
The calls that return two values in rax:rdx today — dma_alloc
(virtual_address + physical_address), msi_bind (address + data), shared_memory_create (virtual_address + handle),
fs_resolve (route tag + node token / backend handle) —
become functions returning a two-u64 struct. The System V ABI returns a
16-byte struct in rax:rdx, so the stub is a plain syscall; ret — the
C-ABI spelling of the existing convention, at zero cost. The one call that
returns three values — ipc_reply_wait (receive_len in rax, badge in
rdx, received capability in r8) — exceeds the two-register return: its
function returns a three-u64 struct, which the ABI passes via a hidden
result pointer, so that one stub stores rax/rdx/r8 through the pointer
after the syscall — a few instructions rather than one.
Grouped as abi.zig groups them:
| Group | Functions |
|---|---|
| process | danos_exit, danos_yield, danos_sleep, danos_spawn, danos_process_enumerate, danos_process_kill, danos_process_exit_reason, danos_process_subscribe, danos_process_signal, danos_signal_bind |
| threads | danos_thread_spawn, danos_thread_exit, danos_current_core, danos_futex_wait, danos_futex_wake, danos_thread_self, danos_thread_join, danos_set_thread_pointer |
| memory | danos_mmap, danos_munmap, danos_dma_alloc, danos_dma_free, danos_shared_memory_create, danos_shared_memory_map, danos_shared_memory_physical |
| ipc | danos_endpoint_create, danos_ipc_call, danos_ipc_reply_wait, danos_ipc_send (naming is not a syscall: a provider binds its contract at the registry and a client resolves /protocol/<name> — see protocol-namespace.md) |
| devices | danos_device_enumerate, danos_device_claim, danos_device_register, danos_mmio_map, danos_irq_bind, danos_irq_ack, danos_msi_bind, danos_io_read, danos_io_write |
| time | danos_clock, danos_wall_clock, danos_timer_bind |
| diagnostics | danos_debug_write (leveled, kernel-stamped records), danos_klog_read, danos_klog_status |
| filesystem naming | danos_fs_resolve, danos_fs_node, danos_fs_mount, danos_fs_unmount (naming only — file DATA still crosses the vfs-protocol IPC, see below) |
The constants that ride alongside the calls — mmap protection bits, DMA
flags, notification badge bits, ExitReason, Signal, page_size, the IPC
message maximum — move to the public header too:
they are wire values a Rust program needs verbatim. What stays private in
abi.zig is exactly the thing the vDSO exists to hide: the SystemCall
numbers and the trap convention.
Errors: the errno space
A failed call returns -errno. The runtime detects failure the way Linux
does — a return value in the top 4096 — so every code stays inside 1..4095.
These are public: unlike the call numbers, a caller must be able to read
them verbatim, and they are the same vocabulary whether the number came from
the kernel or from a user-space provider answering over IPC.
They are defined once in system/abi.zig. The kernel restates them in
system/kernel/ipc-synchronous.zig and the envelope restates the
provider-facing subset in library/protocol/envelope/envelope.zig (the
protocol package deliberately depends on nothing, so it cannot import
abi); a comptime check in library/device/driver/driver.zig makes drift a
compile error.
| # | Name | Meaning |
|---|---|---|
| 1 | EBADF |
bad handle |
| 2 | E2BIG |
an argument exceeds its maximum (a message, a descriptor's resource count) |
| 3 | EFAULT |
buffer unmapped, or outside the user half |
| 4 | ENOENT |
no such name |
| 5 | ENOSPC |
a kernel table is full (handles, devices) |
| 6 | ENOMEM |
out of memory |
| 7 | EPEER |
the peer died before replying — its process exited or was killed |
| 8 | ESRCH |
no such process |
| 9 | EPERM |
not permitted: the caller is not the owner or supervisor |
| 10 | ENOSYS |
this protocol has no such operation |
| 11 | EPROTO |
malformed packet: shorter than the verb it names |
| 12 | EBUSY |
the thing asked for is held by someone still alive |
| 13 | ENODEV |
no such device id |
| 14 | ECHILDREN |
this parent already holds as many children as it can |
| 15 | ERANGE |
a resource escapes the window it must fall inside |
| 16 | ECONFINE |
the device could not be placed under IOMMU translation |
EPEER is the one with no POSIX counterpart and it is worth stating plainly:
synchronous IPC blocks the caller until the server replies, so the caller
needs an answer for "the server died while I was waiting." It is not a
transport error and not a refusal — the request may well have been carried
out — it says only that no reply is coming. A client that treats it as
"retry" can duplicate work; the honest response is to re-resolve the protocol
name, because the provider it held is gone.
A refusal names the rule that refused it. This is a rule and not a
courtesy. device_register alone can fail six ways, and until each got its
own code a bus driver could only report "refused" — which is how an AMD
desktop came to boot with a working display, no USB and no storage, with
three independent causes indistinguishable in the log. See
fixed-bounds-audit.md. A new failure mode that
does not fit an existing code gets a new one here rather than borrowing the
nearest.
Enforcement, and an honest threat model
Renumbering only has teeth if the kernel refuses syscalls that don't come
from the vDSO. The check is cheap: on kernel entry, the saved user rip
must lie inside the calling process's vDSO mapping; otherwise the process is
killed with a fault-class exit reason (its supervisor restarts or gives up,
docs/process-lifecycle.md — a foreign-syscall attempt is a bug or an attack,
never something to limp past). Fuchsia enforces exactly this.
What this buys, precisely:
- ABI freedom — the real prize. The numbers can change per release or per boot and nothing outside the kernel image cares. The private ABI stays actually private, permanently.
- A single audited chokepoint for kernel entry, per process, at a randomised address.
- Raised bar for exploits: shellcode can't issue a hard-coded
syscall; it must first discover the per-process vDSO base (ASLR) and call through it.
What it does not buy: an attacker with arbitrary code execution in a process can still call the vDSO functions — they are mapped executable in that process, and return-oriented chains reach them. Syscall randomisation is hardening, not a security boundary; the security boundary remains the capability model (what the process's endpoints and device claims let it do). It is worth building anyway — for the ABI freedom first and the hardening second — but the design should never be sold as more than that.
Migration
Phased so every step ships alone (the M-milestone discipline):
- The blob + the table. Build the vDSO, map it at spawn, deliver the
base via auxv.
library/kernel/system-call.zigbinds through the table when the auxv entry is present, falls back to rawsyscallwhen absent — the whole tree keeps booting during the transition. - Cut the system library over. Delete the raw stubs; the
system-callmodule no longer imports theSystemCallnumbers at all (abi.zig's enum becomes kernel-internal). The QEMU suite passing proves the table carries the whole system. - Enforce + randomise. Add the
rip-range check, then per-boot number randomisation patched into the blob at kernel init. A test boots with randomisation on and runs the full suite. - The other languages. Publish
danos.h; a Rustdanos-syscrate wraps the table. This is also the seamstd.os.danoscalls through when the Zig self-hosting fork lands (docs/zig-self-hosting.md) — the vDSO is what makes that seam stable across kernel versions.
What deliberately stays out
- No dynamic linker, no
/lib/*.so. The vDSO is kernel-injected precisely so danos binaries can stay fully static above it. Sharing library code across processes stays what it is today: a service behind IPC, or source compiled into each binary. - No file/device I/O in the vDSO. The microkernel line doesn't move: the
vDSO wraps the same deliberately tiny table (docs/syscall.md). The kernel
resolves file NAMES (
fs_resolve— the mount table moved in-kernel), but file data is still the filesystem server's business over the vfs-protocol IPC; the kernel never blocks on a userspace filesystem. - No fast-path user-mode implementations yet. Linux's vDSO exists mostly
to answer
gettimeofdaywithout a kernel entry.danos_clockcould one day read the calibrated TSC in user mode the same way — the blob is where such an optimisation would live — but that is an optimisation, not part of this design's contract.