diff --git a/docs/README.md b/docs/README.md index a7de12c..6b76a48 100644 --- a/docs/README.md +++ b/docs/README.md @@ -39,37 +39,42 @@ rather than restate it. Roughly in the order things happen at runtime: endpoints — the backbone the microkernel's isolated servers talk over. 12. **[syscall.md](syscall.md) — system calls.** How ring 3 asks the kernel for something: the `syscall`/`sysret` fast path, the trap frame, and why the table is - deliberately tiny. -13. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an + deliberately tiny. The numbers are a **private** ABI — [vdso.md](vdso.md) designs + the public boundary that will hide them. +13. **[vfs-protocol.md](vfs-protocol.md) — the VFS wire protocol.** The language-neutral + byte-level spec of the file protocol spoken over IPC: request/reply headers, + the operation table, mount routing, and the append-only evolution rules — the + first IPC protocol documented as public ABI. +14. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an ordinary ring-3 process that claims a device, maps its registers, and **sleeps until its hardware interrupts it**. The claim is the capability; `irq_ack` is the unmask. -14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How +15. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How real driver stacks factor into three shapes and how families share code. The three primitives it proposed are long since built (M13 capability passing, M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them — hello, supervision, restart — is built too (device-manager.md, M18). -15. **[process-management.md](process-management.md) — process management.** The +16. **[process-management.md](process-management.md) — process management.** The microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the supervision link as the kill authority, and child-exit notifications over the same endpoints IRQs arrive on. -16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built +17. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built (M17): signals over IPC as the one lifecycle vocabulary every process speaks — the POSIX.1-1990 words with message delivery instead of stack hijack, the stable `runtime.process` interface, exit reasons, published exit events any stateful service can subscribe to (the VFS releasing dead clients' handles), and the two iron rules (cleanup is the kernel's job; kill is not a signal). -17. **[device-manager.md](device-manager.md) — the device manager.** Built (M18, +18. **[device-manager.md](device-manager.md) — the device manager.** Built (M18, through the app surface): the tree, the matcher, and the supervisor. Tree structure lives in the manager, authority stays in the kernel; bus drivers report what they see; drivers are restarted through the lifecycle vocabulary — the plan that turns [resilience.md](resilience.md)'s restart goal into increments. -18. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard, +19. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard, mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish service layered on top. -19. **[display.md](display.md) — the display service.** The display half of the GUI +20. **[display.md](display.md) — the display service.** The display half of the GUI track: a user-space compositor that owns the framebuffer, composes a layer stack into a double buffer, and presents it. Why GOP and the PCI display device are two views of one controller, the device-node + write-combining handoff, and what flicker-free buys @@ -80,7 +85,7 @@ rather than restate it. Roughly in the order things happen at runtime: further out, two research snapshots survey what a *native* driver for real GPU silicon would take as another `.scanout` backend: [nvidia-gpus.md](nvidia-gpus.md) (RTX 3060 / Ampere) and [intel-igpu.md](intel-igpu.md) (Intel iGPU). -20. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and +21. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and how `while (true) hlt` parks the CPU safely once there's nothing left to do. Start with the north star: @@ -99,6 +104,12 @@ Start with the north star: port to **one seam** (`std.os.danos`), so we build `runtime.os` (→ that seam) plus a thin `runtime.fs`, retire the `posix` shim, and follow a phased path to `zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation. +- **[vdso.md](vdso.md) — the vDSO, the public system-call boundary.** A design note + (not built yet) on keeping `abi.zig` genuinely private: a kernel-supplied, C-ABI + entry blob mapped into every process as the *only* way into the kernel — so the + syscall numbers can be renumbered or randomised at will, and Rust/C binaries get a + stable boundary without danos growing a dynamic linker. danos's public ABI = the + vDSO + the documented IPC wire protocols ([vfs-protocol.md](vfs-protocol.md) first). Cutting across all of these: diff --git a/docs/vdso.md b/docs/vdso.md new file mode 100644 index 0000000..34203cd --- /dev/null +++ b/docs/vdso.md @@ -0,0 +1,206 @@ +# The vDSO — the public system-call boundary + +> **Status:** design note, not built. The runtime today issues raw `syscall` +> instructions from `library/runtime/system-call.zig` using the numbers in +> `system/abi.zig`. This note designs the layer that replaces that arrangement: +> a **kernel-supplied, C-ABI entry library** mapped into every process — the +> only supported way into the kernel — so the raw numbers can stay private, +> be renumbered at will, and eventually be randomised per boot. + +## Why: the ABI danos promises, and the one it doesn't + +`system/abi.zig` is the **private** kernel ↔ runtime contract. Its header says +so: the numbers are an implementation detail the runtime hides and may +renumber, the same split as libSystem over the XNU syscalls on macOS or win32 +over the NT syscalls on Windows. Linux — with its world-visible, frozen +syscall table — is the outlier, not the norm. + +That stance has consequences the moment binaries exist that we don't rebuild +ourselves: + +1. **Third-party binaries** (docs/zig-self-hosting.md) must keep working across + kernel updates. If they contain raw `syscall` instructions with today's + numbers baked in, every renumbering breaks the world — the ABI would be + *de facto* public no matter what the header says. Go on macOS made exactly + this mistake: it issued XNU syscalls directly instead of going through + libSystem, and macOS updates repeatedly broke every Go binary until Go + switched to the library like everyone else. +2. **Not everything is Zig.** A Rust or C program can't import the `runtime` + module. The public boundary has to be expressible in the one calling + convention every language speaks: the C ABI. +3. **Randomised syscall numbers** — a hardening option we want open — only + work if no user binary anywhere knows a number at build time. The binding + must happen at *load time*, from something the kernel controls. + +All three point at the same well-known shape: a **vDSO** (virtual dynamic +shared object). The kernel carries a small blob of user-mode code, maps it +into every process at spawn, and that blob — not the application — contains +the `syscall` instructions. Fuchsia works exactly this way: its vDSO is the +*only* kernel entry, version-matched by construction because the kernel itself +injects it. Because the kernel and the blob ship as one artifact, there is +**no version skew, no loader, no search path, and no shared file on disk** — +which is what makes this the resilient way to have a private ABI +(docs/resilience.md), where a conventional `ld.so` + `/lib/libdanos.so` +arrangement would add a loader to every spawn and a single shared point of +failure. + +The public danos ABI then has exactly two layers, neither of which is +`abi.zig`: + +| Layer | Contract | Spoken by | +|-------|----------|-----------| +| **vDSO** | C-ABI functions, this note | every language's thin shim (`runtime.system` for Zig, a `-sys` crate for Rust, a header for C) | +| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](vfs-protocol.md) is the first one documented) | any client that can lay out bytes | + +Everything above those — the heap, `runtime.fs`, the service harness — is +per-language convenience, compiled into each binary from source, exactly as +today. Nothing about the Zig runtime's shape changes; it just stops being the +*only* door. + +## The blob + +A single copy of the vDSO code lives in the kernel image (built by +`build.zig` as a tiny freestanding object, embedded like the AP trampoline). +At boot the kernel finalises it once — this is where randomised numbers would +be patched in — and thereafter maps the **same physical pages** read-execute +into every process's address space. The blob is: + +- **Position-independent.** It is mapped at a per-process randomised base, so + it must be PIC (rip-relative addressing only — no relocations to process). +- **Stateless and re-entrant.** No writable data. Anything stateful belongs to + the process, not the vDSO. +- **Architecture-specific.** The x86-64 blob wraps `syscall`; an aarch64 blob + wraps `svc #0`. It lives beside the other per-architecture kernel sources + (`system/kernel/architecture//`), selected the same way the + `architecture` module is (docs/arch.md). + +### Shape: a function table, not an ELF + +A real `.so` with a dynamic symbol table is the conventional vDSO shape, but +linking against one at load time needs a dynamic linker in every binary — +machinery danos deliberately doesn't have. Instead the v1 shape is the +simplest thing that is still a stable contract — a **function-pointer table** +at the vDSO base: + +``` +offset 0 u64 magic 'danosVDS' — a mapped-the-wrong-thing guard +offset 8 u64 api_level incremented when the table grows +offset 16 u64 count number of table entries that follow +offset 24 u64 table[count] function pointers into the vDSO's own code +``` + +Table *indices* are the public constants (published in a C header, +`danos.h`), assigned once and append-only — the same discipline the IPC +protocols use for operation values. The pointers point at stubs inside the +blob; what those stubs put in `rax` is nobody's business but the kernel's. +A language shim binds in one step: read the base from the init block, check +the magic, keep the table pointer. Feature detection for a binary built +against older headers is `count`/`api_level` — a kernel never removes or +reorders entries. + +(If danos ever grows a real dynamic linker, the same blob can additionally +present an ELF `dynsym` without breaking the table — Fuchsia's vDSO is +likewise both a mappable blob and a linkable `.so`. That is a later +convenience, not a requirement.) + +### Delivery: the auxiliary vector + +The kernel already builds a System V entry block — argc, argv, envp +terminator, **auxiliary vector** — on every new process's stack +(`buildEntryStack`, read by `runtime.start`). The vDSO base rides in a new +auxv entry, exactly Linux's `AT_SYSINFO_EHDR` move. No new syscall, no magic +address, and a language shim finds it the same portable way on every +architecture. + +## The function surface + +One table entry per kernel call, C ABI (System V AMD64), names prefixed +`danos_`. The current `SystemCall` set maps directly; integer arguments and +returns are `u64`, errors return as negative values exactly as today. + +The calls that return two values in `rax:rdx` today — `dma_alloc` +(vaddr + paddr), `msi_bind` (address + data), `shm_create` (vaddr + handle) — +become functions returning a two-`u64` struct. The System V ABI returns a +16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the +C-ABI spelling of the existing convention, at zero cost. + +Grouped as `abi.zig` groups them: + +| Group | Functions | +|-------|-----------| +| process | `danos_exit`, `danos_yield`, `danos_sleep`, `danos_spawn`, `danos_process_enumerate`, `danos_process_kill`, `danos_process_exit_reason`, `danos_process_subscribe`, `danos_process_signal`, `danos_signal_bind` | +| memory | `danos_mmap`, `danos_munmap`, `danos_dma_alloc`, `danos_dma_free`, `danos_shm_create`, `danos_shm_map`, `danos_shm_physical` | +| ipc | `danos_endpoint_create`, `danos_ipc_register`, `danos_ipc_lookup`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` | +| devices | `danos_device_enumerate`, `danos_device_claim`, `danos_device_register`, `danos_mmio_map`, `danos_irq_bind`, `danos_irq_ack`, `danos_msi_bind`, `danos_io_read`, `danos_io_write` | +| time | `danos_clock`, `danos_wall_clock`, `danos_timer_bind` | +| diagnostics | `danos_debug_write`, `danos_klog_read` | + +The constants that ride alongside the calls — mmap protection bits, DMA +flags, notification badge bits, `ExitReason`, `Signal`, well-known service +ids, `page_size`, the IPC message maximum — move to the public header too: +they are wire values a Rust program needs verbatim. What stays private in +`abi.zig` is exactly the thing the vDSO exists to hide: the `SystemCall` +numbers and the trap convention. + +## Enforcement, and an honest threat model + +Renumbering only has teeth if the kernel **refuses syscalls that don't come +from the vDSO**. The check is cheap: on kernel entry, the saved user `rip` +must lie inside the calling process's vDSO mapping; otherwise the process is +killed with a fault-class exit reason (its supervisor restarts or gives up, +docs/process-lifecycle.md — a foreign-syscall attempt is a bug or an attack, +never something to limp past). Fuchsia enforces exactly this. + +What this buys, precisely: + +- **ABI freedom** — the real prize. The numbers can change per release or per + boot and nothing outside the kernel image cares. The private ABI stays + actually private, permanently. +- **A single audited chokepoint** for kernel entry, per process, at a + randomised address. +- **Raised bar for exploits**: shellcode can't issue a hard-coded `syscall`; + it must first discover the per-process vDSO base (ASLR) and call through + it. + +What it does *not* buy: an attacker with arbitrary code execution in a +process can still *call* the vDSO functions — they are mapped executable in +that process, and return-oriented chains reach them. Syscall randomisation is +hardening, not a security boundary; the security boundary remains the +capability model (what the process's endpoints and device claims let it do). +It is worth building anyway — for the ABI freedom first and the hardening +second — but the design should never be sold as more than that. + +## Migration + +Phased so every step ships alone (the M-milestone discipline): + +1. **The blob + the table.** Build the vDSO, map it at spawn, deliver the + base via auxv. `runtime.system-call.zig` binds through the table when the + auxv entry is present, falls back to raw `syscall` when absent — the whole + tree keeps booting during the transition. +2. **Cut the runtime over.** Delete the raw stubs; `runtime` no longer + imports the `SystemCall` numbers at all (`abi.zig`'s enum becomes + kernel-internal). The QEMU suite passing proves the table carries the + whole system. +3. **Enforce + randomise.** Add the `rip`-range check, then per-boot number + randomisation patched into the blob at kernel init. A test boots with + randomisation on and runs the full suite. +4. **The other languages.** Publish `danos.h`; a Rust `danos-sys` crate wraps + the table. This is also the seam `std.os.danos` calls through when the Zig + self-hosting fork lands (docs/zig-self-hosting.md) — the vDSO is what + makes that seam stable across kernel versions. + +## What deliberately stays out + +- **No dynamic linker, no `/lib/*.so`.** The vDSO is kernel-injected precisely + so danos binaries can stay fully static above it. Sharing *library code* + across processes stays what it is today: a service behind IPC, or source + compiled into each binary. +- **No file/device I/O in the vDSO.** The microkernel line doesn't move: the + vDSO wraps the same deliberately tiny table (docs/syscall.md); files are + still the VFS server's business over IPC. +- **No fast-path user-mode implementations yet.** Linux's vDSO exists mostly + to answer `gettimeofday` without a kernel entry. `danos_clock` could one + day read the calibrated TSC in user mode the same way — the blob is where + such an optimisation would live — but that is an optimisation, not part of + this design's contract. diff --git a/docs/vfs-protocol.md b/docs/vfs-protocol.md new file mode 100644 index 0000000..33298db --- /dev/null +++ b/docs/vfs-protocol.md @@ -0,0 +1,170 @@ +# The VFS wire protocol + +> **Status:** built and spoken today between `runtime.fs` (the client) and the +> VFS server (`system/services/vfs`), with mounted backends (the FAT server) +> speaking the same protocol behind the router. The Zig source of truth is +> `system/services/vfs/protocol.zig` (the `vfs-protocol` module), whose unit +> tests pin the sizes and values below. This page is the **language-neutral +> wire specification** of that contract — what a Rust or C client implements +> ([vdso.md](vdso.md) explains why the IPC protocols, not the syscall +> numbers, are danos's public ABI). + +## Transport + +A VFS exchange is one synchronous IPC rendezvous (`ipc_call`, +docs/ipc.md): the client sends one message and blocks; the server replies +with one message. The endpoint is found by well-known service id +(`ipc_lookup`, service id **1** = vfs). + +- A message is at most **256 bytes** (`message_maximum`). +- A request is a fixed 32-byte **Request** header followed by an inline + payload of at most **224 bytes** (`maximum_payload`) — a path, or write + bytes. There is no multi-message request: paths and single reads/writes + must fit, and larger transfers loop (see *read* / *write*). +- A reply is a fixed 24-byte **Reply** header followed by an inline payload — + read bytes, a `FileStatus`, or a `DirectoryEntry`. +- All integers are **little-endian**; layouts are C layout for x86-64 + (`extern struct`), offsets given below so nothing need be inferred. + +The kernel never parses any of this — it only moves the bytes +(docs/syscall.md); files are entirely a user-space affair. + +## Request header — 32 bytes + +| offset | size | field | meaning | +|-------:|-----:|-------|---------| +| 0 | 4 | `operation` | an **Operation** value (below) | +| 4 | 4 | — | padding | +| 8 | 8 | `node` | the server-side open-node id from a prior `open`; 0 for path-based operations | +| 16 | 8 | `offset` | byte position for read/write; entry index (cursor) for readdir; else 0 | +| 24 | 4 | `len` | payload length for path/write operations; requested byte count for read | +| 28 | 4 | `flags` | open flags (below); else 0 | + +## Reply header — 24 bytes + +| offset | size | field | meaning | +|-------:|-----:|-------|---------| +| 0 | 4 | `status` | **0 = success**, negative = failure (signed) | +| 4 | 4 | — | padding | +| 8 | 8 | `node` | the new open-node id (for `open`); else 0 | +| 16 | 4 | `len` | reply payload length in bytes | +| 20 | 4 | — | padding | + +On failure the router replies `status = -1`; a mounted backend's negative +status is forwarded to the client verbatim. A richer errno vocabulary is +future work — clients must treat *any* negative status as failure, not match +on -1. + +## Operations + +Values are append-only and never renumbered (the same evolution rule every +danos protocol follows); an unrecognised operation gets a `status = -1` +reply. + +| value | operation | request payload | reply | +|------:|-----------|-----------------|-------| +| 0 | `open` | the path (`len` = its length), `flags` as below | `node` = open-node id | +| 1 | `close` | — (`node` set) | status only | +| 2 | `read` | — (`node`, `offset`, `len` = wanted count) | `len` bytes read, payload = the bytes; `len` 0 at end of file | +| 3 | `write` | the bytes (`node`, `offset`, `len` = count) | `len` = bytes accepted (may be short — loop) | +| 4 | `status` | — (`node` set) | payload = **FileStatus** (24 bytes) | +| 5 | `readdir` | — (`node` = a directory, `offset` = cursor) | payload = one **DirectoryEntry** + name; `len` 0 at end | +| 6 | `mount` | the mount-point path; the backend endpoint rides as the call's **capability** | status only | +| 7 | `unmount` | the mount-point path | status only | +| 8 | `mkdir` | the path | status only | +| 9 | `unlink` | the path | status only | +| 10 | `rename` | old path, one `0x00`, new path (`len` = total) | status only | + +Notes per operation: + +- **open** — paths are absolute (`/mnt/usb/notes.txt`) or bare names + (`greeting`); bare names resolve in the VFS's flat ramfs, absolute paths + route through the mount table (below). The returned `node` is an id in the + *router's* open table; clients never see a backend's own ids. +- **read / write** — a single exchange moves at most 224 bytes + (`maximum_payload`); the client loops, advancing `offset` by the returned + `len`, until done (read) or the slice is written (write). A `write` reply + shorter than requested is progress, not an error; a `len` of 0 means no + forward progress — stop rather than spin. +- **readdir** — `offset` is a **cursor: the entry index**, not a byte + position. Each call returns exactly one entry; the client increments the + cursor by 1. A reply with `len` 0 is end-of-directory. The directory must + have been opened with the `directory` flag. +- **mount** — the one operation that passes a **capability**: the caller + (a filesystem server, e.g. FAT) sends its own request endpoint as the + `ipc_call` capability argument, and the router forwards everything under + the mount point to it — speaking this same protocol, with paths rewritten + relative to the mount. Prefixes match at path boundaries only + (`/mnt/usb` never captures `/mnt/usbextra`); the longest matching prefix + wins. +- **rename** — same-directory rename only (the router requires old and new to + resolve under one mount). + +## Open flags + +Bitwise OR in `Request.flags`, meaningful for `open` only: + +| bit | name | meaning | +|----:|------|---------| +| 1 | `create` | create the file if it does not exist | +| 2 | `directory` | open a directory node for `readdir` rather than a file | +| 4 | `truncate` | truncate an existing file to zero length on open (replace, don't overwrite in place) | + +## FileStatus — 24 bytes (the `status` reply payload) + +| offset | size | field | meaning | +|-------:|-----:|-------|---------| +| 0 | 8 | `size` | file size in bytes | +| 8 | 4 | `kind` | a **NodeKind** value | +| 12 | 4 | — | padding | +| 16 | 8 | `mtime` | modification time, Unix epoch seconds UTC; 0 if the backend keeps none | + +## DirectoryEntry — 16 bytes + name (the `readdir` reply payload) + +| offset | size | field | meaning | +|-------:|-----:|-------|---------| +| 0 | 4 | `kind` | a **NodeKind** value | +| 4 | 4 | `name_len` | length of the name that follows | +| 8 | 8 | `size` | the entry's size in bytes | +| 16 | `name_len` | name | the entry's name, not NUL-terminated | + +## NodeKind + +Aligned to the FSH file-type table +(docs/danos-file-system-hierarchy-FSH.md): + +| value | kind | +|------:|------| +| 0 | regular file | +| 1 | directory | +| 2 | character device | +| 3 | block device | +| 4 | symbolic link | +| 5 | fifo | +| 6 | socket | + +Clients should map unknown values to *regular* rather than reject — the +table can grow. + +## Lifetimes and trust + +Open-node ids live in the server. A client that dies without closing leaks +nothing permanently: the VFS subscribes to the kernel's published process-exit +events (docs/process-lifecycle.md) and releases a dead client's handles, +closing forwarded backend nodes best-effort. Ids are plain integers, not +capabilities — the VFS trusts its callers with each other's ids today, which +is acceptable while every client is part of the system image and worth +revisiting (per-client id namespaces) before third-party binaries arrive. + +## Evolution rules + +What a non-Zig implementation may rely on, and what it must not: + +- Operation values, flag bits, `NodeKind` values, and struct layouts are + **append-only and frozen once shipped** — the unit tests in `protocol.zig` + pin them exactly so a refactor can't silently move them. +- The 256-byte message ceiling is a property of the current IPC transport, + not a promise; clients should read `maximum_payload`-shaped limits from the + reply lengths they actually get (loop-until-done), not hard-code 224. +- Negative statuses beyond -1 will appear (an errno vocabulary); success is + exactly 0.