docs: vdso + vfs-protocol — the public ABI boundary, linked from the index
This commit is contained in:
+20
-9
@@ -39,37 +39,42 @@ rather than restate it. Roughly in the order things happen at runtime:
|
|||||||
endpoints — the backbone the microkernel's isolated servers talk over.
|
endpoints — the backbone the microkernel's isolated servers talk over.
|
||||||
12. **[syscall.md](syscall.md) — system calls.** How ring 3 asks the kernel for
|
12. **[syscall.md](syscall.md) — system calls.** How ring 3 asks the kernel for
|
||||||
something: the `syscall`/`sysret` fast path, the trap frame, and why the table is
|
something: the `syscall`/`sysret` fast path, the trap frame, and why the table is
|
||||||
deliberately tiny.
|
deliberately tiny. The numbers are a **private** ABI — [vdso.md](vdso.md) designs
|
||||||
13. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an
|
the public boundary that will hide them.
|
||||||
|
13. **[vfs-protocol.md](vfs-protocol.md) — the VFS wire protocol.** The language-neutral
|
||||||
|
byte-level spec of the file protocol spoken over IPC: request/reply headers,
|
||||||
|
the operation table, mount routing, and the append-only evolution rules — the
|
||||||
|
first IPC protocol documented as public ABI.
|
||||||
|
14. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an
|
||||||
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
||||||
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
||||||
unmask.
|
unmask.
|
||||||
14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
|
15. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
|
||||||
real driver stacks factor into three shapes and how families share code. The
|
real driver stacks factor into three shapes and how families share code. The
|
||||||
three primitives it proposed are long since built (M13 capability passing,
|
three primitives it proposed are long since built (M13 capability passing,
|
||||||
M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them —
|
M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them —
|
||||||
hello, supervision, restart — is built too (device-manager.md, M18).
|
hello, supervision, restart — is built too (device-manager.md, M18).
|
||||||
15. **[process-management.md](process-management.md) — process management.** The
|
16. **[process-management.md](process-management.md) — process management.** The
|
||||||
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
|
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
|
||||||
supervision link as the kill authority, and child-exit notifications over the
|
supervision link as the kill authority, and child-exit notifications over the
|
||||||
same endpoints IRQs arrive on.
|
same endpoints IRQs arrive on.
|
||||||
16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
|
17. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
|
||||||
(M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
|
(M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
|
||||||
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
|
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
|
||||||
`runtime.process` interface, exit reasons, published exit events any stateful
|
`runtime.process` interface, exit reasons, published exit events any stateful
|
||||||
service can subscribe to (the VFS releasing dead clients' handles), and the two
|
service can subscribe to (the VFS releasing dead clients' handles), and the two
|
||||||
iron rules (cleanup is the kernel's job; kill is not a signal).
|
iron rules (cleanup is the kernel's job; kill is not a signal).
|
||||||
17. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
|
18. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
|
||||||
through the app surface): the
|
through the app surface): the
|
||||||
tree, the matcher, and the supervisor. Tree structure lives in the manager,
|
tree, the matcher, and the supervisor. Tree structure lives in the manager,
|
||||||
authority stays in the kernel; bus drivers report what they see; drivers are
|
authority stays in the kernel; bus drivers report what they see; drivers are
|
||||||
restarted through the lifecycle vocabulary — the plan that turns
|
restarted through the lifecycle vocabulary — the plan that turns
|
||||||
[resilience.md](resilience.md)'s restart goal into increments.
|
[resilience.md](resilience.md)'s restart goal into increments.
|
||||||
18. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard,
|
19. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard,
|
||||||
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
|
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
|
||||||
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
|
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
|
||||||
service layered on top.
|
service layered on top.
|
||||||
19. **[display.md](display.md) — the display service.** The display half of the GUI
|
20. **[display.md](display.md) — the display service.** The display half of the GUI
|
||||||
track: a user-space compositor that owns the framebuffer, composes a layer stack into
|
track: a user-space compositor that owns the framebuffer, composes a layer stack into
|
||||||
a double buffer, and presents it. Why GOP and the PCI display device are two views of
|
a double buffer, and presents it. Why GOP and the PCI display device are two views of
|
||||||
one controller, the device-node + write-combining handoff, and what flicker-free buys
|
one controller, the device-node + write-combining handoff, and what flicker-free buys
|
||||||
@@ -80,7 +85,7 @@ rather than restate it. Roughly in the order things happen at runtime:
|
|||||||
further out, two research snapshots survey what a *native* driver for real GPU silicon
|
further out, two research snapshots survey what a *native* driver for real GPU silicon
|
||||||
would take as another `.scanout` backend: [nvidia-gpus.md](nvidia-gpus.md) (RTX 3060 /
|
would take as another `.scanout` backend: [nvidia-gpus.md](nvidia-gpus.md) (RTX 3060 /
|
||||||
Ampere) and [intel-igpu.md](intel-igpu.md) (Intel iGPU).
|
Ampere) and [intel-igpu.md](intel-igpu.md) (Intel iGPU).
|
||||||
20. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
21. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||||
|
|
||||||
Start with the north star:
|
Start with the north star:
|
||||||
@@ -99,6 +104,12 @@ Start with the north star:
|
|||||||
port to **one seam** (`std.os.danos`), so we build `runtime.os` (→ that seam) plus a
|
port to **one seam** (`std.os.danos`), so we build `runtime.os` (→ that seam) plus a
|
||||||
thin `runtime.fs`, retire the `posix` shim, and follow a phased path to
|
thin `runtime.fs`, retire the `posix` shim, and follow a phased path to
|
||||||
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
|
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
|
||||||
|
- **[vdso.md](vdso.md) — the vDSO, the public system-call boundary.** A design note
|
||||||
|
(not built yet) on keeping `abi.zig` genuinely private: a kernel-supplied, C-ABI
|
||||||
|
entry blob mapped into every process as the *only* way into the kernel — so the
|
||||||
|
syscall numbers can be renumbered or randomised at will, and Rust/C binaries get a
|
||||||
|
stable boundary without danos growing a dynamic linker. danos's public ABI = the
|
||||||
|
vDSO + the documented IPC wire protocols ([vfs-protocol.md](vfs-protocol.md) first).
|
||||||
|
|
||||||
Cutting across all of these:
|
Cutting across all of these:
|
||||||
|
|
||||||
|
|||||||
+206
@@ -0,0 +1,206 @@
|
|||||||
|
# The vDSO — the public system-call boundary
|
||||||
|
|
||||||
|
> **Status:** design note, not built. The runtime today issues raw `syscall`
|
||||||
|
> instructions from `library/runtime/system-call.zig` using the numbers in
|
||||||
|
> `system/abi.zig`. This note designs the layer that replaces that arrangement:
|
||||||
|
> a **kernel-supplied, C-ABI entry library** mapped into every process — the
|
||||||
|
> only supported way into the kernel — so the raw numbers can stay private,
|
||||||
|
> be renumbered at will, and eventually be randomised per boot.
|
||||||
|
|
||||||
|
## Why: the ABI danos promises, and the one it doesn't
|
||||||
|
|
||||||
|
`system/abi.zig` is the **private** kernel ↔ runtime contract. Its header says
|
||||||
|
so: the numbers are an implementation detail the runtime hides and may
|
||||||
|
renumber, the same split as libSystem over the XNU syscalls on macOS or win32
|
||||||
|
over the NT syscalls on Windows. Linux — with its world-visible, frozen
|
||||||
|
syscall table — is the outlier, not the norm.
|
||||||
|
|
||||||
|
That stance has consequences the moment binaries exist that we don't rebuild
|
||||||
|
ourselves:
|
||||||
|
|
||||||
|
1. **Third-party binaries** (docs/zig-self-hosting.md) must keep working across
|
||||||
|
kernel updates. If they contain raw `syscall` instructions with today's
|
||||||
|
numbers baked in, every renumbering breaks the world — the ABI would be
|
||||||
|
*de facto* public no matter what the header says. Go on macOS made exactly
|
||||||
|
this mistake: it issued XNU syscalls directly instead of going through
|
||||||
|
libSystem, and macOS updates repeatedly broke every Go binary until Go
|
||||||
|
switched to the library like everyone else.
|
||||||
|
2. **Not everything is Zig.** A Rust or C program can't import the `runtime`
|
||||||
|
module. The public boundary has to be expressible in the one calling
|
||||||
|
convention every language speaks: the C ABI.
|
||||||
|
3. **Randomised syscall numbers** — a hardening option we want open — only
|
||||||
|
work if no user binary anywhere knows a number at build time. The binding
|
||||||
|
must happen at *load time*, from something the kernel controls.
|
||||||
|
|
||||||
|
All three point at the same well-known shape: a **vDSO** (virtual dynamic
|
||||||
|
shared object). The kernel carries a small blob of user-mode code, maps it
|
||||||
|
into every process at spawn, and that blob — not the application — contains
|
||||||
|
the `syscall` instructions. Fuchsia works exactly this way: its vDSO is the
|
||||||
|
*only* kernel entry, version-matched by construction because the kernel itself
|
||||||
|
injects it. Because the kernel and the blob ship as one artifact, there is
|
||||||
|
**no version skew, no loader, no search path, and no shared file on disk** —
|
||||||
|
which is what makes this the resilient way to have a private ABI
|
||||||
|
(docs/resilience.md), where a conventional `ld.so` + `/lib/libdanos.so`
|
||||||
|
arrangement would add a loader to every spawn and a single shared point of
|
||||||
|
failure.
|
||||||
|
|
||||||
|
The public danos ABI then has exactly two layers, neither of which is
|
||||||
|
`abi.zig`:
|
||||||
|
|
||||||
|
| Layer | Contract | Spoken by |
|
||||||
|
|-------|----------|-----------|
|
||||||
|
| **vDSO** | C-ABI functions, this note | every language's thin shim (`runtime.system` for Zig, a `-sys` crate for Rust, a header for C) |
|
||||||
|
| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](vfs-protocol.md) is the first one documented) | any client that can lay out bytes |
|
||||||
|
|
||||||
|
Everything above those — the heap, `runtime.fs`, the service harness — is
|
||||||
|
per-language convenience, compiled into each binary from source, exactly as
|
||||||
|
today. Nothing about the Zig runtime's shape changes; it just stops being the
|
||||||
|
*only* door.
|
||||||
|
|
||||||
|
## The blob
|
||||||
|
|
||||||
|
A single copy of the vDSO code lives in the kernel image (built by
|
||||||
|
`build.zig` as a tiny freestanding object, embedded like the AP trampoline).
|
||||||
|
At boot the kernel finalises it once — this is where randomised numbers would
|
||||||
|
be patched in — and thereafter maps the **same physical pages** read-execute
|
||||||
|
into every process's address space. The blob is:
|
||||||
|
|
||||||
|
- **Position-independent.** It is mapped at a per-process randomised base, so
|
||||||
|
it must be PIC (rip-relative addressing only — no relocations to process).
|
||||||
|
- **Stateless and re-entrant.** No writable data. Anything stateful belongs to
|
||||||
|
the process, not the vDSO.
|
||||||
|
- **Architecture-specific.** The x86-64 blob wraps `syscall`; an aarch64 blob
|
||||||
|
wraps `svc #0`. It lives beside the other per-architecture kernel sources
|
||||||
|
(`system/kernel/architecture/<arch>/`), selected the same way the
|
||||||
|
`architecture` module is (docs/arch.md).
|
||||||
|
|
||||||
|
### Shape: a function table, not an ELF
|
||||||
|
|
||||||
|
A real `.so` with a dynamic symbol table is the conventional vDSO shape, but
|
||||||
|
linking against one at load time needs a dynamic linker in every binary —
|
||||||
|
machinery danos deliberately doesn't have. Instead the v1 shape is the
|
||||||
|
simplest thing that is still a stable contract — a **function-pointer table**
|
||||||
|
at the vDSO base:
|
||||||
|
|
||||||
|
```
|
||||||
|
offset 0 u64 magic 'danosVDS' — a mapped-the-wrong-thing guard
|
||||||
|
offset 8 u64 api_level incremented when the table grows
|
||||||
|
offset 16 u64 count number of table entries that follow
|
||||||
|
offset 24 u64 table[count] function pointers into the vDSO's own code
|
||||||
|
```
|
||||||
|
|
||||||
|
Table *indices* are the public constants (published in a C header,
|
||||||
|
`danos.h`), assigned once and append-only — the same discipline the IPC
|
||||||
|
protocols use for operation values. The pointers point at stubs inside the
|
||||||
|
blob; what those stubs put in `rax` is nobody's business but the kernel's.
|
||||||
|
A language shim binds in one step: read the base from the init block, check
|
||||||
|
the magic, keep the table pointer. Feature detection for a binary built
|
||||||
|
against older headers is `count`/`api_level` — a kernel never removes or
|
||||||
|
reorders entries.
|
||||||
|
|
||||||
|
(If danos ever grows a real dynamic linker, the same blob can additionally
|
||||||
|
present an ELF `dynsym` without breaking the table — Fuchsia's vDSO is
|
||||||
|
likewise both a mappable blob and a linkable `.so`. That is a later
|
||||||
|
convenience, not a requirement.)
|
||||||
|
|
||||||
|
### Delivery: the auxiliary vector
|
||||||
|
|
||||||
|
The kernel already builds a System V entry block — argc, argv, envp
|
||||||
|
terminator, **auxiliary vector** — on every new process's stack
|
||||||
|
(`buildEntryStack`, read by `runtime.start`). The vDSO base rides in a new
|
||||||
|
auxv entry, exactly Linux's `AT_SYSINFO_EHDR` move. No new syscall, no magic
|
||||||
|
address, and a language shim finds it the same portable way on every
|
||||||
|
architecture.
|
||||||
|
|
||||||
|
## The function surface
|
||||||
|
|
||||||
|
One table entry per kernel call, C ABI (System V AMD64), names prefixed
|
||||||
|
`danos_`. The current `SystemCall` set maps directly; integer arguments and
|
||||||
|
returns are `u64`, errors return as negative values exactly as today.
|
||||||
|
|
||||||
|
The calls that return two values in `rax:rdx` today — `dma_alloc`
|
||||||
|
(vaddr + paddr), `msi_bind` (address + data), `shm_create` (vaddr + handle) —
|
||||||
|
become functions returning a two-`u64` struct. The System V ABI returns a
|
||||||
|
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
|
||||||
|
C-ABI spelling of the existing convention, at zero cost.
|
||||||
|
|
||||||
|
Grouped as `abi.zig` groups them:
|
||||||
|
|
||||||
|
| Group | Functions |
|
||||||
|
|-------|-----------|
|
||||||
|
| process | `danos_exit`, `danos_yield`, `danos_sleep`, `danos_spawn`, `danos_process_enumerate`, `danos_process_kill`, `danos_process_exit_reason`, `danos_process_subscribe`, `danos_process_signal`, `danos_signal_bind` |
|
||||||
|
| memory | `danos_mmap`, `danos_munmap`, `danos_dma_alloc`, `danos_dma_free`, `danos_shm_create`, `danos_shm_map`, `danos_shm_physical` |
|
||||||
|
| ipc | `danos_endpoint_create`, `danos_ipc_register`, `danos_ipc_lookup`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` |
|
||||||
|
| devices | `danos_device_enumerate`, `danos_device_claim`, `danos_device_register`, `danos_mmio_map`, `danos_irq_bind`, `danos_irq_ack`, `danos_msi_bind`, `danos_io_read`, `danos_io_write` |
|
||||||
|
| time | `danos_clock`, `danos_wall_clock`, `danos_timer_bind` |
|
||||||
|
| diagnostics | `danos_debug_write`, `danos_klog_read` |
|
||||||
|
|
||||||
|
The constants that ride alongside the calls — mmap protection bits, DMA
|
||||||
|
flags, notification badge bits, `ExitReason`, `Signal`, well-known service
|
||||||
|
ids, `page_size`, the IPC message maximum — move to the public header too:
|
||||||
|
they are wire values a Rust program needs verbatim. What stays private in
|
||||||
|
`abi.zig` is exactly the thing the vDSO exists to hide: the `SystemCall`
|
||||||
|
numbers and the trap convention.
|
||||||
|
|
||||||
|
## Enforcement, and an honest threat model
|
||||||
|
|
||||||
|
Renumbering only has teeth if the kernel **refuses syscalls that don't come
|
||||||
|
from the vDSO**. The check is cheap: on kernel entry, the saved user `rip`
|
||||||
|
must lie inside the calling process's vDSO mapping; otherwise the process is
|
||||||
|
killed with a fault-class exit reason (its supervisor restarts or gives up,
|
||||||
|
docs/process-lifecycle.md — a foreign-syscall attempt is a bug or an attack,
|
||||||
|
never something to limp past). Fuchsia enforces exactly this.
|
||||||
|
|
||||||
|
What this buys, precisely:
|
||||||
|
|
||||||
|
- **ABI freedom** — the real prize. The numbers can change per release or per
|
||||||
|
boot and nothing outside the kernel image cares. The private ABI stays
|
||||||
|
actually private, permanently.
|
||||||
|
- **A single audited chokepoint** for kernel entry, per process, at a
|
||||||
|
randomised address.
|
||||||
|
- **Raised bar for exploits**: shellcode can't issue a hard-coded `syscall`;
|
||||||
|
it must first discover the per-process vDSO base (ASLR) and call through
|
||||||
|
it.
|
||||||
|
|
||||||
|
What it does *not* buy: an attacker with arbitrary code execution in a
|
||||||
|
process can still *call* the vDSO functions — they are mapped executable in
|
||||||
|
that process, and return-oriented chains reach them. Syscall randomisation is
|
||||||
|
hardening, not a security boundary; the security boundary remains the
|
||||||
|
capability model (what the process's endpoints and device claims let it do).
|
||||||
|
It is worth building anyway — for the ABI freedom first and the hardening
|
||||||
|
second — but the design should never be sold as more than that.
|
||||||
|
|
||||||
|
## Migration
|
||||||
|
|
||||||
|
Phased so every step ships alone (the M-milestone discipline):
|
||||||
|
|
||||||
|
1. **The blob + the table.** Build the vDSO, map it at spawn, deliver the
|
||||||
|
base via auxv. `runtime.system-call.zig` binds through the table when the
|
||||||
|
auxv entry is present, falls back to raw `syscall` when absent — the whole
|
||||||
|
tree keeps booting during the transition.
|
||||||
|
2. **Cut the runtime over.** Delete the raw stubs; `runtime` no longer
|
||||||
|
imports the `SystemCall` numbers at all (`abi.zig`'s enum becomes
|
||||||
|
kernel-internal). The QEMU suite passing proves the table carries the
|
||||||
|
whole system.
|
||||||
|
3. **Enforce + randomise.** Add the `rip`-range check, then per-boot number
|
||||||
|
randomisation patched into the blob at kernel init. A test boots with
|
||||||
|
randomisation on and runs the full suite.
|
||||||
|
4. **The other languages.** Publish `danos.h`; a Rust `danos-sys` crate wraps
|
||||||
|
the table. This is also the seam `std.os.danos` calls through when the Zig
|
||||||
|
self-hosting fork lands (docs/zig-self-hosting.md) — the vDSO is what
|
||||||
|
makes that seam stable across kernel versions.
|
||||||
|
|
||||||
|
## What deliberately stays out
|
||||||
|
|
||||||
|
- **No dynamic linker, no `/lib/*.so`.** The vDSO is kernel-injected precisely
|
||||||
|
so danos binaries can stay fully static above it. Sharing *library code*
|
||||||
|
across processes stays what it is today: a service behind IPC, or source
|
||||||
|
compiled into each binary.
|
||||||
|
- **No file/device I/O in the vDSO.** The microkernel line doesn't move: the
|
||||||
|
vDSO wraps the same deliberately tiny table (docs/syscall.md); files are
|
||||||
|
still the VFS server's business over IPC.
|
||||||
|
- **No fast-path user-mode implementations yet.** Linux's vDSO exists mostly
|
||||||
|
to answer `gettimeofday` without a kernel entry. `danos_clock` could one
|
||||||
|
day read the calibrated TSC in user mode the same way — the blob is where
|
||||||
|
such an optimisation would live — but that is an optimisation, not part of
|
||||||
|
this design's contract.
|
||||||
@@ -0,0 +1,170 @@
|
|||||||
|
# The VFS wire protocol
|
||||||
|
|
||||||
|
> **Status:** built and spoken today between `runtime.fs` (the client) and the
|
||||||
|
> VFS server (`system/services/vfs`), with mounted backends (the FAT server)
|
||||||
|
> speaking the same protocol behind the router. The Zig source of truth is
|
||||||
|
> `system/services/vfs/protocol.zig` (the `vfs-protocol` module), whose unit
|
||||||
|
> tests pin the sizes and values below. This page is the **language-neutral
|
||||||
|
> wire specification** of that contract — what a Rust or C client implements
|
||||||
|
> ([vdso.md](vdso.md) explains why the IPC protocols, not the syscall
|
||||||
|
> numbers, are danos's public ABI).
|
||||||
|
|
||||||
|
## Transport
|
||||||
|
|
||||||
|
A VFS exchange is one synchronous IPC rendezvous (`ipc_call`,
|
||||||
|
docs/ipc.md): the client sends one message and blocks; the server replies
|
||||||
|
with one message. The endpoint is found by well-known service id
|
||||||
|
(`ipc_lookup`, service id **1** = vfs).
|
||||||
|
|
||||||
|
- A message is at most **256 bytes** (`message_maximum`).
|
||||||
|
- A request is a fixed 32-byte **Request** header followed by an inline
|
||||||
|
payload of at most **224 bytes** (`maximum_payload`) — a path, or write
|
||||||
|
bytes. There is no multi-message request: paths and single reads/writes
|
||||||
|
must fit, and larger transfers loop (see *read* / *write*).
|
||||||
|
- A reply is a fixed 24-byte **Reply** header followed by an inline payload —
|
||||||
|
read bytes, a `FileStatus`, or a `DirectoryEntry`.
|
||||||
|
- All integers are **little-endian**; layouts are C layout for x86-64
|
||||||
|
(`extern struct`), offsets given below so nothing need be inferred.
|
||||||
|
|
||||||
|
The kernel never parses any of this — it only moves the bytes
|
||||||
|
(docs/syscall.md); files are entirely a user-space affair.
|
||||||
|
|
||||||
|
## Request header — 32 bytes
|
||||||
|
|
||||||
|
| offset | size | field | meaning |
|
||||||
|
|-------:|-----:|-------|---------|
|
||||||
|
| 0 | 4 | `operation` | an **Operation** value (below) |
|
||||||
|
| 4 | 4 | — | padding |
|
||||||
|
| 8 | 8 | `node` | the server-side open-node id from a prior `open`; 0 for path-based operations |
|
||||||
|
| 16 | 8 | `offset` | byte position for read/write; entry index (cursor) for readdir; else 0 |
|
||||||
|
| 24 | 4 | `len` | payload length for path/write operations; requested byte count for read |
|
||||||
|
| 28 | 4 | `flags` | open flags (below); else 0 |
|
||||||
|
|
||||||
|
## Reply header — 24 bytes
|
||||||
|
|
||||||
|
| offset | size | field | meaning |
|
||||||
|
|-------:|-----:|-------|---------|
|
||||||
|
| 0 | 4 | `status` | **0 = success**, negative = failure (signed) |
|
||||||
|
| 4 | 4 | — | padding |
|
||||||
|
| 8 | 8 | `node` | the new open-node id (for `open`); else 0 |
|
||||||
|
| 16 | 4 | `len` | reply payload length in bytes |
|
||||||
|
| 20 | 4 | — | padding |
|
||||||
|
|
||||||
|
On failure the router replies `status = -1`; a mounted backend's negative
|
||||||
|
status is forwarded to the client verbatim. A richer errno vocabulary is
|
||||||
|
future work — clients must treat *any* negative status as failure, not match
|
||||||
|
on -1.
|
||||||
|
|
||||||
|
## Operations
|
||||||
|
|
||||||
|
Values are append-only and never renumbered (the same evolution rule every
|
||||||
|
danos protocol follows); an unrecognised operation gets a `status = -1`
|
||||||
|
reply.
|
||||||
|
|
||||||
|
| value | operation | request payload | reply |
|
||||||
|
|------:|-----------|-----------------|-------|
|
||||||
|
| 0 | `open` | the path (`len` = its length), `flags` as below | `node` = open-node id |
|
||||||
|
| 1 | `close` | — (`node` set) | status only |
|
||||||
|
| 2 | `read` | — (`node`, `offset`, `len` = wanted count) | `len` bytes read, payload = the bytes; `len` 0 at end of file |
|
||||||
|
| 3 | `write` | the bytes (`node`, `offset`, `len` = count) | `len` = bytes accepted (may be short — loop) |
|
||||||
|
| 4 | `status` | — (`node` set) | payload = **FileStatus** (24 bytes) |
|
||||||
|
| 5 | `readdir` | — (`node` = a directory, `offset` = cursor) | payload = one **DirectoryEntry** + name; `len` 0 at end |
|
||||||
|
| 6 | `mount` | the mount-point path; the backend endpoint rides as the call's **capability** | status only |
|
||||||
|
| 7 | `unmount` | the mount-point path | status only |
|
||||||
|
| 8 | `mkdir` | the path | status only |
|
||||||
|
| 9 | `unlink` | the path | status only |
|
||||||
|
| 10 | `rename` | old path, one `0x00`, new path (`len` = total) | status only |
|
||||||
|
|
||||||
|
Notes per operation:
|
||||||
|
|
||||||
|
- **open** — paths are absolute (`/mnt/usb/notes.txt`) or bare names
|
||||||
|
(`greeting`); bare names resolve in the VFS's flat ramfs, absolute paths
|
||||||
|
route through the mount table (below). The returned `node` is an id in the
|
||||||
|
*router's* open table; clients never see a backend's own ids.
|
||||||
|
- **read / write** — a single exchange moves at most 224 bytes
|
||||||
|
(`maximum_payload`); the client loops, advancing `offset` by the returned
|
||||||
|
`len`, until done (read) or the slice is written (write). A `write` reply
|
||||||
|
shorter than requested is progress, not an error; a `len` of 0 means no
|
||||||
|
forward progress — stop rather than spin.
|
||||||
|
- **readdir** — `offset` is a **cursor: the entry index**, not a byte
|
||||||
|
position. Each call returns exactly one entry; the client increments the
|
||||||
|
cursor by 1. A reply with `len` 0 is end-of-directory. The directory must
|
||||||
|
have been opened with the `directory` flag.
|
||||||
|
- **mount** — the one operation that passes a **capability**: the caller
|
||||||
|
(a filesystem server, e.g. FAT) sends its own request endpoint as the
|
||||||
|
`ipc_call` capability argument, and the router forwards everything under
|
||||||
|
the mount point to it — speaking this same protocol, with paths rewritten
|
||||||
|
relative to the mount. Prefixes match at path boundaries only
|
||||||
|
(`/mnt/usb` never captures `/mnt/usbextra`); the longest matching prefix
|
||||||
|
wins.
|
||||||
|
- **rename** — same-directory rename only (the router requires old and new to
|
||||||
|
resolve under one mount).
|
||||||
|
|
||||||
|
## Open flags
|
||||||
|
|
||||||
|
Bitwise OR in `Request.flags`, meaningful for `open` only:
|
||||||
|
|
||||||
|
| bit | name | meaning |
|
||||||
|
|----:|------|---------|
|
||||||
|
| 1 | `create` | create the file if it does not exist |
|
||||||
|
| 2 | `directory` | open a directory node for `readdir` rather than a file |
|
||||||
|
| 4 | `truncate` | truncate an existing file to zero length on open (replace, don't overwrite in place) |
|
||||||
|
|
||||||
|
## FileStatus — 24 bytes (the `status` reply payload)
|
||||||
|
|
||||||
|
| offset | size | field | meaning |
|
||||||
|
|-------:|-----:|-------|---------|
|
||||||
|
| 0 | 8 | `size` | file size in bytes |
|
||||||
|
| 8 | 4 | `kind` | a **NodeKind** value |
|
||||||
|
| 12 | 4 | — | padding |
|
||||||
|
| 16 | 8 | `mtime` | modification time, Unix epoch seconds UTC; 0 if the backend keeps none |
|
||||||
|
|
||||||
|
## DirectoryEntry — 16 bytes + name (the `readdir` reply payload)
|
||||||
|
|
||||||
|
| offset | size | field | meaning |
|
||||||
|
|-------:|-----:|-------|---------|
|
||||||
|
| 0 | 4 | `kind` | a **NodeKind** value |
|
||||||
|
| 4 | 4 | `name_len` | length of the name that follows |
|
||||||
|
| 8 | 8 | `size` | the entry's size in bytes |
|
||||||
|
| 16 | `name_len` | name | the entry's name, not NUL-terminated |
|
||||||
|
|
||||||
|
## NodeKind
|
||||||
|
|
||||||
|
Aligned to the FSH file-type table
|
||||||
|
(docs/danos-file-system-hierarchy-FSH.md):
|
||||||
|
|
||||||
|
| value | kind |
|
||||||
|
|------:|------|
|
||||||
|
| 0 | regular file |
|
||||||
|
| 1 | directory |
|
||||||
|
| 2 | character device |
|
||||||
|
| 3 | block device |
|
||||||
|
| 4 | symbolic link |
|
||||||
|
| 5 | fifo |
|
||||||
|
| 6 | socket |
|
||||||
|
|
||||||
|
Clients should map unknown values to *regular* rather than reject — the
|
||||||
|
table can grow.
|
||||||
|
|
||||||
|
## Lifetimes and trust
|
||||||
|
|
||||||
|
Open-node ids live in the server. A client that dies without closing leaks
|
||||||
|
nothing permanently: the VFS subscribes to the kernel's published process-exit
|
||||||
|
events (docs/process-lifecycle.md) and releases a dead client's handles,
|
||||||
|
closing forwarded backend nodes best-effort. Ids are plain integers, not
|
||||||
|
capabilities — the VFS trusts its callers with each other's ids today, which
|
||||||
|
is acceptable while every client is part of the system image and worth
|
||||||
|
revisiting (per-client id namespaces) before third-party binaries arrive.
|
||||||
|
|
||||||
|
## Evolution rules
|
||||||
|
|
||||||
|
What a non-Zig implementation may rely on, and what it must not:
|
||||||
|
|
||||||
|
- Operation values, flag bits, `NodeKind` values, and struct layouts are
|
||||||
|
**append-only and frozen once shipped** — the unit tests in `protocol.zig`
|
||||||
|
pin them exactly so a refactor can't silently move them.
|
||||||
|
- The 256-byte message ceiling is a property of the current IPC transport,
|
||||||
|
not a promise; clients should read `maximum_payload`-shaped limits from the
|
||||||
|
reply lengths they actually get (loop-until-done), not hard-code 224.
|
||||||
|
- Negative statuses beyond -1 will appear (an errno vocabulary); success is
|
||||||
|
exactly 0.
|
||||||
Reference in New Issue
Block a user