docs: vdso + vfs-protocol — the public ABI boundary, linked from the index
This commit is contained in:
+20
-9
@@ -39,37 +39,42 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
endpoints — the backbone the microkernel's isolated servers talk over.
|
||||
12. **[syscall.md](syscall.md) — system calls.** How ring 3 asks the kernel for
|
||||
something: the `syscall`/`sysret` fast path, the trap frame, and why the table is
|
||||
deliberately tiny.
|
||||
13. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an
|
||||
deliberately tiny. The numbers are a **private** ABI — [vdso.md](vdso.md) designs
|
||||
the public boundary that will hide them.
|
||||
13. **[vfs-protocol.md](vfs-protocol.md) — the VFS wire protocol.** The language-neutral
|
||||
byte-level spec of the file protocol spoken over IPC: request/reply headers,
|
||||
the operation table, mount routing, and the append-only evolution rules — the
|
||||
first IPC protocol documented as public ABI.
|
||||
14. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an
|
||||
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
||||
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
||||
unmask.
|
||||
14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
|
||||
15. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
|
||||
real driver stacks factor into three shapes and how families share code. The
|
||||
three primitives it proposed are long since built (M13 capability passing,
|
||||
M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them —
|
||||
hello, supervision, restart — is built too (device-manager.md, M18).
|
||||
15. **[process-management.md](process-management.md) — process management.** The
|
||||
16. **[process-management.md](process-management.md) — process management.** The
|
||||
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
|
||||
supervision link as the kill authority, and child-exit notifications over the
|
||||
same endpoints IRQs arrive on.
|
||||
16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
|
||||
17. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
|
||||
(M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
|
||||
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
|
||||
`runtime.process` interface, exit reasons, published exit events any stateful
|
||||
service can subscribe to (the VFS releasing dead clients' handles), and the two
|
||||
iron rules (cleanup is the kernel's job; kill is not a signal).
|
||||
17. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
|
||||
18. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
|
||||
through the app surface): the
|
||||
tree, the matcher, and the supervisor. Tree structure lives in the manager,
|
||||
authority stays in the kernel; bus drivers report what they see; drivers are
|
||||
restarted through the lifecycle vocabulary — the plan that turns
|
||||
[resilience.md](resilience.md)'s restart goal into increments.
|
||||
18. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard,
|
||||
19. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard,
|
||||
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
|
||||
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
|
||||
service layered on top.
|
||||
19. **[display.md](display.md) — the display service.** The display half of the GUI
|
||||
20. **[display.md](display.md) — the display service.** The display half of the GUI
|
||||
track: a user-space compositor that owns the framebuffer, composes a layer stack into
|
||||
a double buffer, and presents it. Why GOP and the PCI display device are two views of
|
||||
one controller, the device-node + write-combining handoff, and what flicker-free buys
|
||||
@@ -80,7 +85,7 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
further out, two research snapshots survey what a *native* driver for real GPU silicon
|
||||
would take as another `.scanout` backend: [nvidia-gpus.md](nvidia-gpus.md) (RTX 3060 /
|
||||
Ampere) and [intel-igpu.md](intel-igpu.md) (Intel iGPU).
|
||||
20. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
21. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||
|
||||
Start with the north star:
|
||||
@@ -99,6 +104,12 @@ Start with the north star:
|
||||
port to **one seam** (`std.os.danos`), so we build `runtime.os` (→ that seam) plus a
|
||||
thin `runtime.fs`, retire the `posix` shim, and follow a phased path to
|
||||
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
|
||||
- **[vdso.md](vdso.md) — the vDSO, the public system-call boundary.** A design note
|
||||
(not built yet) on keeping `abi.zig` genuinely private: a kernel-supplied, C-ABI
|
||||
entry blob mapped into every process as the *only* way into the kernel — so the
|
||||
syscall numbers can be renumbered or randomised at will, and Rust/C binaries get a
|
||||
stable boundary without danos growing a dynamic linker. danos's public ABI = the
|
||||
vDSO + the documented IPC wire protocols ([vfs-protocol.md](vfs-protocol.md) first).
|
||||
|
||||
Cutting across all of these:
|
||||
|
||||
|
||||
+206
@@ -0,0 +1,206 @@
|
||||
# The vDSO — the public system-call boundary
|
||||
|
||||
> **Status:** design note, not built. The runtime today issues raw `syscall`
|
||||
> instructions from `library/runtime/system-call.zig` using the numbers in
|
||||
> `system/abi.zig`. This note designs the layer that replaces that arrangement:
|
||||
> a **kernel-supplied, C-ABI entry library** mapped into every process — the
|
||||
> only supported way into the kernel — so the raw numbers can stay private,
|
||||
> be renumbered at will, and eventually be randomised per boot.
|
||||
|
||||
## Why: the ABI danos promises, and the one it doesn't
|
||||
|
||||
`system/abi.zig` is the **private** kernel ↔ runtime contract. Its header says
|
||||
so: the numbers are an implementation detail the runtime hides and may
|
||||
renumber, the same split as libSystem over the XNU syscalls on macOS or win32
|
||||
over the NT syscalls on Windows. Linux — with its world-visible, frozen
|
||||
syscall table — is the outlier, not the norm.
|
||||
|
||||
That stance has consequences the moment binaries exist that we don't rebuild
|
||||
ourselves:
|
||||
|
||||
1. **Third-party binaries** (docs/zig-self-hosting.md) must keep working across
|
||||
kernel updates. If they contain raw `syscall` instructions with today's
|
||||
numbers baked in, every renumbering breaks the world — the ABI would be
|
||||
*de facto* public no matter what the header says. Go on macOS made exactly
|
||||
this mistake: it issued XNU syscalls directly instead of going through
|
||||
libSystem, and macOS updates repeatedly broke every Go binary until Go
|
||||
switched to the library like everyone else.
|
||||
2. **Not everything is Zig.** A Rust or C program can't import the `runtime`
|
||||
module. The public boundary has to be expressible in the one calling
|
||||
convention every language speaks: the C ABI.
|
||||
3. **Randomised syscall numbers** — a hardening option we want open — only
|
||||
work if no user binary anywhere knows a number at build time. The binding
|
||||
must happen at *load time*, from something the kernel controls.
|
||||
|
||||
All three point at the same well-known shape: a **vDSO** (virtual dynamic
|
||||
shared object). The kernel carries a small blob of user-mode code, maps it
|
||||
into every process at spawn, and that blob — not the application — contains
|
||||
the `syscall` instructions. Fuchsia works exactly this way: its vDSO is the
|
||||
*only* kernel entry, version-matched by construction because the kernel itself
|
||||
injects it. Because the kernel and the blob ship as one artifact, there is
|
||||
**no version skew, no loader, no search path, and no shared file on disk** —
|
||||
which is what makes this the resilient way to have a private ABI
|
||||
(docs/resilience.md), where a conventional `ld.so` + `/lib/libdanos.so`
|
||||
arrangement would add a loader to every spawn and a single shared point of
|
||||
failure.
|
||||
|
||||
The public danos ABI then has exactly two layers, neither of which is
|
||||
`abi.zig`:
|
||||
|
||||
| Layer | Contract | Spoken by |
|
||||
|-------|----------|-----------|
|
||||
| **vDSO** | C-ABI functions, this note | every language's thin shim (`runtime.system` for Zig, a `-sys` crate for Rust, a header for C) |
|
||||
| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](vfs-protocol.md) is the first one documented) | any client that can lay out bytes |
|
||||
|
||||
Everything above those — the heap, `runtime.fs`, the service harness — is
|
||||
per-language convenience, compiled into each binary from source, exactly as
|
||||
today. Nothing about the Zig runtime's shape changes; it just stops being the
|
||||
*only* door.
|
||||
|
||||
## The blob
|
||||
|
||||
A single copy of the vDSO code lives in the kernel image (built by
|
||||
`build.zig` as a tiny freestanding object, embedded like the AP trampoline).
|
||||
At boot the kernel finalises it once — this is where randomised numbers would
|
||||
be patched in — and thereafter maps the **same physical pages** read-execute
|
||||
into every process's address space. The blob is:
|
||||
|
||||
- **Position-independent.** It is mapped at a per-process randomised base, so
|
||||
it must be PIC (rip-relative addressing only — no relocations to process).
|
||||
- **Stateless and re-entrant.** No writable data. Anything stateful belongs to
|
||||
the process, not the vDSO.
|
||||
- **Architecture-specific.** The x86-64 blob wraps `syscall`; an aarch64 blob
|
||||
wraps `svc #0`. It lives beside the other per-architecture kernel sources
|
||||
(`system/kernel/architecture/<arch>/`), selected the same way the
|
||||
`architecture` module is (docs/arch.md).
|
||||
|
||||
### Shape: a function table, not an ELF
|
||||
|
||||
A real `.so` with a dynamic symbol table is the conventional vDSO shape, but
|
||||
linking against one at load time needs a dynamic linker in every binary —
|
||||
machinery danos deliberately doesn't have. Instead the v1 shape is the
|
||||
simplest thing that is still a stable contract — a **function-pointer table**
|
||||
at the vDSO base:
|
||||
|
||||
```
|
||||
offset 0 u64 magic 'danosVDS' — a mapped-the-wrong-thing guard
|
||||
offset 8 u64 api_level incremented when the table grows
|
||||
offset 16 u64 count number of table entries that follow
|
||||
offset 24 u64 table[count] function pointers into the vDSO's own code
|
||||
```
|
||||
|
||||
Table *indices* are the public constants (published in a C header,
|
||||
`danos.h`), assigned once and append-only — the same discipline the IPC
|
||||
protocols use for operation values. The pointers point at stubs inside the
|
||||
blob; what those stubs put in `rax` is nobody's business but the kernel's.
|
||||
A language shim binds in one step: read the base from the init block, check
|
||||
the magic, keep the table pointer. Feature detection for a binary built
|
||||
against older headers is `count`/`api_level` — a kernel never removes or
|
||||
reorders entries.
|
||||
|
||||
(If danos ever grows a real dynamic linker, the same blob can additionally
|
||||
present an ELF `dynsym` without breaking the table — Fuchsia's vDSO is
|
||||
likewise both a mappable blob and a linkable `.so`. That is a later
|
||||
convenience, not a requirement.)
|
||||
|
||||
### Delivery: the auxiliary vector
|
||||
|
||||
The kernel already builds a System V entry block — argc, argv, envp
|
||||
terminator, **auxiliary vector** — on every new process's stack
|
||||
(`buildEntryStack`, read by `runtime.start`). The vDSO base rides in a new
|
||||
auxv entry, exactly Linux's `AT_SYSINFO_EHDR` move. No new syscall, no magic
|
||||
address, and a language shim finds it the same portable way on every
|
||||
architecture.
|
||||
|
||||
## The function surface
|
||||
|
||||
One table entry per kernel call, C ABI (System V AMD64), names prefixed
|
||||
`danos_`. The current `SystemCall` set maps directly; integer arguments and
|
||||
returns are `u64`, errors return as negative values exactly as today.
|
||||
|
||||
The calls that return two values in `rax:rdx` today — `dma_alloc`
|
||||
(vaddr + paddr), `msi_bind` (address + data), `shm_create` (vaddr + handle) —
|
||||
become functions returning a two-`u64` struct. The System V ABI returns a
|
||||
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
|
||||
C-ABI spelling of the existing convention, at zero cost.
|
||||
|
||||
Grouped as `abi.zig` groups them:
|
||||
|
||||
| Group | Functions |
|
||||
|-------|-----------|
|
||||
| process | `danos_exit`, `danos_yield`, `danos_sleep`, `danos_spawn`, `danos_process_enumerate`, `danos_process_kill`, `danos_process_exit_reason`, `danos_process_subscribe`, `danos_process_signal`, `danos_signal_bind` |
|
||||
| memory | `danos_mmap`, `danos_munmap`, `danos_dma_alloc`, `danos_dma_free`, `danos_shm_create`, `danos_shm_map`, `danos_shm_physical` |
|
||||
| ipc | `danos_endpoint_create`, `danos_ipc_register`, `danos_ipc_lookup`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` |
|
||||
| devices | `danos_device_enumerate`, `danos_device_claim`, `danos_device_register`, `danos_mmio_map`, `danos_irq_bind`, `danos_irq_ack`, `danos_msi_bind`, `danos_io_read`, `danos_io_write` |
|
||||
| time | `danos_clock`, `danos_wall_clock`, `danos_timer_bind` |
|
||||
| diagnostics | `danos_debug_write`, `danos_klog_read` |
|
||||
|
||||
The constants that ride alongside the calls — mmap protection bits, DMA
|
||||
flags, notification badge bits, `ExitReason`, `Signal`, well-known service
|
||||
ids, `page_size`, the IPC message maximum — move to the public header too:
|
||||
they are wire values a Rust program needs verbatim. What stays private in
|
||||
`abi.zig` is exactly the thing the vDSO exists to hide: the `SystemCall`
|
||||
numbers and the trap convention.
|
||||
|
||||
## Enforcement, and an honest threat model
|
||||
|
||||
Renumbering only has teeth if the kernel **refuses syscalls that don't come
|
||||
from the vDSO**. The check is cheap: on kernel entry, the saved user `rip`
|
||||
must lie inside the calling process's vDSO mapping; otherwise the process is
|
||||
killed with a fault-class exit reason (its supervisor restarts or gives up,
|
||||
docs/process-lifecycle.md — a foreign-syscall attempt is a bug or an attack,
|
||||
never something to limp past). Fuchsia enforces exactly this.
|
||||
|
||||
What this buys, precisely:
|
||||
|
||||
- **ABI freedom** — the real prize. The numbers can change per release or per
|
||||
boot and nothing outside the kernel image cares. The private ABI stays
|
||||
actually private, permanently.
|
||||
- **A single audited chokepoint** for kernel entry, per process, at a
|
||||
randomised address.
|
||||
- **Raised bar for exploits**: shellcode can't issue a hard-coded `syscall`;
|
||||
it must first discover the per-process vDSO base (ASLR) and call through
|
||||
it.
|
||||
|
||||
What it does *not* buy: an attacker with arbitrary code execution in a
|
||||
process can still *call* the vDSO functions — they are mapped executable in
|
||||
that process, and return-oriented chains reach them. Syscall randomisation is
|
||||
hardening, not a security boundary; the security boundary remains the
|
||||
capability model (what the process's endpoints and device claims let it do).
|
||||
It is worth building anyway — for the ABI freedom first and the hardening
|
||||
second — but the design should never be sold as more than that.
|
||||
|
||||
## Migration
|
||||
|
||||
Phased so every step ships alone (the M-milestone discipline):
|
||||
|
||||
1. **The blob + the table.** Build the vDSO, map it at spawn, deliver the
|
||||
base via auxv. `runtime.system-call.zig` binds through the table when the
|
||||
auxv entry is present, falls back to raw `syscall` when absent — the whole
|
||||
tree keeps booting during the transition.
|
||||
2. **Cut the runtime over.** Delete the raw stubs; `runtime` no longer
|
||||
imports the `SystemCall` numbers at all (`abi.zig`'s enum becomes
|
||||
kernel-internal). The QEMU suite passing proves the table carries the
|
||||
whole system.
|
||||
3. **Enforce + randomise.** Add the `rip`-range check, then per-boot number
|
||||
randomisation patched into the blob at kernel init. A test boots with
|
||||
randomisation on and runs the full suite.
|
||||
4. **The other languages.** Publish `danos.h`; a Rust `danos-sys` crate wraps
|
||||
the table. This is also the seam `std.os.danos` calls through when the Zig
|
||||
self-hosting fork lands (docs/zig-self-hosting.md) — the vDSO is what
|
||||
makes that seam stable across kernel versions.
|
||||
|
||||
## What deliberately stays out
|
||||
|
||||
- **No dynamic linker, no `/lib/*.so`.** The vDSO is kernel-injected precisely
|
||||
so danos binaries can stay fully static above it. Sharing *library code*
|
||||
across processes stays what it is today: a service behind IPC, or source
|
||||
compiled into each binary.
|
||||
- **No file/device I/O in the vDSO.** The microkernel line doesn't move: the
|
||||
vDSO wraps the same deliberately tiny table (docs/syscall.md); files are
|
||||
still the VFS server's business over IPC.
|
||||
- **No fast-path user-mode implementations yet.** Linux's vDSO exists mostly
|
||||
to answer `gettimeofday` without a kernel entry. `danos_clock` could one
|
||||
day read the calibrated TSC in user mode the same way — the blob is where
|
||||
such an optimisation would live — but that is an optimisation, not part of
|
||||
this design's contract.
|
||||
@@ -0,0 +1,170 @@
|
||||
# The VFS wire protocol
|
||||
|
||||
> **Status:** built and spoken today between `runtime.fs` (the client) and the
|
||||
> VFS server (`system/services/vfs`), with mounted backends (the FAT server)
|
||||
> speaking the same protocol behind the router. The Zig source of truth is
|
||||
> `system/services/vfs/protocol.zig` (the `vfs-protocol` module), whose unit
|
||||
> tests pin the sizes and values below. This page is the **language-neutral
|
||||
> wire specification** of that contract — what a Rust or C client implements
|
||||
> ([vdso.md](vdso.md) explains why the IPC protocols, not the syscall
|
||||
> numbers, are danos's public ABI).
|
||||
|
||||
## Transport
|
||||
|
||||
A VFS exchange is one synchronous IPC rendezvous (`ipc_call`,
|
||||
docs/ipc.md): the client sends one message and blocks; the server replies
|
||||
with one message. The endpoint is found by well-known service id
|
||||
(`ipc_lookup`, service id **1** = vfs).
|
||||
|
||||
- A message is at most **256 bytes** (`message_maximum`).
|
||||
- A request is a fixed 32-byte **Request** header followed by an inline
|
||||
payload of at most **224 bytes** (`maximum_payload`) — a path, or write
|
||||
bytes. There is no multi-message request: paths and single reads/writes
|
||||
must fit, and larger transfers loop (see *read* / *write*).
|
||||
- A reply is a fixed 24-byte **Reply** header followed by an inline payload —
|
||||
read bytes, a `FileStatus`, or a `DirectoryEntry`.
|
||||
- All integers are **little-endian**; layouts are C layout for x86-64
|
||||
(`extern struct`), offsets given below so nothing need be inferred.
|
||||
|
||||
The kernel never parses any of this — it only moves the bytes
|
||||
(docs/syscall.md); files are entirely a user-space affair.
|
||||
|
||||
## Request header — 32 bytes
|
||||
|
||||
| offset | size | field | meaning |
|
||||
|-------:|-----:|-------|---------|
|
||||
| 0 | 4 | `operation` | an **Operation** value (below) |
|
||||
| 4 | 4 | — | padding |
|
||||
| 8 | 8 | `node` | the server-side open-node id from a prior `open`; 0 for path-based operations |
|
||||
| 16 | 8 | `offset` | byte position for read/write; entry index (cursor) for readdir; else 0 |
|
||||
| 24 | 4 | `len` | payload length for path/write operations; requested byte count for read |
|
||||
| 28 | 4 | `flags` | open flags (below); else 0 |
|
||||
|
||||
## Reply header — 24 bytes
|
||||
|
||||
| offset | size | field | meaning |
|
||||
|-------:|-----:|-------|---------|
|
||||
| 0 | 4 | `status` | **0 = success**, negative = failure (signed) |
|
||||
| 4 | 4 | — | padding |
|
||||
| 8 | 8 | `node` | the new open-node id (for `open`); else 0 |
|
||||
| 16 | 4 | `len` | reply payload length in bytes |
|
||||
| 20 | 4 | — | padding |
|
||||
|
||||
On failure the router replies `status = -1`; a mounted backend's negative
|
||||
status is forwarded to the client verbatim. A richer errno vocabulary is
|
||||
future work — clients must treat *any* negative status as failure, not match
|
||||
on -1.
|
||||
|
||||
## Operations
|
||||
|
||||
Values are append-only and never renumbered (the same evolution rule every
|
||||
danos protocol follows); an unrecognised operation gets a `status = -1`
|
||||
reply.
|
||||
|
||||
| value | operation | request payload | reply |
|
||||
|------:|-----------|-----------------|-------|
|
||||
| 0 | `open` | the path (`len` = its length), `flags` as below | `node` = open-node id |
|
||||
| 1 | `close` | — (`node` set) | status only |
|
||||
| 2 | `read` | — (`node`, `offset`, `len` = wanted count) | `len` bytes read, payload = the bytes; `len` 0 at end of file |
|
||||
| 3 | `write` | the bytes (`node`, `offset`, `len` = count) | `len` = bytes accepted (may be short — loop) |
|
||||
| 4 | `status` | — (`node` set) | payload = **FileStatus** (24 bytes) |
|
||||
| 5 | `readdir` | — (`node` = a directory, `offset` = cursor) | payload = one **DirectoryEntry** + name; `len` 0 at end |
|
||||
| 6 | `mount` | the mount-point path; the backend endpoint rides as the call's **capability** | status only |
|
||||
| 7 | `unmount` | the mount-point path | status only |
|
||||
| 8 | `mkdir` | the path | status only |
|
||||
| 9 | `unlink` | the path | status only |
|
||||
| 10 | `rename` | old path, one `0x00`, new path (`len` = total) | status only |
|
||||
|
||||
Notes per operation:
|
||||
|
||||
- **open** — paths are absolute (`/mnt/usb/notes.txt`) or bare names
|
||||
(`greeting`); bare names resolve in the VFS's flat ramfs, absolute paths
|
||||
route through the mount table (below). The returned `node` is an id in the
|
||||
*router's* open table; clients never see a backend's own ids.
|
||||
- **read / write** — a single exchange moves at most 224 bytes
|
||||
(`maximum_payload`); the client loops, advancing `offset` by the returned
|
||||
`len`, until done (read) or the slice is written (write). A `write` reply
|
||||
shorter than requested is progress, not an error; a `len` of 0 means no
|
||||
forward progress — stop rather than spin.
|
||||
- **readdir** — `offset` is a **cursor: the entry index**, not a byte
|
||||
position. Each call returns exactly one entry; the client increments the
|
||||
cursor by 1. A reply with `len` 0 is end-of-directory. The directory must
|
||||
have been opened with the `directory` flag.
|
||||
- **mount** — the one operation that passes a **capability**: the caller
|
||||
(a filesystem server, e.g. FAT) sends its own request endpoint as the
|
||||
`ipc_call` capability argument, and the router forwards everything under
|
||||
the mount point to it — speaking this same protocol, with paths rewritten
|
||||
relative to the mount. Prefixes match at path boundaries only
|
||||
(`/mnt/usb` never captures `/mnt/usbextra`); the longest matching prefix
|
||||
wins.
|
||||
- **rename** — same-directory rename only (the router requires old and new to
|
||||
resolve under one mount).
|
||||
|
||||
## Open flags
|
||||
|
||||
Bitwise OR in `Request.flags`, meaningful for `open` only:
|
||||
|
||||
| bit | name | meaning |
|
||||
|----:|------|---------|
|
||||
| 1 | `create` | create the file if it does not exist |
|
||||
| 2 | `directory` | open a directory node for `readdir` rather than a file |
|
||||
| 4 | `truncate` | truncate an existing file to zero length on open (replace, don't overwrite in place) |
|
||||
|
||||
## FileStatus — 24 bytes (the `status` reply payload)
|
||||
|
||||
| offset | size | field | meaning |
|
||||
|-------:|-----:|-------|---------|
|
||||
| 0 | 8 | `size` | file size in bytes |
|
||||
| 8 | 4 | `kind` | a **NodeKind** value |
|
||||
| 12 | 4 | — | padding |
|
||||
| 16 | 8 | `mtime` | modification time, Unix epoch seconds UTC; 0 if the backend keeps none |
|
||||
|
||||
## DirectoryEntry — 16 bytes + name (the `readdir` reply payload)
|
||||
|
||||
| offset | size | field | meaning |
|
||||
|-------:|-----:|-------|---------|
|
||||
| 0 | 4 | `kind` | a **NodeKind** value |
|
||||
| 4 | 4 | `name_len` | length of the name that follows |
|
||||
| 8 | 8 | `size` | the entry's size in bytes |
|
||||
| 16 | `name_len` | name | the entry's name, not NUL-terminated |
|
||||
|
||||
## NodeKind
|
||||
|
||||
Aligned to the FSH file-type table
|
||||
(docs/danos-file-system-hierarchy-FSH.md):
|
||||
|
||||
| value | kind |
|
||||
|------:|------|
|
||||
| 0 | regular file |
|
||||
| 1 | directory |
|
||||
| 2 | character device |
|
||||
| 3 | block device |
|
||||
| 4 | symbolic link |
|
||||
| 5 | fifo |
|
||||
| 6 | socket |
|
||||
|
||||
Clients should map unknown values to *regular* rather than reject — the
|
||||
table can grow.
|
||||
|
||||
## Lifetimes and trust
|
||||
|
||||
Open-node ids live in the server. A client that dies without closing leaks
|
||||
nothing permanently: the VFS subscribes to the kernel's published process-exit
|
||||
events (docs/process-lifecycle.md) and releases a dead client's handles,
|
||||
closing forwarded backend nodes best-effort. Ids are plain integers, not
|
||||
capabilities — the VFS trusts its callers with each other's ids today, which
|
||||
is acceptable while every client is part of the system image and worth
|
||||
revisiting (per-client id namespaces) before third-party binaries arrive.
|
||||
|
||||
## Evolution rules
|
||||
|
||||
What a non-Zig implementation may rely on, and what it must not:
|
||||
|
||||
- Operation values, flag bits, `NodeKind` values, and struct layouts are
|
||||
**append-only and frozen once shipped** — the unit tests in `protocol.zig`
|
||||
pin them exactly so a refactor can't silently move them.
|
||||
- The 256-byte message ceiling is a property of the current IPC transport,
|
||||
not a promise; clients should read `maximum_payload`-shaped limits from the
|
||||
reply lengths they actually get (loop-until-done), not hard-code 224.
|
||||
- Negative statuses beyond -1 will appear (an errno vocabulary); success is
|
||||
exactly 0.
|
||||
Reference in New Issue
Block a user