library: five protocols speak the envelope
The folded header stops being a rule in a document and becomes the layout on the wire. Verbs number from sixteen, leaving describe, enumerate, subscribe and unsubscribe reserved and answered the same way by every provider — none of them writes a line to do it. What each protocol used to carry in a field of its own now travels in the header: a vfs node and a display layer are the packet's target, and a reply opens with a status the envelope stamps rather than one each protocol spelled for itself. Display gains the most. One forty-byte request had served eleven verbs, so attach_scanout smuggled stride through x, refresh through y and format through colour, and every coordinate crossed as a bitcast. Per-operation structs end all three: the fields have their own names and their own signs, and the tile payload grows to 224 bytes because the prefix shrank. Scanout loses a message maximum of 64 it had no business declaring — it answers calls, and the floor for a call is 256 — and virtio-gpu stops hard-coding that number at its harness. Two changes are semantic rather than notational. A directory now ends at an entry with no name, because the fixed part of a reply always travels and a zero-length reply no longer exists to mean anything. And input joins the service harness, the last loop in the tree that answered no ping and heard no terminate; its subscriber table, its pruning and its fan-out are the same code, and a shutdown now asks it to stop instead of killing it. A new conformance case reads the registry's own listing and asks every protocol it finds for its name, its version and its verb count, then offers a verb nobody defines and requires -ENOSYS — the envelope's promise, checked against providers rather than against itself. What it cannot reach in that boot it names on the serial line instead of passing quietly. Suite 110/110.
This commit is contained in:
@@ -6,7 +6,10 @@
|
||||
> serves it directly (the read-only /system initrd mount, via `fs_node`) or
|
||||
> redirects the caller to the owning backend's endpoint plus the rewritten
|
||||
> mount-relative path — after which the client speaks THIS protocol to the
|
||||
> backend, unchanged. The Zig source of truth is `library/protocol/vfs/vfs-protocol.zig`
|
||||
> backend, unchanged. Since P4a the contract is expressed through
|
||||
> `envelope.Define` (docs/os-development/protocol-namespace.md), so every
|
||||
> packet begins with the universal 16-byte prefix and the open-node id rides
|
||||
> in it. The Zig source of truth is `library/protocol/vfs/vfs-protocol.zig`
|
||||
> (the `vfs-protocol` module), whose unit test pins a sample of the sizes
|
||||
> and values below. This page is the **language-neutral wire specification**
|
||||
> of that contract — what a Rust or C client implements ([vdso.md](../os-development/vdso.md)
|
||||
@@ -22,12 +25,17 @@ also hands back the path rewritten relative to the mount — not from a
|
||||
registry lookup. (Service id 1, the old userspace router, is retired.)
|
||||
|
||||
- A message is at most **256 bytes** (`message_maximum`).
|
||||
- A request is a fixed 32-byte **Request** header followed by an inline
|
||||
payload of at most **224 bytes** (`maximum_payload`) — a path, or write
|
||||
bytes. There is no multi-message request: paths and single reads/writes
|
||||
must fit, and larger transfers loop (see *read* / *write*).
|
||||
- A reply is a fixed 24-byte **Reply** header followed by an inline payload —
|
||||
read bytes, a `FileStatus`, or a `DirectoryEntry`.
|
||||
- Every packet begins with the 16-byte **envelope prefix**
|
||||
([protocol-namespace.md](../os-development/protocol-namespace.md)): a
|
||||
`Header` on a request, a `Status` on a reply. The prefix is **folded, not
|
||||
stacked** — the verb and the object being addressed live in it, and no
|
||||
request or reply below repeats either.
|
||||
- A request is the header, then the verb's own fixed part (0–16 bytes), then
|
||||
an inline tail of at most **224 bytes** (`maximum_payload`) — a path, or
|
||||
write bytes. There is no multi-message request: paths and single
|
||||
reads/writes must fit, and larger transfers loop (see *read* / *write*).
|
||||
- A reply is the status, then the verb's own fixed part, then an inline tail
|
||||
— read bytes, or a directory entry's name.
|
||||
- All integers are **little-endian**; layouts are C layout for x86-64
|
||||
(`extern struct`), offsets given below so nothing need be inferred.
|
||||
|
||||
@@ -37,91 +45,108 @@ With clients holding backend node ids directly, a backend records each open
|
||||
handle's owner and sweeps a dead client's handles via the published process
|
||||
exit events.
|
||||
|
||||
## Request header — 32 bytes
|
||||
## Request header — 16 bytes
|
||||
|
||||
The envelope's `Header`, identical in every danos protocol:
|
||||
|
||||
| offset | size | field | meaning |
|
||||
|-------:|-----:|-------|---------|
|
||||
| 0 | 4 | `operation` | an **Operation** value (below) |
|
||||
| 0 | 4 | `operation` | an **Operation** value (below); 0–15 are the reserved universal verbs |
|
||||
| 4 | 4 | — | padding |
|
||||
| 8 | 8 | `node` | the server-side open-node id from a prior `open`; 0 for path-based operations |
|
||||
| 16 | 8 | `offset` | byte position for read/write; entry index (cursor) for readdir; else 0 |
|
||||
| 24 | 4 | `len` | payload length for path/write operations; requested byte count for read |
|
||||
| 28 | 4 | `flags` | open flags (below); else 0 |
|
||||
| 8 | 8 | `target` | **the open-node id** from a prior `open`; 0 for `open` itself and the path-based verbs |
|
||||
|
||||
## Reply header — 24 bytes
|
||||
## Reply header — 16 bytes
|
||||
|
||||
The envelope's `Status`:
|
||||
|
||||
| offset | size | field | meaning |
|
||||
|-------:|-----:|-------|---------|
|
||||
| 0 | 4 | `status` | **0 = success**, negative = failure (signed) |
|
||||
| 4 | 4 | — | padding |
|
||||
| 8 | 8 | `node` | the new open-node id (for `open`); else 0 |
|
||||
| 16 | 4 | `len` | reply payload length in bytes |
|
||||
| 20 | 4 | — | padding |
|
||||
| 8 | 4 | `len` | reply bytes following this header: the verb's fixed part plus its tail |
|
||||
| 12 | 4 | — | padding |
|
||||
|
||||
On failure the backend replies `status = -1`, and that reply reaches the
|
||||
client directly — there is no party between them on the wire. (Kernel-served
|
||||
paths produce no wire replies at all: `fs_resolve`/`fs_node` failures are
|
||||
syscall register statuses.) A richer errno vocabulary is future work —
|
||||
clients must treat *any* negative status as failure, not match on -1.
|
||||
A failing backend replies with the status alone (`len` = 0) and no fixed
|
||||
part, and that reply reaches the client directly — there is no party between
|
||||
them on the wire. (Kernel-served paths produce no wire replies at all:
|
||||
`fs_resolve`/`fs_node` failures are syscall register statuses.) The errno
|
||||
vocabulary is the kernel's, continued by the envelope: `ENOENT` = 4 is what a
|
||||
backend answers for anything it cannot find or cannot do, `ENOSYS` = 10 for a
|
||||
verb it does not implement, `EPROTO` = 11 for a packet shorter than the verb
|
||||
it names. Clients must treat *any* negative status as failure rather than
|
||||
matching a particular one.
|
||||
|
||||
## Operations
|
||||
|
||||
Values are append-only and never renumbered (the same evolution rule every
|
||||
danos protocol follows). Send only values from this table: the shipped server
|
||||
decodes the operation into an exhaustive enum, so an out-of-range value is
|
||||
not answered with a `status = -1` reply — it trips a safety check in safe
|
||||
builds and is undefined otherwise. (The `-1` replies cover recognised but
|
||||
refused operations, such as `mount` sent to a backend.)
|
||||
Values number from 16 (`first_protocol_operation`) in declaration order, and
|
||||
are frozen once shipped. Values 0–15 are the envelope's reserved universal
|
||||
verbs, which mean the same thing at every provider in the system: `describe`
|
||||
(0) answers the protocol's name and version and is implemented by the
|
||||
envelope itself, so every backend answers it. A verb outside this table is
|
||||
answered `-ENOSYS`; it is never a safety check any more, because the
|
||||
dispatch compares numbers rather than decoding an enum.
|
||||
|
||||
| value | operation | request payload | reply |
|
||||
|------:|-----------|-----------------|-------|
|
||||
| 0 | `open` | the path (`len` = its length), `flags` as below | `node` = open-node id |
|
||||
| 1 | `close` | — (`node` set) | status only |
|
||||
| 2 | `read` | — (`node`, `offset`, `len` = wanted count) | `len` bytes read, payload = the bytes; `len` 0 at end of file |
|
||||
| 3 | `write` | the bytes (`node`, `offset`, `len` = count) | `len` = bytes accepted (may be short — loop) |
|
||||
| 4 | `status` | — (`node` set) | payload = **FileStatus** (24 bytes) |
|
||||
| 5 | `readdir` | — (`node` = a directory, `offset` = cursor) | payload = one **DirectoryEntry** + name; `len` 0 at end |
|
||||
| 6 | `mount` | the mount-point path; the backend endpoint rides as the call's **capability** | status only |
|
||||
| 7 | `unmount` | the mount-point path | status only |
|
||||
| 8 | `mkdir` | the path | status only |
|
||||
| 9 | `unlink` | the path | status only |
|
||||
| 10 | `rename` | old path, one `0x00`, new path (`len` = total) | status only |
|
||||
Each row's *request* and *reply* name the bytes **after** the 16-byte prefix.
|
||||
|
||||
| value | operation | request | tail | reply | reply tail |
|
||||
|------:|-----------|---------|------|-------|-----------|
|
||||
| 16 | `open` | `flags` (4 bytes, below) | the path | `node` (8 bytes) = the open-node id | — |
|
||||
| 17 | `close` | — | — | — | — |
|
||||
| 18 | `read` | `offset` (8), `len` (4) = wanted count | — | — | the bytes read; `Status.len` 0 at end of file |
|
||||
| 19 | `write` | `offset` (8), `len` (4) = count | the bytes | `count` (4) = bytes accepted (may be short — loop) | — |
|
||||
| 20 | `status` | — | — | **FileStatus** (24 bytes) | — |
|
||||
| 21 | `readdir` | `cursor` (8) | — | one **DirectoryEntry** (16 bytes) | the name |
|
||||
| 22 | `mount` | — | the mount-point path; the backend endpoint rides as the call's **capability** | — | — |
|
||||
| 23 | `unmount` | — | the mount-point path | — | — |
|
||||
| 24 | `mkdir` | — | the path | — | — |
|
||||
| 25 | `unlink` | — | the path | — | — |
|
||||
| 26 | `rename` | — | old path, one `0x00`, new path | — | — |
|
||||
| 27 | `bind` | — | the contract name; the provider's endpoint rides as the call's **capability** | — | — |
|
||||
|
||||
Notes per operation:
|
||||
|
||||
- **open** — the path is the mount-relative path `fs_resolve` handed back
|
||||
(absolute-shaped: `/notes.txt` under fat's `/mnt/usb` mount). Bare names
|
||||
(absolute-shaped: `/notes.txt` under fat's `/volumes/usb` mount). Bare names
|
||||
(`greeting`) resolve nowhere — the flat ramfs is retired, and `fs_resolve`
|
||||
refuses non-absolute paths. The returned `node` is the *backend's* own
|
||||
open-node id: with the router in the kernel there is no forwarding table,
|
||||
and clients hold backend ids directly (see *Lifetimes and trust*).
|
||||
and clients hold backend ids directly (see *Lifetimes and trust*). Every
|
||||
later packet carries it in `Header.target` — the path is spoken once, here,
|
||||
and integers do the rest.
|
||||
- **read / write** — a single exchange moves at most 224 bytes
|
||||
(`maximum_payload`); the client loops, advancing `offset` by the returned
|
||||
`len`, until done (read) or the slice is written (write). A `write` reply
|
||||
shorter than requested is progress, not an error; a `len` of 0 means no
|
||||
forward progress — stop rather than spin.
|
||||
- **readdir** — `offset` is a **cursor: the entry index**, not a byte
|
||||
position. Each call returns exactly one entry; the client increments the
|
||||
cursor by 1. A reply with `len` 0 is end-of-directory. The directory must
|
||||
have been opened with the `directory` flag.
|
||||
(`maximum_payload`); the client loops, advancing its own offset by what
|
||||
came back, until done (read) or the slice is written (write). A `write`
|
||||
reply shorter than requested is progress, not an error; a count of 0 means
|
||||
no forward progress — stop rather than spin.
|
||||
- **readdir** — `cursor` is the **entry index**, not a byte position. Each
|
||||
call returns exactly one entry; the client increments the cursor by 1. **A
|
||||
`name_len` of 0 is end-of-directory** — the reply's own length cannot say
|
||||
so, because the envelope always sends the fixed reply part. The directory
|
||||
must have been opened with the `directory` flag.
|
||||
- **mount / unmount** — RETIRED from the wire: mounting is the `fs_mount`
|
||||
syscall now (a filesystem server passes its endpoint handle; possession is
|
||||
the capability, exactly the trust of the old cap-passing op). The op
|
||||
the capability, exactly the trust of the old cap-passing op). The verb
|
||||
numbers stay reserved. Mount-prefix semantics are unchanged: prefixes
|
||||
match at path boundaries only (`/mnt/usb` never captures `/mnt/usbextra`),
|
||||
the longest matching prefix wins, and an optional backend-side rewrite
|
||||
prefix maps a mount into the backend's namespace (fat serves `/mnt/usb`
|
||||
from its volume root and `/var` from its `/var` subtree).
|
||||
match at path boundaries only (`/volumes/usb` never captures
|
||||
`/volumes/usbextra`), the longest matching prefix wins, and an optional
|
||||
backend-side rewrite prefix maps a mount into the backend's namespace (fat
|
||||
serves `/volumes/usb` from its volume root and `/system/logs` from its
|
||||
`/system/logs` subtree).
|
||||
- **rename** — same-directory rename only: the backend compares the old and
|
||||
new parent paths and refuses a mismatch. The client (`file_system`) refuses
|
||||
earlier when the two paths resolve to different backend endpoints, but that
|
||||
check is coarser than "one mount" — one endpoint can serve several mounts
|
||||
(fat serves `/mnt/usb` and `/var`), so a cross-mount rename reaches the
|
||||
backend and fails on its same-directory check.
|
||||
(fat serves `/volumes/usb`, `/system/configuration` and `/system/logs`), so
|
||||
a cross-mount rename reaches the backend and fails on its same-directory
|
||||
check.
|
||||
- **bind** — the protocol registry's claim verb, implemented only by the
|
||||
synthetic `/protocol` backend inside PID 1
|
||||
([protocol-namespace.md](../os-development/protocol-namespace.md)). A file
|
||||
backend answers `-ENOSYS`.
|
||||
|
||||
## Open flags
|
||||
|
||||
Bitwise OR in `Request.flags`, meaningful for `open` only:
|
||||
Bitwise OR in `open`'s `flags`, meaningful for `open` only:
|
||||
|
||||
| bit | name | meaning |
|
||||
|----:|------|---------|
|
||||
@@ -129,7 +154,7 @@ Bitwise OR in `Request.flags`, meaningful for `open` only:
|
||||
| 2 | `directory` | open a directory node for `readdir` rather than a file |
|
||||
| 4 | `truncate` | truncate an existing file to zero length on open (replace, don't overwrite in place) |
|
||||
|
||||
## FileStatus — 24 bytes (the `status` reply payload)
|
||||
## FileStatus — 24 bytes (the `status` reply's fixed part)
|
||||
|
||||
| offset | size | field | meaning |
|
||||
|-------:|-----:|-------|---------|
|
||||
@@ -138,12 +163,12 @@ Bitwise OR in `Request.flags`, meaningful for `open` only:
|
||||
| 12 | 4 | — | padding |
|
||||
| 16 | 8 | `mtime` | modification time, Unix epoch seconds UTC; 0 if the backend keeps none |
|
||||
|
||||
## DirectoryEntry — 16 bytes + name (the `readdir` reply payload)
|
||||
## DirectoryEntry — 16 bytes + name (the `readdir` reply)
|
||||
|
||||
| offset | size | field | meaning |
|
||||
|-------:|-----:|-------|---------|
|
||||
| 0 | 4 | `kind` | a **NodeKind** value |
|
||||
| 4 | 4 | `name_len` | length of the name that follows |
|
||||
| 4 | 4 | `name_len` | length of the name that follows; **0 means end of directory** |
|
||||
| 8 | 8 | `size` | the entry's size in bytes |
|
||||
| 16 | `name_len` | name | the entry's name, not NUL-terminated |
|
||||
|
||||
@@ -174,7 +199,7 @@ volumes, reserved) remain part of the design.
|
||||
## An open reply may carry a capability
|
||||
|
||||
`open` rides `ipc_call`, whose reply direction can hand back an endpoint
|
||||
capability alongside the `Reply` header. A file backend never uses it — FAT
|
||||
capability alongside the reply. A file backend never uses it — FAT
|
||||
answers with a node id and nothing else — but a **synthetic** backend does:
|
||||
opening a `protocol` node returns the provider's endpoint, and possession of
|
||||
that endpoint *is* the channel. The convention is per-backend, not
|
||||
@@ -198,11 +223,15 @@ What a non-Zig implementation may rely on, and what it must not:
|
||||
|
||||
- Operation values, flag bits, `NodeKind` values, and struct layouts are
|
||||
**append-only and frozen once shipped**. The unit test in
|
||||
`library/protocol/vfs/vfs-protocol.zig` pins a sample of them (the `DirectoryEntry`
|
||||
size, `NodeKind` 0–1 and 6–7, `Operation` values 0, 4 and 5); this page is
|
||||
the full record of the frozen values.
|
||||
`library/protocol/vfs/vfs-protocol.zig` pins a sample of them (the
|
||||
`DirectoryEntry` size, `NodeKind` 0–1 and 6–7, `Operation` values 16–21, 26
|
||||
and 27); this page is the full record of the frozen values.
|
||||
*The one renumbering this contract has had was the rebase onto the envelope
|
||||
(P4a), which moved every verb above the reserved range — a deliberate
|
||||
flag-day across a system with no third-party clients yet, not a precedent.*
|
||||
- The 256-byte message ceiling is a property of the current IPC transport,
|
||||
not a promise; clients should read `maximum_payload`-shaped limits from the
|
||||
reply lengths they actually get (loop-until-done), not hard-code 224.
|
||||
- Negative statuses beyond -1 will appear (an errno vocabulary); success is
|
||||
exactly 0.
|
||||
- Success is exactly 0, and the negative statuses come from one system-wide
|
||||
errno vocabulary (the kernel's, continued by the envelope) rather than from
|
||||
this protocol.
|
||||
|
||||
Reference in New Issue
Block a user