diff --git a/docs/README.md b/docs/README.md index c1412d1..5663e48 100644 --- a/docs/README.md +++ b/docs/README.md @@ -208,7 +208,7 @@ the whole reason for the arrangement ([vision.md](vision.md)). danos is a **monorepo of sub-projects**. Each service or driver is a directory that is its own Zig module — it can hold as many files as it needs, and other sub-projects reach it *by module name*, never by a path into its files. The source tree deliberately -**mirrors the runtime FHS** ([danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md)): +**mirrors the runtime file-system hierarchy** ([file-system-hierarchy.md](file-system-development/file-system-hierarchy.md)): what you see under `system/` in the source is what a running danos represents under `/system`. @@ -216,7 +216,7 @@ what you see under `system/` in the source is what a running danos represents un name.** `system/services/init/` contains `init.zig` (its root), and produces a binary addressed as **`system/services/init`** — the repeated leaf resolves away: -| Source (root file) | Addressed as (module / binary / FHS path) | +| Source (root file) | Addressed as (module / binary / hierarchy path) | |----------------------------------------|--------------------------------------------| | `system/services/init/init.zig` | `system/services/init` → `/system/services/init` | | `system/drivers/ps2-bus/ps2-bus.zig` | `system/drivers/ps2-bus` → `/system/drivers/ps2-bus` | diff --git a/docs/character-devices-and-tty.md b/docs/character-devices-and-tty.md index 9c2635f..a4e945d 100644 --- a/docs/character-devices-and-tty.md +++ b/docs/character-devices-and-tty.md @@ -38,8 +38,11 @@ So the design is small: the existing VFS wire protocol, whose read/write have stream semantics.** No device numbers, no `/dev` special casing, no new syscalls, no new protocol — -a service mounts itself at a path (per the FSH, e.g. `/device/console`), clients -open it with `runtime.fs` like any file, and `FileStatus.kind` says what it is. +a service is reachable at a path, clients open it with `runtime.fs` like any +file, and the node kind says what it is. (Since the protocol namespace landed +in design, that path is `/protocol/console` — a protocol node, see +[os-development/protocol-namespace.md](os-development/protocol-namespace.md) — +rather than a mounted device file; the stream semantics below are unchanged.) ### Stream semantics (the actual contract change) @@ -160,7 +163,12 @@ Phase 1 both list; neither track repeats them. - [zig-self-hosting.md](zig-self-hosting.md) — ditto ("stdio as fds"). - [file-system-development/vfs-protocol.md](file-system-development/vfs-protocol.md) — the wire protocol this note extends. -- [file-system-development/danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md) - — where `/device/console` lives. +- [file-system-development/file-system-hierarchy.md](file-system-development/file-system-hierarchy.md) + — the tree the console surfaces in. +- [os-development/protocol-namespace.md](os-development/protocol-namespace.md) — + supersedes this note's device-node naming: the console lands as a protocol + (`/protocol/console`, a protocol node), not a `/dev`-style device file. The + stream semantics designed here (line discipline, cooked/raw modes) carry over + unchanged. - [device-driver-development/input.md](device-driver-development/input.md) — the `InputEvent` stream the console cooks. diff --git a/docs/device-driver-development/ipc.md b/docs/device-driver-development/ipc.md index aae29d6..86274c4 100644 --- a/docs/device-driver-development/ipc.md +++ b/docs/device-driver-development/ipc.md @@ -1,34 +1,70 @@ -# IPC: message-passing channels +# IPC: the kernel-ipc transport Inter-process communication is the **backbone of a microkernel**. Once drivers and services run isolated in their own address spaces ([vision](../vision.md)), they can't -just call each other — a request becomes a **message**. In a microkernel, whatever +just call each other — a request becomes bytes on a wire. In a microkernel, whatever was a function call across a monolithic kernel is IPC, so it's a first-class concern, not an afterthought. -There are two layers, built a milestone apart: +This document describes **one transport** — the bottom layer (L0) of the +communication stack defined in +[communication.md](../os-development/communication.md), which owns the model +and the vocabulary (*protocol*, *channel*, *packet*, *signal*, *endpoint*). +kernel-ipc is the **first** transport, not the only possible one: in +buffer-plus-doorbell terms it is a kernel-owned mailbox with the scheduler as +the doorbell. Its distinguishing properties, which the layers above may rely +on where they say so: -- **`system/kernel/ipc.zig`** — a bounded blocking channel between *kernel threads*, - described below. The primitive, and where the blocking discipline was worked out. -- **`system/kernel/ipc-synchronous.zig`** — synchronous call/reply between *processes*, across - address spaces. What user-space servers and drivers actually talk over. It's the - second half of this document. +- **Rendezvous.** A call is a synchronous meeting, copied sender-page to + receiver-page — natural backpressure, no queue to size. +- **Capability carriage.** The *only* transport that can move a handle + between processes. Channels are therefore always established over + kernel-ipc, and it remains every channel's control path even when bulk + data is negotiated onto a fatter transport (a shared-memory ring). +- **Verified source.** Every delivery carries the kernel-stamped badge — the + identity the channel layer attaches to received packets. +- **Bounded packets.** 256 bytes call/reply, 64 pushed — the floor every + protocol may assume on any transport. -## The channel +Three properties keep the networking analogy honest — kernel-ipc is +networking-*shaped*, not TCP: -The first form is a **bounded blocking channel** (`system/kernel/ipc.zig`): a fixed-size -ring buffer of messages with a producer/consumer rendezvous, built on the -scheduler's [wait queues](../os-development/scheduling.md). +- **Channels over it are RPC-shaped, not streams.** Packets, call/reply, + datagram pushes — closer to UDP plus RPC than to a byte stream. Ordering + exists per exchange (a reply answers its call), not across a channel. +- **Possession is the connection.** There is no handshake state in the + kernel: holding the capability *is* having the channel. A provider's one + endpoint terminates every client's channel at once, demultiplexed by badge + — like every client sharing the server's listening socket, with + per-connection state living in the provider, keyed by badge. A *private* + channel (a dedicated endpoint pair) is built when wanted: that is exactly + what `subscribe` does. +- **Packets never fragment.** If it doesn't fit in a packet, it isn't a + packet: bulk data lives in shared memory and a packet (or signal) is the + doorbell. The display path already works this way. + +The rest of this document is the implementation, bottom-up: the kernel-thread +queue the blocking discipline was worked out on, then endpoints — this +transport's termination points. + +## The kernel-thread queue + +The first form is a **bounded blocking queue** (`system/kernel/ipc.zig`): a +fixed-size ring buffer of messages with a producer/consumer rendezvous, built +on the scheduler's [wait queues](../os-development/scheduling.md). (Its type +is still named `Channel(T, capacity)` — it predates the vocabulary above, and +is a *queue between kernel threads in one address space*, not a channel in +the model's sense; a rename can ride a later flag-day.) `Channel(T, capacity)` is generic over the message type and buffer size. It holds a ring buffer, a count, and two wait queues: -- **`send(msg)`** — if the channel is full, block on the *not-full* queue; otherwise +- **`send(msg)`** — if the queue is full, block on the *not-full* queue; otherwise write the message, bump the count, and wake a waiting receiver. -- **`receive()`** — if the channel is empty, block on the *not-empty* queue; otherwise +- **`receive()`** — if the queue is empty, block on the *not-empty* queue; otherwise take a message, drop the count, and wake a waiting sender. -Neither side busy-waits: a full channel parks the sender, an empty one parks the +Neither side busy-waits: a full queue parks the sender, an empty one parks the receiver, and each operation wakes the other side when it makes progress possible. Two details make it correct: @@ -45,47 +81,52 @@ Two details make it correct: CPU. `waitLocked` / `wakeLocked` are the variants that assume the caller already holds that critical section. -## Verifying it +### Verifying it The `ipc` test (see [testing.md](../testing.md)) runs a producer and a consumer passing -**100 messages through a 4-slot channel**. The small buffer means the channel goes +**100 messages through a 4-slot queue**. The small buffer means the queue goes full and empty over and over, so both the blocking-send and blocking-receive paths are exercised heavily. The messages arrive intact and in order (their sum is the expected `5050`), and neither task busy-waits — they block and wake each other. -## Endpoints: call/reply across address spaces +## Endpoints: the termination points -A channel connects two kernel threads sharing one address space. Real servers are -*processes*, so the payload has to cross an address-space boundary. That's +A queue connects two kernel threads sharing one address space. Real providers are +*processes*, so a packet has to cross an address-space boundary. That's `system/kernel/ipc-synchronous.zig`, and its shape is L4's: a synchronous **rendezvous** at an -`Endpoint`, with the message copied directly from the sender's pages to the receiver's +`Endpoint`, with the packet copied directly from the sender's pages to the receiver's (`copyAcross` walks both sets of page tables through the physmap — no CR3 switch, no bounce buffer). -Two syscalls carry it: +Two syscalls carry the request/reply exchange: -- **`ipc_call(h, msg, reply)`** — copy `msg` to the server, block until it replies. +- **`ipc_call(h, msg, reply)`** — copy the request packet to the provider, block + until the reply packet comes back. - **`ipc_reply_wait(h, reply, recv)`** — reply to the client you're still holding (if - any), then block for the next request. One syscall, because a server's steady state + any), then block for the next request. One syscall, because a provider's steady state is *always* "finish the last one, wait for the next". An endpoint is reached by **handle** — a small integer index into the process's handle -table (`Task.handles`), exactly like a file descriptor, and just as unforgeable. The -bootstrap problem (how do you get the first handle?) is solved by a tiny name registry: -a server calls `ipc_register(service_id, h)` under a well-known small integer, and a -client calls `ipc_lookup(service_id)`. +table (`Task.handles`), exactly like a file descriptor, and just as unforgeable. +The provider never learns the client's identity beyond the **badge** delivered +alongside each packet: the caller's task id, stamped by the kernel — +unforgeable source addressing, a property a network's source field lacks. -The server never learns the client's identity beyond a **badge**, delivered alongside -the message: the caller's task id. +The bootstrap problem — how a channel is first established — is the subject of +[protocol-namespace.md](../os-development/protocol-namespace.md): a protocol is +resolved by name and the channel arrives as a capability. (The mechanism this +replaces, `ipc_register`/`ipc_lookup` under compile-time `ServiceId` integers, +is retired by that design.) -### Interrupts are messages too +### Interrupts are signals -`notifyFromIsr` posts an *asynchronous* notification to an endpoint — no payload, no +`notifyFromIsr` posts an *asynchronous* signal to an endpoint — no payload, no reply owed — and wakes whoever is blocked in `reply_wait`. Its badge has the top bit set (`notify_badge_bit`), which is how a driver's single event loop distinguishes "a -client wants something" from "the hardware wants something". Notifications sit in a +client wants something" from "the hardware wants something". Signals sit in a small coalescing ring on the endpoint, so an interrupt taken while the driver was busy -elsewhere is not lost. +elsewhere is not lost — coalesced, never dropped, which is exactly a signal's +contract (the *count* may collapse; the *fact* may not). This is what makes a user-space driver possible at all, and it's the subject of [drivers.md](drivers.md). @@ -93,40 +134,44 @@ This is what makes a user-space driver possible at all, and it's the subject of ## What's next (partly done since) - **Priority inheritance** through IPC — still open: a high-priority client - blocked on a low-priority server suffers unbounded priority inversion. + blocked on a low-priority provider suffers unbounded priority inversion. - **Handle transfer.** *Landed as cap-passing (M13)*: `ipc_call` and `ipc_reply_wait` carry an optional capability alongside the bytes (`send_cap`), - copying an endpoint or shared-memory handle into the peer's table. First user: - [input](input.md) subscribers register by handing over their own endpoint, and - class drivers get a private channel to one device. + copying an endpoint or shared-memory handle into the peer's table — the + mechanism by which channels are established and private channels built. First + user: [input](input.md) subscribers register by handing over their own + endpoint, and class drivers get a private channel to one device. - **Asynchronous / buffered send** for the cases where a rendezvous is the wrong - shape (logging, notifications between servers). *Landed as `ipc_send`* — a - non-blocking post to an endpoint's bounded payload queue, delivered through - `reply_wait` as a buffered message (badge bit `notify_message_bit`). Built for, and - first used by, the [input service](input.md)'s keyboard-event broadcast, where a - synchronous push would let one dead subscriber hang the fan-out. A full queue drops - the oldest (discrete messages, not a coalescing level like the notification ring). -- **A bounded reply** — half landed. The copy is still 256 bytes - (`MESSAGE_MAXIMUM`) under the big kernel lock, but bulk transfer got its shared + shape (logging, event fan-out). *Landed as `ipc_send`* — a + non-blocking post of an event packet (≤ 64 bytes) to an endpoint's bounded + queue, delivered through `reply_wait` (badge bit `notify_message_bit`). Built + for, and first used by, the [input service](input.md)'s keyboard-event + broadcast, where a synchronous push would let one dead subscriber hang the + fan-out. A full queue drops the oldest — event packets are droppable by + design ([protocol-namespace.md](../os-development/protocol-namespace.md)'s + wiring section states the rule). +- **A bounded reply** — half landed. The copy is still one packet + (256 bytes) under the big kernel lock, but bulk transfer got its shared pages: `shared_memory_create`/`map`/`physical`, the region handle delegated as - a capability (above). virtio-gpu's scanout surface is the first user + a capability (above) — the packets-never-fragment rule in practice. + virtio-gpu's scanout surface is the first user ([display-v2.md](display-v2.md)). ## Lifecycle conventions over IPC (M17) Three conventions from [process-lifecycle.md](../os-development/process-lifecycle.md) ride the -notification mechanism: +signal mechanism: -- **Signals** arrive as notifications on the endpoint a process nominated with - `signal_bind` (`process.bindSignals`): badge = the signal bit plus the - coalesced pending mask (`process.signalsFrom` decodes). Statements, +- **Process signals** arrive as endpoint signals on the endpoint a process + nominated with `signal_bind` (`process.bindSignals`): badge = the signal bit + plus the coalesced pending mask (`process.signalsFrom` decodes). Statements, never questions; no payload, no reply. - **One-shot timers** (`timer_bind`, `time.timerOnce`) land as a - timer-bit notification — the timed wait: a service arms a deadline and keeps + timer-bit signal — the timed wait: a service arms a deadline and keeps serving, instead of blocking in sleep. - **The universal ping**: a **zero-length request is the liveness probe**, answered with a zero-length reply by the service harness itself (`service.run`). No protocol's requests start at length zero, so the encoding cannot collide, and a wedged service simply fails to answer — which is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service - protocol message. + protocol packet. diff --git a/docs/file-system-development/danos-file-system-hierarchy-FSH.md b/docs/file-system-development/danos-file-system-hierarchy-FSH.md deleted file mode 100644 index 1139577..0000000 --- a/docs/file-system-development/danos-file-system-hierarchy-FSH.md +++ /dev/null @@ -1,128 +0,0 @@ -# DanOS Filesystem Hierarchy Standard (DFHS) - -Most modern Unix and Unix-like operating systems follow the FHS. DanOS has its own FHS structure which extends the unix FHS. Root path resolution is provided by the kernel-resident VFS root (`fs_resolve`, `system/kernel/vfs.zig`); mounted filesystem servers serve the subtrees they own. - -## Directory structure - -| Path | Description | -|------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| / | Primary hierarchy root and root directory of the entire file system hierarchy. | -| /bin | Essential command binaries that need to be available in single-user mode, including to bring up the system or repair it, for all users (e.g., cat, ls, cp). | -| /boot | Boot loader files (e.g., EFI, initial-ramdisk.img ). | -| /dev | POSIX Device files (e.g., /dev/null, /dev/disk0, /dev/tty, /dev/random). | -| /etc | Host-specific system-wide configuration files. | -| /home | Users' home directories, containing saved files, personal settings, etc. | -| /lib | Libraries essential for the binaries in /bin and /sbin. eg realtime, system, ipc etc. | -| /sbin | Essential system binaries (e.g init) | -| /srv | Site-specific data served by this system, such as data and scripts for web servers, data offered by FTP servers, and repositories for version control systems | -| /system | DanOS operating system files (similar idea to C:\Windows). A true representation of danos — its layout mirrors the source tree, so `/system` is what danos *is*. | -| /system/devices | danos virtual device tree e.g. similar to /sys on linux but with danos device tree conventions (the structures in the devices module) | -| /system/drivers | driver binaries, one sub-project each (e.g. /system/drivers/pci-bus, /system/drivers/ps2-bus) | -| /system/services | system-service binaries — init, the FAT server, and other user-mode servers (e.g. /system/services/init, /system/services/fat) | -| /system/kernel | the kernel image | -| /test | Test fixtures for the QEMU integration suite. Read-only and initrd-backed like /system, and its layout likewise mirrors the source tree (the repo's test/ directory). Present on development and test images; a volume without it still boots. | -| /test/system/services | test-fixture binaries (e.g. /test/system/services/vfs-test, /test/system/services/thread-test) — the same path in the repo source tree and on the boot volume | -| /tmp | Directory for temporary files (see also /var/tmp). Often not preserved between system reboots and may be severely size-restricted. | -| /usr | Secondary hierarchy for read-only user data; contains the majority of (multi-)user utilities and applications. Should be shareable and read-only. | -| /var | Variable files: files whose content is expected to continually change during normal operation of the system, such as logs, spool files, and temporary e-mail files. | - -## File types - -POSIX specifies the long format of the ls command to represent the Unix file type as the first letter for an entry. - -| type | symbol | Description | -|-------------------|--------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| regular | - | An ordinary file holding an uninterpreted byte stream. Reads and writes are positional, and the file grows on demand (e.g., a binary in /bin, a config file in /etc). | -| directory | d | A container mapping names to other files. It may only be modified through directory operations, never written to directly. | -| symbolic link | l | A file whose contents are a path that is resolved in its place. The target need not exist, and may cross mount points. | -| FIFO special | p | A named pipe: an in-order byte stream between processes, where writers block until a reader opens the other end. | -| block special | b | A device node addressed in fixed-size blocks with the kernel free to buffer and reorder access (e.g., /dev/disk0). | -| character special | c | A device node addressed as an unbuffered byte stream, delivered to the driver in order (e.g., /dev/tty, /dev/null). | -| socket | s | A named endpoint for bidirectional message-passing between processes, bound to a path rather than an address. | - -## /dev - -`/dev` holds the names through which processes reach devices. It is deliberately not -the device tree: the tree — every node discovered by ACPI or PCI enumeration, with its -resources and its parent — lives under [/system/devices](#directory-structure) and is -addressed by device id. `/dev` is the much smaller set of devices that have a driver -willing to serve them, addressed by name. - -A device node is not a file the VFS can read. The bytes live in a driver process -([drivers.md](../device-driver-development/drivers.md)), so opening a `/dev` name has to resolve to that driver's -IPC endpoint, and subsequent reads and writes are calls against it. Resolve-to-endpoint -is exactly what the kernel's `fs_resolve` already does for any mounted backend, and -`FileStatus.kind` is the field that marks a device node; **what is not implemented today -is `/dev` itself** — no service mounts it. (The flat eight-node ramfs this section once -described is retired: the kernel-resident VFS root in `system/kernel/vfs.zig` serves a -read-only initrd mount per top-level tree — `/system`, and `/test` on images that carry -the fixtures — with real directories and node kinds, and filesystem -backends such as the FAT server mount the rest.) The three sections below describe the -intended shape, and are honest about which parts the kernel can already support. - -### Character devices - -A character device is a byte stream with no addressable position: bytes are delivered -to the driver in the order written, and a read consumes what is there. Terminals, -serial lines, keyboards and mice are all of this shape. These are the natural first -device nodes in danos, because a character driver needs nothing the kernel doesn't -already provide — it claims its device, maps its registers with `mmio_map`, and blocks -on `replyWait` for either an interrupt or a client request. `system/drivers/ps2-bus/ps2-bus.zig` -is already that program, minus the file-node client half. - -The obstacle was never the file type; it is which hardware a ring-3 driver can reach. -Direct `in`/`out` from user space is still a #GP (no TSS I/O bitmap, IOPL never raised), -but a driver no longer needs it: **`io_read`/`io_write`** grant port access the same way -`mmio_map` grants memory — gated by `device_claim` and the device's discovered `io_port` -resource. So the 16550 UART at `0x3F8` and the PS/2 controller at `0x60`/`0x64` (and thus -`/dev/ttyS0` and a keyboard node) are now writable as ordinary ring-3 drivers; the -low-rate legacy hardware that needs port I/O is fine with a syscall per access. A -memory-mapped device such as the framebuffer, needing no port I/O at all, remains the -easiest first entry. - -### Block devices - -A block device is addressed in fixed-size blocks and, unlike a character device, the -layer above is free to buffer, reorder, coalesce and retry requests against it. Disks -and other persistent storage are the whole population of this class. - -A block driver is now **writable, but not yet memory-safe.** Every storage controller -worth naming is a bus master: it is programmed by handing it the physical address of a -descriptor ring and left to read and write memory on its own. That ring is exactly what -**`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its -physical address disclosed — and **`/lib/device/mmio`**'s barriers order the descriptor writes -against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver -can be written today (the M14/M15 work in [driver-model.md](../device-driver-development/driver-model.md); the earlier -"cannot host a block driver at all" is no longer true). - -What is *not* yet true is that it is safe. A device programmed with an arbitrary physical -address writes to arbitrary physical memory, and page tables do not sit between a device -and RAM — an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains -are programmed, so granting a DMA-capable device to a driver process is still equivalent -to granting ring 0. Until per-device domains confine a driver's DMA to the buffers it -`dma_alloc`'d, a block driver works but forfeits the isolation that motivates user-space -drivers — enforcement is the next step, and lands with that first driver. A ramdisk over -the initial ramdisk remains the one block-shaped thing that needs no driver process at all. - -### Pseudo-devices - -A pseudo-device has the interface of a device and no hardware behind it: `/dev/null` -discarding writes and reading as end-of-file, `/dev/zero` reading as an endless run of -zero bytes, `/dev/full` failing writes with `ENOSPC`, `/dev/random` and `/dev/urandom` -yielding unpredictable bytes. - -These are the only `/dev` entries danos can implement immediately, and they are the -sensible place to start, because they are exactly the entries that need no driver -process, no `device_claim`, no MMIO grant and no interrupt. A future pseudo-device -service would answer them out of its own address space — `null` and `zero` are a few -lines each in its `read` and `write` handlers — and mount itself at `/dev` the way the -FAT server mounts `/mnt/usb`. The two pieces of structure every later device node -depends on (and that the flat ramfs of the time lacked) exist now: directories, so that -`/dev/null` is a path rather than a name; and a populated `FileStatus.kind`, so that a -caller can tell a character device from a regular file. - -`/dev/random` is the one that is not free. It needs an entropy source, and the honest -options on this kernel are `RDRAND`/`RDSEED` where CPUID advertises them, and the HPET -counter's low bits as a poor fallback. Neither is a seeded CSPRNG, and a `/dev/random` -that is merely unpredictable-looking is worse than none — nothing should be keyed from -it until it is a real one. diff --git a/docs/file-system-development/file-system-hierarchy.md b/docs/file-system-development/file-system-hierarchy.md new file mode 100644 index 0000000..4c0c979 --- /dev/null +++ b/docs/file-system-development/file-system-hierarchy.md @@ -0,0 +1,90 @@ +# The danos file-system hierarchy + +danos is not unix, and its tree does not follow the unix FHS. Paths are the +system's universal namespace — files, the device inventory, and protocol +endpoints all live in one tree — but what a path *yields* differs by subtree: +bytes, facts, or a connection. Root path resolution is provided by the +kernel-resident VFS root (`fs_resolve`, `system/kernel/vfs.zig`); mounted +backends serve the subtrees they own. + +Naming follows the codebase conventions: kebab-case, full words, no +abbreviations. Every top-level name says what its subtree *is*. + +## The tree + +| Path | What it is | +|-------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------| +| `/` | The root of the one namespace. | +| `/applications` | Installed applications, one directory per application — the directory is the identity, the same rule as source sub-projects. *(Planned; empty today.)* | +| `/protocol` | The contract namespace: one protocol node per contract, grouped into directories by domain (`/protocol/display`, `/protocol/networking/ip`). Synthetic — no bytes; opening a name yields a connection to the current provider. See [protocol-namespace.md](../os-development/protocol-namespace.md). | +| `/system` | The operating system — what danos *is*. Its program subtrees mirror the source tree exactly. | +| `/system/kernel` | The kernel image. | +| `/system/drivers` | Driver binaries, one per sub-project (`/system/drivers/pci-bus`, `/system/drivers/ps2-bus`). | +| `/system/services` | System-service binaries (`/system/services/init`, `/system/services/fat`). | +| `/system/devices` | The device inventory: every node hardware discovery found, with its resources and parent — the structures of the devices module, as a browsable virtual tree. Informational only; you *read about* hardware here and *talk to* it through `/protocol`. *(Planned; served by device-manager.)* | +| `/system/configuration` | Machine configuration (`init.csv`, `devices.csv`). Writable, served from the boot volume. | +| `/system/logs` | Per-boot logs: `/system/logs//.log`. Writable, served from the boot volume. | +| `/test` | Test fixtures for the QEMU integration suite. Read-only and initrd-backed like the program subtrees of `/system`, mirroring the repo's `test/` directory. Present on development and test images; a volume without it still boots. | +| `/volumes` | Attached storage volumes, one directory per volume (`/volumes/usb`). A volume's own tree appears beneath its name. | + +Read-only and writable halves of `/system`: the program subtrees (`kernel`, +`drivers`, `services`) and the future `devices` are immutable at runtime — +initrd-backed or synthetic — while `configuration` and `logs` are mutable +machine state served by the boot-volume FAT backend. The kernel's +reserved-prefix rule (no mount may shadow `/system`, `/test`, or `/protocol`) +needs a carve-out for exactly these two writable subtrees; that lands with the +path migration below. + +Deliberately not defined yet: a temporary-files location and per-application +mutable storage. Both belong to the `/applications` design and will be +specified there, not guessed at here. + +## Node kinds + +What a path resolves to. These fill `FileStatus.kind` and +`DirectoryEntry.kind` in the [vfs protocol](vfs-protocol.md) +(`library/protocol/vfs/vfs-protocol.zig`); enum values are append-only. + +| Kind | Meaning | +|--------------------|-------------------------------------------------------------------------------------------------------------------------------------------------| +| `regular` | An ordinary file: an uninterpreted byte stream, positional reads and writes, grows on demand. | +| `directory` | A container mapping names to nodes; modified only through directory operations. | +| `character_device` | A node whose read/write have **stream semantics**: unseekable, reads block until bytes exist, size is meaningless. The console and every tty-shaped node ([character-devices-and-tty.md](../character-devices-and-tty.md)); what a POSIX layer's `isatty` detects. | +| `block_device` | A node addressed in fixed-size sectors — a raw volume. Reserved: recognized, nothing serves one yet. | +| `symbolic_link` | Reserved: a recognized value, not implemented by any backend. | +| `fifo` | Reserved for the future pipe object (wanted by the POSIX compatibility layer); not implemented. | +| `protocol` | A node naming a contract: `open` yields an IPC connection (an endpoint capability) instead of a file id — the kind of every leaf under `/protocol`. *(Being added; see protocol-namespace.md.)* | + +Note the layering: `protocol` says what *opening the name* does (you get a +conversation); `character_device`/`block_device` say what *read and write +mean* on a node a provider serves you. The two compose — `/protocol/console` +is a protocol node in the registry, and the node opened over that connection +reports `character_device`, which is what gives it stream semantics. Only +`socket` is retired (its value stays reserved for wire stability): a named +rendezvous point is exactly what a protocol node is. + +## What is deliberately absent + +There is no `/bin`, `/boot`, `/dev`, `/etc`, `/home`, `/lib`, `/mnt`, `/sbin`, +`/srv`, `/tmp`, `/usr`, or `/var`. These encode unix history — the +binary/library split of small disks, configuration-as-scattered-text, devices +as magic files — that danos does not carry. A POSIX compatibility layer (the +Python track's mini-libc) may *present* whichever of these its programs +expect, mapped onto the real tree; the tree itself stays danos-native. + +## Migration + +The tree above is the specification; some code still writes the unix paths it +replaced. The flag-day converting them: + +| Today (in code) | Becomes | Where | +|------------------------------------------|-------------------------------------------|-----------------------------------------------------------------| +| `/etc/init.csv` | `/system/configuration/init.csv` | `system/services/init/init.zig` | +| `/etc/devices.csv` | `/system/configuration/devices.csv` | `system/services/device-manager/device-manager.zig` | +| `/var/log/...` | `/system/logs/...` | `system/services/logger/logger.zig`, the FAT server's `/var` mount | +| `/mnt/usb` | `/volumes/usb` | `system/services/fat/fat.zig`, the fat/vfs tests | +| `ServiceId` lookup | resolve + open under `/protocol` | every service and client; [protocol-namespace.md](../os-development/protocol-namespace.md) | + +The boot-image builder and the on-volume directory layout move in the same +change, so a freshly written image and the paths the services expect never +disagree. diff --git a/docs/file-system-development/vfs-protocol.md b/docs/file-system-development/vfs-protocol.md index 32182b8..b65c3b9 100644 --- a/docs/file-system-development/vfs-protocol.md +++ b/docs/file-system-development/vfs-protocol.md @@ -149,8 +149,8 @@ Bitwise OR in `Request.flags`, meaningful for `open` only: ## NodeKind -Aligned to the FSH file-type table -(docs/danos-file-system-hierarchy-FSH.md): +Aligned to the node-kind table in the file-system hierarchy +(docs/file-system-development/file-system-hierarchy.md): | value | kind | |------:|------| @@ -163,7 +163,12 @@ Aligned to the FSH file-type table | 6 | socket | Clients should map unknown values to *regular* rather than reject — the -table can grow. +table can grow. Kind 6 (`socket`) keeps its wire value but is retired from +the design — a named rendezvous point is exactly what a `protocol` node is, +planned as value 7 with the protocol namespace +(docs/os-development/protocol-namespace.md). `character_device` (stream +semantics — the tty/console shape) and `block_device` (raw sector-addressed +volumes, reserved) remain part of the design. ## Lifetimes and trust diff --git a/docs/os-development/communication.md b/docs/os-development/communication.md new file mode 100644 index 0000000..2e241b8 --- /dev/null +++ b/docs/os-development/communication.md @@ -0,0 +1,120 @@ +# Communication: the four layers + +*Design, agreed 2026-07-31. The model document — the vocabulary and layering +every other communication document speaks.* + +danos separates **what is said** from **how the bytes move**, so that the +mechanism is replaceable. The shape is a network stack's, cut into four +layers; a program only ever touches the top two. + +``` +L3 namespace /protocol/... names establishment points protocol-namespace.md +L2 protocol the language: packet schemas, verbs, targets the envelope, library/protocol/* +L1 channel two ends exchanging packets and signals the client library's Channel +L0 transport a buffer + a doorbell: moves the bytes ipc.md (kernel-ipc), later shm-ring, … +``` + +## Vocabulary + +| Term | Meaning | +|---|---| +| **protocol** | The language: which packets exist, what their fields mean, which verbs a provider answers. Defined transport-independently in a `library/protocol/*` module. | +| **channel** | An open conversation between two processes, speaking one protocol. Established by opening a `/protocol/...` name; both ends can send and receive. | +| **packet** | The unit a protocol transmits: a bounded, atomic header+payload. Never fragmented — if it doesn't fit, it isn't a packet; bulk data rides shared memory with a packet as the doorbell. | +| **signal** | A payload-less poke below the packet layer: "something happened, come look." Coalescing — the count may collapse, the fact may not. | +| **transport** | What moves the bytes of one channel: a buffer plus a doorbell. Chosen (and upgradable) at establishment, invisible above L1. | +| **endpoint** | A termination point where a transport delivers. The kernel-ipc transport's endpoint is its kernel mailbox object. | + +## Addressing: parties by channel, objects by target + +There are no network-style addresses in a packet. The two questions addresses +answer are answered at different layers: + +- **Who am I talking to?** The **channel**, decided once at establishment. + Opening `/protocol/input` yields a channel; every packet sent on it goes to + the peer. Nothing to route per-packet — like TCP, where no HTTP request + carries the server's IP. +- **Who sent this?** Attached to every received packet **by the channel + layer**, from identity the transport can verify — under kernel-ipc, the + kernel-stamped badge. The sender never writes a source field, which is what + makes source unforgeable (the property a network's spoofable source header + lacks). +- **Which of your things?** The packet's **`target`** field: *object* + addressing within the already-chosen peer — the vfs protocol's node id, the + display protocol's layer id, a block volume. `target = 0` addresses the + provider itself; a protocol without objects never uses it. + +`target` is how instance multiplicity stays out of the namespace. Ten USB +sticks and the namespace still holds exactly one name, `/protocol/block`: a +channel to the provider, `enumerate` lists the current volumes as targets, a +`targets_changed` signal announces hotplug, and a read names its volume in +`target`. The unix `/dev/sda`,`/dev/sdb` problem is dissolved, not renamed. + +If a future transport genuinely routes between machines, *it* carries real +source/destination addressing internally at L0 — the way IP runs under TCP — +and none of it surfaces into the packet header. Protocols stay ignorant of +distance. + +## The transport (L0): a buffer and a doorbell + +Strip any transport to its skeleton and the same two parts remain: + +| Transport | Buffer | Doorbell | Status | +|---|---|---|---| +| **kernel-ipc** | kernel-owned mailbox (the `Endpoint`) | the scheduler (rendezvous wake) | the first transport — [ipc.md](../device-driver-development/ipc.md) | +| **shm-ring** | user-owned shared-memory ring | a signal | exists ad hoc (display bulk); to be formalized — the unlock for the 256-byte ceiling | +| network | NIC queue | an interrupt | someday, when danos networks | + +Transports differ in their **properties**, which the channel layer exposes and +the protocol layer may depend on: + +- **packet ceiling** — kernel-ipc: 256 bytes request/reply, 64 pushed. An + shm-ring's ceiling is its slot size. Kernel-ipc's 256 is the *floor* every + protocol may assume everywhere. +- **synchrony** — kernel-ipc's call is a rendezvous: natural backpressure, no + queue to size. An asynchronous transport buffers, so a channel over one + needs explicit flow control. Backpressure is a *transport property*, not a + channel guarantee — protocols that rely on it say so. +- **droppability** — pushed event packets may drop when a ring fills; + request/reply may not. +- **capability carriage** — **only kernel-ipc can move a capability.** + Handles are kernel objects; a user-space ring cannot transfer one. So + kernel-ipc is always the *establishment and control* transport — channels + are born on it, capabilities ride it — even when a channel's data is + negotiated onto something fatter. + +That negotiation is the upgrade path: a channel starts on kernel-ipc; the +protocol's handshake may then delegate a shared-memory region (as a +capability, over kernel-ipc) and move its bulk traffic there. The display +path already does exactly this by hand; formalizing it in the channel layer +makes it every protocol's option. + +## The channel (L1) + +A channel has two ends, and **the ends are peers**: each may send packets, +each may receive, each may signal. Request/reply is a *pattern* over the +channel — a send with a correlated receive, which the kernel-ipc transport +happens to accelerate as a single rendezvous — not the definition of it. The +event stream (subscribe, then pushes) and the change signal (poke, then +re-read) are the other two patterns; all three are catalogued in +[protocol-namespace.md](protocol-namespace.md)'s wiring section. + +The channel layer's obligations: deliver packets whole, attach the verified +source to every receive, expose the transport's properties, and hide the +transport's mechanics. The client library's `Channel` type is this layer made +concrete — a program holds channels that speak protocols and never touches a +raw handle. + +## The protocol (L2) and the namespace (L3) + +A protocol defines its packets through the envelope — every packet begins +`{operation, target}`, reserved verbs (`describe`, `enumerate`, `subscribe`, +`unsubscribe`) mean the same thing in every protocol, and `Define` checks +every packet against the transport floor at compile time. The full treatment, +including how names are granted, resolved, and restricted per process, is +[protocol-namespace.md](protocol-namespace.md). + +Establishment points are named by contract — `/protocol/display`, never +`/protocol/ipc-1` — because the name must outlive the mechanism: a +transport named in the namespace could never be swapped, which would defeat +this document's premise. diff --git a/docs/os-development/protocol-namespace.md b/docs/os-development/protocol-namespace.md new file mode 100644 index 0000000..ff3ed92 --- /dev/null +++ b/docs/os-development/protocol-namespace.md @@ -0,0 +1,472 @@ +# The protocol namespace + +*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. Not yet implemented — +the migration plan at the end is the work list.* + +How a program finds, connects to, and is restricted from the things it talks to. +Three ideas, kept deliberately separate: + +1. **Naming** — a path under `/protocol` names a *contract*, not a service. +2. **Access** — resolving that path yields an endpoint *capability*; what a process + cannot resolve, it cannot reach. +3. **Transport** — unchanged: packets over channels, moved by whichever + transport the channel rides (kernel-ipc first). + This document is layers **L3** (the namespace) and **L2** (the protocol + and its envelope) of the communication stack; + [communication.md](communication.md) owns the model and the vocabulary + (*protocol* the language, *channel* the conversation, *packet* the + transmitted unit, *signal* the payload-less poke, *transport* the + replaceable mechanism), and + [ipc.md](../device-driver-development/ipc.md) is the first transport. + +## Why ServiceId has to go + +Today a service calls `ipc_register(service_id, endpoint)` and a client calls +`ipc_lookup(service_id)`, where `ServiceId` is a compile-time enum in `abi.zig` +backed by a flat 16-slot table in the kernel. Three defects, in rising order: + +- **Static.** The id space is baked into the ABI at compile time. A third-party + program can never introduce a service; the one place danos is *less* dynamic + than its own design. +- **Ungated.** `ipc_register` is callable by any process and *replaces* an + existing registration. Any process can hijack `.fat` or `.display` and + impersonate it. `ipc_lookup` is equally ambient. +- **Unrestrictable.** Because lookup is a syscall available to everyone, there is + no point at which "this process may not talk to the display" can be enforced. + Any future file-access restriction would be bypassable by speaking to the FAT + server directly. + +## Naming: contracts, not services + +`/protocol/` names a protocol — the contract a conversation follows — and +resolving it connects you to whatever process currently provides that contract. +The client never cared *which* binary answers; it cares that its messages are +understood. Naming the contract makes that explicit, and buys: + +- **Swappable providers.** Replace the display server; `/protocol/display` + routes to the new one; clients notice nothing. +- **Test fakes.** Spawn a program whose namespace wires `/protocol/display` to a + mock. The name promises the protocol; the mock speaks it. +- **One vocabulary.** The names mirror `library/protocol/`: a program imports + the `display-protocol` module, then opens `/protocol/display`. What you + compiled against and what you ask the namespace for are the same word. + +A leaf names one contract — kebab-case, full words, matching the +`library/protocol/` module that defines its wire format — and related +contracts group into directories: `/protocol/networking/ip`, +`/protocol/networking/bluetooth`. Directories organize *contracts only*; +they never encode addressing (see below), so a directory appears because a +domain has several contracts, never because hardware multiplied. The module +tree mirrors the namespace (`library/protocol/networking/ip` ↔ +`/protocol/networking/ip`), and registrar grants scope naturally to subtrees +— an application installed at `/applications/foo` can be granted +`/protocol/applications/foo/...` and nothing above it. `/protocol` is +top level, beside `/system` and `/applications`, because the boundary it names +is spoken on both sides: applications talk to protocols as much as the OS does +(see [file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md)). + +**Addressing lives inside the protocol, never in the path.** Which volume, which +layer, which input device — that is a destination field in the messages, the way +TCP carries a destination address, and the way danos protocols already work (the +display protocol multiplexes layer ids; the vfs protocol addresses node ids). +The namespace answers exactly one question — *may this process speak this +protocol at all* — so `/protocol/block` is one name no matter how many disks are +attached. The source address is never in the message either: it is the IPC +badge, stamped by the kernel per message, unforgeable — a property TCP's source +address does not have. + +`/system/devices` (the device inventory) stays purely informational: facts for +diagnosis, never a routing mechanism. Unix conflated the two in `/dev`; danos +does not. You *read about* hardware in `/system/devices`; you *talk to* it +through `/protocol`. + +## Resolution: a protocol node in the VFS + +The kernel VFS router already does the hard part: `fs_resolve` matches a mount +prefix and installs the backend's endpoint capability in the caller's handle +table. The registry is just a backend mounted at `/protocol` — ring 3, like FAT. +Connecting is a normal vfs-protocol `open` with one twist in the reply: + +``` +client kernel router registry backend + │ fs_resolve("/protocol/display") │ + │──────────────────────────▶│ prefix match: /protocol │ + │◀── registry endpoint ─────│ (capability installed) │ + │ vfs open("display") ──────────────────────────────────────▶│ + │◀───────────────── Reply + capability = provider endpoint ──│ + │ ipc_call(provider, display-protocol messages...) │ +``` + +Both capability moves use machinery the kernel already has: request-direction +and reply-direction `send_cap` on `call`/`replyWait`. The vfs protocol needs two +additions, both append-only: + +- `NodeKind.protocol` — a node that names a contract; its `open` establishes + a **channel** (delivered as an endpoint capability) instead of returning a + file id. The node is the protocol, the channel is the conversation, and the + addressing inside the packets decides where within the provider each one + lands. `readdir` over `/protocol` lists protocol nodes like any others, so + the tree stays browsable for diagnosis. +- The convention that an `open` reply may carry a capability. File backends + (FAT) never use it; synthetic backends (the registry, later the device + inventory) do. + +The path lookup happens once, at connect time. The hot path — `ipc_call` on the +cached endpoint — is untouched. A provider crash turns the cached endpoint dead +(`-EPEER`), and the client's recovery is to re-resolve: the restart story falls +out of the naming layer for free. + +## Registration: the registrar, held by init + +The registry backend is **init**. It is already PID 1, already spawns every +service from its manifest, and already holds the supervision link to each — it +is the process that *knows* which binary is which. (If init grows +uncomfortable, the same design lifts into a dedicated registry service that +init spawns first and delegates to; nothing below changes.) + +- **Binding.** A service creates its endpoint and sends the registry a `bind` + request with the protocol name as payload and the endpoint attached as the + call's capability. +- **Authorization.** Init's manifest gains a column: the protocols each spawned + binary may bind. A `bind` from any process not granted that name is refused + (`-EPERM`) — the badge identifies the caller, the supervision records map + badge to binary. This is the registrar authority; it never leaves init. +- **Collision is an error.** A name already bound refuses a second bind — never + last-writer-wins. When a provider dies, init (its supervisor) unbinds its + names; the restarted instance binds again. +- **Provenance.** The registry records name → task id → binary path, so a + diagnostic listing answers "who serves this?" at a glance: + + ``` + /protocol/display pid 12 /system/services/display + /protocol/input pid 7 /system/services/input + ``` + +`ipc_register` and `ipc_lookup` retire; the `ServiceId` enum leaves `abi.zig`. +The kernel keeps one residual rule: `/protocol` becomes a reserved prefix like +`/system` — `fs_mount` refuses to shadow it, and init's boot-time mount is the +only one it will ever hold. (Full gating of `fs_mount` is a separate item on +the security track; the reserved prefix closes the hole for this namespace +without waiting for it.) + +## Restriction: per-process namespaces, not ACLs + +danos has no users and no principals, deliberately. Restriction is therefore +**delegation**: what a process may open is decided by whoever spawned it, and +enforcement is absence — a protocol you cannot resolve does not exist for you. +"Permission denied" and "not found" are the same answer, which is the same +discipline the device layer already follows: the claim is the capability; here, +the resolvable name is the capability. + +Two stages, deliberately ordered so the useful half lands first: + +**Stage one — the registry filters by badge.** Init is both the spawner and the +registry, so its manifest already knows which binary may *open* which protocols +(a second manifest column, beside the bind grants). An `open` from a process +whose binary is not granted that protocol is refused. No new kernel mechanism +at all; the display driver's view can be narrowed to nothing, a future +downloaded application's to `display` and `input`, today. + +**Stage two — spawn passes the namespace.** `spawn` gains an initial +capability: the child's connection to *its* registry view, chosen by the +spawner. A newly spawned process starts with an empty handle table and this one +handle — its world is whatever its parent wired in. This removes the last +ambient reach (`fs_resolve` finding `/protocol` globally), lets any supervisor +— not just init — narrow or fake a child's view (an application launcher +granting an app only what its manifest declares; a test harness substituting +every provider), and composes down the supervision tree. Stage one's manifest +column becomes the *content* of the view init builds, so nothing is thrown +away. + +### A worked example: the microphone prompt + +The scenario stage two exists for: an application opens +`/protocol/audio-input`, and the user should be asked. The supervisor is an +ordinary user process — an application launcher — and the flow needs no new +security concepts: + +1. The launcher spawned the app with a namespace channel that terminates at + **the launcher itself**. The app's whole world is a conversation with its + supervisor. +2. The app's `open("audio-input")` packet lands in the launcher, + badge-stamped. The launcher spawned the app, so badge → binary path + (`/applications/foo`) is its own supervision record — "remember my choice" + needs no identity system. +3. Grant unknown → the launcher parks the request and shows a prompt (it is a + user process with display access; init never does UI). Blocking an open on + a human is architecturally fine: opens are connect-time, never hot-path. +4. **Yes** → the launcher opens `/protocol/audio-input` in *its own* + namespace and attaches the resulting channel to the parked reply. The app + cannot tell a prompt happened — a consented open is indistinguishable from + a direct one, merely slower. +5. **No** → refuse the open, indistinguishable from "no such protocol" — or + hand the app a **fake**: a silence-generating provider. The test-fake + mechanism doubles as a privacy feature. + +The capability discipline holds throughout: the launcher can only grant what +it holds — if init never gave the launcher `audio-input`, no prompt can +conjure it. Consent is delegation flowing down the supervision tree, never a +global ACL edit. And the provider still sees the app's badge on every packet, +so a coarser second check at the audio service remains possible. + +Two mechanical requirements this scenario pins on stage two: + +- **Parked replies.** A prompt takes seconds, and the service loop holds one + outstanding reply today — the launcher must park request A, keep serving B + and C, and reply to A later (by badge). The kernel already tracks owed + replies (that is how death delivers `-EPEER`); multiple parked replies is + the extension, in the harness and, if needed, the kernel. +- **Granted channels are dedicated, hence revocable.** Once the app holds a + channel capability, nobody reaches into its handle table — so a + prompt-granted channel must be one that can be *killed*: a dedicated + endpoint pair (or per-client session at the provider) whose death turns + the app's capability into `-EPEER`. Revoking microphone access is then + killing that channel, using machinery that already exists. + +One adjacent problem, named and deferred: **trusted UI**. The prompt is only +meaningful if the app cannot draw a convincing fake or overlay the real one — +a display-layer question (a reserved surface for the supervisor chain), owned +by the display track, not this one. + +Fine-grained restriction *within* a protocol (this process may use volume A but +not volume B) is not the namespace's job. The capability-shaped answer, when it +is needed: the supervisor pre-opens a connection scoped to one target and passes +that connection to the child, which never opens `/protocol/block` at all. +Delegation again, not ACLs. + +## The envelope: one addressing scheme for every protocol + +Every protocol module today hand-rolls its `Request`/`Reply` with an +`operation` first field. That convention becomes a library, so addressing is +uniform and the rules are enforced by construction rather than by review. New +module: **`library/protocol/envelope`** (the one protocol-layer module that is +not itself a protocol). + +```zig +/// Every packet a danos protocol transmits begins with this header. +pub const Header = extern struct { + operation: u32, // the verb; values 0..15 are reserved universal verbs + _padding: u32 = 0, + /// Object addressing, never party addressing: which of the peer's + /// objects this packet operates on — a volume, layer, node, device. + /// 0 addresses the provider itself. Parties are addressed by the + /// channel; the protocol defines target's meaning; the field's place + /// and width are universal. + target: u64 = 0, +}; + +/// Reserved verbs, answered by every provider. +pub const operation_describe: u32 = 0; // -> protocol name, version, target kinds +pub const operation_enumerate: u32 = 1; // -> the current targets, one per reply page +pub const operation_subscribe: u32 = 2; // capability = the subscriber's endpoint +pub const operation_unsubscribe: u32 = 3; +pub const first_protocol_operation: u32 = 16; + +/// Every reply begins with this. +pub const Status = extern struct { + status: i32, // 0 or a negative errno + _padding: u32 = 0, + len: u32 = 0, // payload bytes following the header + _padding2: u32 = 0, +}; +``` + +A protocol is then *defined through* the envelope, not beside it: + +```zig +pub const Protocol = envelope.Define(.{ + .name = "display", + .version = 1, + .operations = &.{ + .{ .name = "configure_layer", .request = ConfigureLayer, .reply = void }, + .{ .name = "blit", .request = Blit, .reply = void }, + ... + }, +}); +``` + +`Define` is comptime and is where the enforcement lives: + +- verbs are numbered automatically from `first_protocol_operation`, so no + protocol can collide with the reserved range; +- every packet is size-checked at compile time against the kernel-ipc floor + — `packet_maximum` (256) for request/reply, `post_maximum` (64) for event + packets. Ceilings are transport properties + ([communication.md](communication.md)); the floor is what every protocol + may assume on any transport. The errors that today surface as runtime + truncation become compile errors, and packets-never-fragment is enforced + at the source; +- the generated type carries encode/decode helpers and a provider-side dispatch + table, so a provider answers `describe` automatically and unknown operations + with `-ENOSYS` uniformly; +- the service harness (`library/kernel/service.zig`) accepts the generated + dispatch type, which is what makes the envelope *enforced*: a protocol that + bypasses `Define` does not plug into the harness. + +Universal conventions that ride on the reserved verbs: + +- **`describe`** is the version handshake. Version lives in the handshake, not + in every message — the 256-byte budget is too small to spend per call. +- **`enumerate`** is how multi-target protocols expose their targets, and the + standard `targets_changed` notification (a notify bit) tells subscribers to + re-enumerate — arrival and removal of volumes, layers, devices all take the + same shape. Hotplug fits the notification ring far better than a filesystem + tree ever did. +- **Source is the badge.** No protocol defines a "sender" field; the kernel's + per-message badge is the only source identity, and providers key per-client + state on it. + +### Paths resolve once; integers do the work + +A rule the envelope makes official: **a path appears in a conversation at most +once — at resolve or open — and everything after it addresses integers.** The +namespace resolves `/protocol/display` to an endpoint; a backend's `open` +resolves a path payload to a node id; from then on every packet carries the +integer in `target`. Integers compare in one instruction and fit the fixed +header, and the 256-byte message budget never re-carries path strings on the +hot path. This is already the system's shape — vfs node ids, display layer ids +— and the envelope pins it as the required shape for every protocol. + +Two integer identities, not to be confused: + +- **An open handle** — what vfs `open` returns today: transient, meaningful + only within one client's session with one provider, swept when the client + exits. Cheap, and all a protocol usually needs. Handles must be **scoped per + client** — validated against the badge, or drawn from a per-client id + namespace. (Today the FAT server's node ids are guessable small integers + honoured across clients; that hole closes with this rule.) +- **A persistent node identity** — a unix inode number, stable across opens + and renames. danos deliberately does not promise this, because FAT cannot + deliver it: a FAT file's identity is its directory entry, and rename or + truncation moves every candidate anchor. If a future filesystem or a cache + layer needs stable identity, that is the backend's promise to make, never + the protocol's assumption. + +The five existing protocol modules (`vfs`, `display`, `input`, `power`, +`block`, plus `scanout`, `usb-transfer`, `device-manager`) rebase onto the +envelope during the migration flag-day. `input-protocol`'s subscribe/publish +split and `vfs-protocol`'s node addressing both map cleanly (`node` and layer +ids become `target`). + +## Wiring: how conversations flow + +The patterns below are channel-layer (L1) shapes; the delivery mechanics are +the kernel-ipc transport's, described here because it is the transport every +channel starts on. Kernel-ipc provides exactly three delivery shapes, and +every one is unicast. An endpoint is a mailbox owned by one process — its +creator receives; anyone holding its capability sends into it. That direction +never reverses: + +1. **Synchronous call** — request/reply. The kernel parks the caller and + `replyWait` delivers the reply straight back, so the provider answers + without holding any capability to the client. Badge-stamped, blocking, and + the *only* shape that carries capabilities (in the request, and in the + reply — which is how a reverse path is bootstrapped). +2. **Asynchronous send** — an event packet pushed into the receiver's post + ring, at most `post_maximum` (64) bytes, no reply owed, never blocks the + sender. Strictly one-way: to be pushed to, you must first hand the pusher + your endpoint. +3. **Signals** — payload-less notification bits, below the packet layer, + coalescing: "something changed, come look." + +A bidirectional link is therefore always **a pair of endpoints**, one per +direction, each delivered by cap-passing. Three conversation patterns are +built from these, and the envelope names all three: + +- **Request/response** — the synchronous call. The default, and the only + place capabilities move. +- **Event stream** — `subscribe` (a synchronous call whose attached + capability is the subscriber's own endpoint), after which the provider + pushes events asynchronously; `unsubscribe` or subscriber exit ends it. + Listened-to, not blocked-on. +- **Change signal** — a signal plus re-read: `targets_changed` → + `enumerate`. For state whose truth lives with the provider. + +**Broadcast is a provider pattern, never a kernel primitive.** The kernel +does not know subscriber sets — a service does. The input service is the +model: sources *publish* (a unicast call to the service), the service +*broadcasts* (a fan-out loop of asynchronous sends over its subscriber list, +so one dead subscriber can never stall the rest). One fan-out point per event +domain, owned by the service that defines the event. + +The harness owns the machinery: the subscriber table, the dead-subscriber +sweep (via process-exit notifications), and the fan-out loop — all written by +hand in `input.zig` today, lifted into the service harness so every protocol +gets identical semantics. `Define` declares a protocol's events (`.events`), +and each event type is checked against `post_maximum` at compile time, +generalizing the assert `input-protocol` already carries. + +**Event packets are droppable.** A slow subscriber's ring fills, and the +provider must not block on it — so an event stream is a hint or a coalescing +signal, never a ledger. Anything that must not be lost is either re-readable +state (the change-signal pattern) or bulk data in shared memory with a +packet as the doorbell, which is how the display path already works — the +packets-never-fragment rule and this one are the same rule seen from two +sides. + +**Source direction (open point).** Today event sources are *clients*: an +input driver resolves `/protocol/input` and delivers each event as a +synchronous `publish` call — one capability, obtained by resolution, covers +everything, and the badge tells the service exactly who each event came from. +The inversion — the service subscribing to each driver — would require every +driver to be individually discoverable and its endpoint ferried to the +service, machinery whose payoff (the service choosing its sources) the +namespace already provides more cheaply: only a process granted open on +`/protocol/input` can publish into it. Sources stay clients for now; +revisited at restriction stage two, when a supervisor can wire capabilities +at spawn time. + +## What this deliberately does not solve + +The wider security track, for which this namespace is the foundation, not the +whole: + +- **File access restriction** — the point of the exercise. The same stage-two + namespace mechanism extends from protocol names to file paths: the spawner + decides which subtrees resolve. Designed separately once this lands. +- `fs_mount` gating beyond the reserved prefixes; `system_spawn` gating; + `klog_read` being world-readable; backends checking the badge on per-node + operations (the FAT server honours node ids across clients today). +- Kernel hardening items already noted in-tree: SMEP/SMAP and SYSRET + canonical-RIP, now designed in [smep-smap.md](smep-smap.md). +- Pipes/FIFOs for the POSIX layer — a byte-stream object *beside* message IPC, + wanted by the Python track, unrelated to naming. +- **Trusted UI** — a permission prompt an application cannot fake or overlay + (see the microphone example). A display-track concern: the supervisor chain + needs a reserved surface. + +## Migration plan + +Flag-day per phase, in the style of the DMA-capability conversion — no +dual-stack periods, the QEMU suite green at each phase boundary. + +**P1 — mechanics, no behavior change.** The `envelope` module with its comptime +`Define`, unit tests; `NodeKind.protocol` and the open-reply-capability +convention in `vfs-protocol`; existing protocols untouched. + +**P2 — the registry.** Init serves `/protocol` (bind with manifest +authorization, collision refusal, unbind on provider death, provenance); +kernel reserves the `/protocol` prefix; every service converts from +`ipc_register` to `bind`, every client from `ipc_lookup` to resolve-and-open; +`ServiceId`, `ipc_register`, `ipc_lookup` deleted. Tests: unauthorized bind +refused, collision refused, provider restart re-binds and a client re-resolves. + +**P3 — restriction, stage one.** The open-grant column in init's manifest; +registry refuses ungranted opens. Test: a fixture process denied a protocol its +neighbour is granted. + +**P4 — protocol rebase.** Existing protocol modules re-expressed through +`Define`; providers move onto the generated dispatch; `describe`/`enumerate` +answered everywhere; the conformance test fixture exercises the reserved verbs +against every registered provider. + +**P5 — restriction, stage two.** Spawn's initial capability; namespace views +built by the spawner; ambient resolution of `/protocol` retired. Includes the +two requirements the microphone example pins: **parked replies** (a +supervisor parks an open, keeps serving, replies later by badge) and +**dedicated, killable granted channels** (revocation = channel death → +`-EPEER`). Scoped separately — it touches `spawn`, the loader contract, and +every supervisor — and lands together with the file-path half of namespacing. + +The unix-path migration ([file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md#migration)) +is independent of P1–P5 and can land before or after. diff --git a/docs/os-development/smep-smap.md b/docs/os-development/smep-smap.md new file mode 100644 index 0000000..7c7a8e0 --- /dev/null +++ b/docs/os-development/smep-smap.md @@ -0,0 +1,131 @@ +# SMEP and SMAP — supervisor-mode hardening + +*Design, 2026-07-31. Not yet implemented. Companion to +[protocol-namespace.md](protocol-namespace.md) on the security track — this is +the hardware half; that is the namespace half.* + +Two CR4 bits that make the CPU refuse the two things a kernel should never do +with user memory: + +- **SMEP** (Supervisor Mode Execution Prevention, CR4 bit 20): instruction + fetch in ring 0 from a page whose U/S bit says *user* → #PF. Kills the + classic ret2usr exploit shape — a kernel bug that redirects control flow + can no longer land in attacker-prepared user code. +- **SMAP** (Supervisor Mode Access Prevention, CR4 bit 21): data read/write + in ring 0 to a user page → #PF, unless `EFLAGS.AC` is set. `stac`/`clac` + open and close deliberate access windows; danos's design needs no windows + at all (below). + +Detection is CPUID leaf 7, subleaf 0, EBX bit 7 (SMEP) and bit 20 (SMAP). +Both bits are per-core state: the BSP and every AP must set them. + +## Why, in danos terms + +Every syscall argument is an attacker-controlled integer, and several take +pointers. A kernel bug that dereferences a crafted pointer reads, writes, or +executes memory of the attacker's choosing — the exact bug class the +isolation tracks exist to prevent. SMEP/SMAP turn that class from "silent +compromise" into "immediate, attributable #PF with a kernel RIP in the log." + +The second benefit matters as much as the first: **SMAP is a permanent +tripwire.** Once it is on, any *future* syscall that touches user memory +directly — instead of going through the checked copy layer — faults the +first time the QEMU suite runs it. The discipline stops depending on review. + +## Where danos already stands + +The design is closer than it looks, because the IPC layer was built right: + +- **The copy layer is already SMAP-proof.** `copyAcross` and `copyFromUser` + (`system/kernel/ipc-synchronous.zig:305,333`) never dereference a user + virtual address: they walk the page tables and move bytes through the + physmap — kernel mappings throughout. SMAP cannot object. +- **Syscall entry already clears AC.** `SFMASK = 0x4_0700` clears IF, TF, + DF, **AC** on every `syscall` + (`system/kernel/architecture/x86_64/per-cpu.zig:76`). The syscall path is + SMAP-clean from day one. +- **The interrupt path is not.** Hardware does *not* clear AC on IDT + delivery, and ring 3 can set AC with `popfq` — so a hostile process could + take an interrupt with AC=1 and have the handler run with SMAP suspended. + `isr_common` (`system/kernel/architecture/x86_64/isr.s:366`) needs a + `clac` beside its `swapgs`. +- **CR4 today:** the BSP inherits firmware CR4 (no kernel write anywhere); + APs set PAE/OSFXSR/OSXMMEXCPT in `trampoline.s:62-68`. Neither path sets + SMEP/SMAP yet, and both must. +- **The stragglers.** A handful of syscalls still dereference user pointers + raw after a bounds check — every one is a SMAP #PF waiting to happen, and + every one is *already* a latent kernel fault today (an unmapped-but-in- + range user page oopses the kernel instead of failing the call). The audit + list, from the 2026-07-31 survey of `system/kernel/process.zig`: + + | Syscall | Raw access | + |---|---| + | `system_spawn` | name + argument blob (`:972`, `:980`) | + | `fs_resolve` | path in, result out (`:1780`, `:1797`) | + | `fs_mount` / `fs_unmount` | prefix + rewrite strings (`:1864`) | + | `fs_node` | read buffer out (`:1820`) | + | `debug_write` | message bytes (`:1700`) | + | `klog_read` | log bytes out (`:1740`) | + | `process_enumerate` | descriptor array out (`:1132`) | + | `device_enumerate` | descriptor array out (`:388`) | + + (Some paths already do it right — the futex word and the device-register + descriptor go through `copyFromUser` (`:1087`, `:924`). The write + direction has no helper yet.) + +- **One known gap inside the copy layer itself:** the walk checks presence, + not the leaf U/S and writable bits (`ipc-synchronous.zig:20-22` flags + this). Today that is nearly moot — the user half contains only mappings + the kernel itself created for that process — but it must close before + shared or copy-on-write mappings exist, and closing it is part of making + the copy layer the single trusted door. + +## The plan + +**H1 — copy discipline (the real work).** A `user-memory` kernel module: +`copyFromUser` / `copyToUser` (the missing write direction) via the physmap +walk, with U/S and writable leaf checks closing the in-tree TODO. Convert +the eight stragglers. This fixes the latent unmapped-page kernel fault on +its own — it is worth doing even if SMEP/SMAP never shipped. QEMU suite +green; no behavior change visible to correct programs. + +**H2 — SMEP.** A leaf-7 feature probe (the kernel has per-leaf `cpuid` +helpers in `apic.zig` to generalize); set CR4.SMEP during per-CPU bring-up +on BSP and APs — prefer the Zig-side per-CPU init over the trampoline +assembly, so one code path covers every core and the trampoline stays +minimal. Audit first that ring 0 never executes user-mapped pages: kernel +text lives in the kernel half, `jump_to_user` is kernel code, and the AP +trampoline page is kernel-mapped — expected clean, verify before flipping. + +**H3 — SMAP.** Add `clac` at `isr_common` entry. `clac` is #UD on CPUs +without SMAP, so the instruction is a 3-byte NOP in the image, patched to +`clac` at boot when CPUID advertises SMAP (one-time patch beats a +conditional branch in the hottest path in the kernel). Then set CR4.SMAP in +the same per-CPU init. From this point the whole QEMU suite doubles as the +enforcement test: any missed raw dereference is a vector-14 with a kernel +RIP and a user CR2 — loud and attributable. + +**H4 — keep it honest.** A line in the coding standards: kernel code +touches user memory only through `user-memory`; there is no `stac` anywhere +in the tree, and a PR that adds one is wrong by definition. SMAP enforces +the rule mechanically at test time. + +Feature-gating follows the timekeeping rule (work on any VM, real Intel, +real AMD): both bits are probed, absence is logged and tolerated — like the +IOMMU's fail-open, the machine still boots, just unhardened. QEMU: TCG +implements both; KVM inherits the host (Intel Ivy Bridge+ for SMEP, +Broadwell+ for SMAP; AMD Zen+ for both). The test images should run with +`-cpu max` so the suite always exercises the enabled paths. + +## Adjacent, deliberately separate + +- **SYSRET canonical-RIP hardening** (`isr.s:192-194` documents it): a + non-canonical return RIP makes `sysretq` #GP *in ring 0* on Intel. Same + hardening bucket, independent fix (validate RCX before `sysretq`, fall + back to `iretq`), should ride the same branch as H2/H3 but is not + SMEP/SMAP. +- **KPTI / Meltdown-class leaks are out of scope.** SMEP/SMAP police + architectural accesses, not speculative ones. danos runs one kernel + mapping in every address space and accepts that on affected hardware; + revisit only if the threat model ever includes hostile native code on + shared machines. diff --git a/docs/os-development/system-image.md b/docs/os-development/system-image.md index a429af9..4f85ec7 100644 --- a/docs/os-development/system-image.md +++ b/docs/os-development/system-image.md @@ -11,7 +11,7 @@ sequential pass and hands the bytes to the kernel unmodified. The capsule is a *performance artifact*, not a source of truth. The boot volume's `/system` and `/test` file trees remain the canonical layout (see -[danos-file-system-hierarchy-FSH.md](../file-system-development/danos-file-system-hierarchy-FSH.md)); +[file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md)); the capsule is a pre-baked snapshot of the same binaries, derived from the same build graph, so the running system is identical whether the loader read the capsule or walked the tree. @@ -36,11 +36,11 @@ so it need be no fancier. Little-endian throughout: ``` Header magic: u32 = "DNR2" (0x32524E44), count: u32 -Entry × count name: [64]u8 (NUL-padded FHS path), offset: u64, len: u64 +Entry × count name: [64]u8 (NUL-padded hierarchy path), offset: u64, len: u64 blobs... each entry's file bytes, at its offset within the image ``` -- **Names are full FHS paths** (`/system/services/init`), not basenames — that +- **Names are full hierarchy paths** (`/system/services/init`), not basenames — that is what "v2" means. The 64-byte capacity matches `abi.maximum_process_name`, so a task named after its binary path is never truncated. Paths longer than 63 bytes are a build error (`pack-system-image.py` rejects them). @@ -54,14 +54,14 @@ blobs... each entry's file bytes, at its offset within the image ## How it is built -`build.zig` maintains one `bundled` list — every user binary and its FHS home. +`build.zig` maintains one `bundled` list — every user binary and its hierarchy home. Three artifacts are derived from that same list, in the same build graph, so they cannot drift apart: -1. **The tree**: each binary installed at its FHS path (`zig-out/system/...` +1. **The tree**: each binary installed at its hierarchy path (`zig-out/system/...` and `zig-out/test/...`, mirrored onto the FAT boot volume by `tools/make-fat-image.py`). -2. **The manifest** (`system/manifest`): the FHS path of every bundled binary, +2. **The manifest** (`system/manifest`): the hierarchy path of every bundled binary, one per line — the loader's per-file fallback input. 3. **The capsule**: `tools/pack-system-image.py` packs the same binaries into the v2 container, installed at `zig-out/boot/system.img` and placed on the @@ -103,7 +103,7 @@ the kernel (`kernel.zig`) then publishes the same bytes twice, to two consumers: - **The process layer** (`process.zig`): `system_spawn` looks binaries up in - the ramdisk via `Reader.find` — exact FHS path, or unique basename for + the ramdisk via `Reader.find` — exact hierarchy path, or unique basename for pre-path callers — and loads them as fresh ring-3 processes. The stored path becomes the task's name. - **The VFS root** (`vfs.zig`, `setInitialRamdisk`): the image is mounted as @@ -111,7 +111,7 @@ consumers: paths, so `/system` and, when the fixtures are bundled, `/test`. Directory nodes are derived from the entry paths (the unique parents), so the trees are listable and their files readable over the normal VFS protocol — the - FHS boot tree every process sees comes straight out of the capsule bytes. + boot tree every process sees comes straight out of the capsule bytes. The image is never copied after the handoff and never mutated: the initrd is immutable, which is what makes the VFS's node serving lock-free. diff --git a/docs/python-on-danos-milestones.md b/docs/python-on-danos-milestones.md index 3a02886..4ade118 100644 --- a/docs/python-on-danos-milestones.md +++ b/docs/python-on-danos-milestones.md @@ -82,7 +82,7 @@ return real answers. - `--disable-shared`; static `Modules/Setup`: `posix errno _io _codecs _weakref time math _stat _collections itertools _functools _locale _sre` plus what the interpreter core insists on; threadless build (WASI precedent). -- `Lib/` on the FAT image under the FSH (e.g. `/system/python/lib`); +- `Lib/` on the FAT image under the hierarchy (e.g. `/system/python/lib`); `PYTHONHOME` set accordingly; `.pyc` written with **checked-hash invalidation** (FAT's 2-second mtime granularity makes mtime-based validation lie during fast edit-run cycles). diff --git a/docs/python-on-danos.md b/docs/python-on-danos.md index fc464c1..cea6f6a 100644 --- a/docs/python-on-danos.md +++ b/docs/python-on-danos.md @@ -257,5 +257,5 @@ in practice cares. the spawn-argv work extends. - [device-driver-development/ipc.md](device-driver-development/ipc.md) — the IPC surface Python services speak. -- [file-system-development/danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md) +- [file-system-development/file-system-hierarchy.md](file-system-development/file-system-hierarchy.md) — where `Lib/` and `site-packages` land on the image. diff --git a/docs/zig-self-hosting.md b/docs/zig-self-hosting.md index 48cb09c..43efa14 100644 --- a/docs/zig-self-hosting.md +++ b/docs/zig-self-hosting.md @@ -347,7 +347,7 @@ Two current decisions fall out of this roadmap: - [syscall.md](os-development/syscall.md) — the kernel↔runtime ABI `runtime.os` is built on. - [sysv.md](os-development/sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs. - [ipc.md](device-driver-development/ipc.md) — the IPC the VFS/FAT operations travel over. -- [danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md) — the +- [file-system-hierarchy.md](file-system-development/file-system-hierarchy.md) — the filesystem layout the file surface serves. - [coding-standards.md](coding-standards.md) — danos naming (why the compat spellings are confined, and now retired). diff --git a/library/protocol/vfs/vfs-protocol.zig b/library/protocol/vfs/vfs-protocol.zig index d03cda5..fe48fb6 100644 --- a/library/protocol/vfs/vfs-protocol.zig +++ b/library/protocol/vfs/vfs-protocol.zig @@ -33,8 +33,8 @@ pub const Operation = enum(u32) { rename, // rename(old\0new payload) -> status }; -/// The type of a filesystem node, aligned to the FSH file-type table -/// (docs/danos-file-system-hierarchy-FSH.md). Fills `FileStatus.kind` and +/// The type of a filesystem node, aligned to the node-kind table +/// (docs/file-system-development/file-system-hierarchy.md). Fills `FileStatus.kind` and /// `DirectoryEntry.kind`; `regular = 0` keeps the historical hardcoded value. pub const NodeKind = enum(u32) { regular = 0,