docs: the communication stack — /protocol namespace, layered model, non-unix hierarchy, SMEP/SMAP plan
The security-track design set. communication.md is the model: four layers (namespace / protocol / channel / transport), packets and signals, parties addressed by the channel and objects by target, transports as replaceable buffer+doorbell mechanisms. protocol-namespace.md is L3+L2: /protocol names contracts, resolution establishes a channel, init is the registrar, restriction is per-process namespace delegation (with the microphone-prompt worked example), the envelope is the universal packet header, migration P1-P5 retires ServiceId. file-system-hierarchy.md replaces the unix-FHS spec with the danos-native tree (/applications, /protocol, /system, /volumes) and its migration table. ipc.md is re-cut as the kernel-ipc transport document. smep-smap.md designs the kernel hardening: copy discipline for the eight raw user-pointer syscalls, then SMEP, then SMAP as a permanent tripwire.
This commit is contained in:
+2
-2
@@ -208,7 +208,7 @@ the whole reason for the arrangement ([vision.md](vision.md)).
|
||||
danos is a **monorepo of sub-projects**. Each service or driver is a directory that is
|
||||
its own Zig module — it can hold as many files as it needs, and other sub-projects
|
||||
reach it *by module name*, never by a path into its files. The source tree deliberately
|
||||
**mirrors the runtime FHS** ([danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md)):
|
||||
**mirrors the runtime file-system hierarchy** ([file-system-hierarchy.md](file-system-development/file-system-hierarchy.md)):
|
||||
what you see under `system/` in the source is what a running danos represents under
|
||||
`/system`.
|
||||
|
||||
@@ -216,7 +216,7 @@ what you see under `system/` in the source is what a running danos represents un
|
||||
name.** `system/services/init/` contains `init.zig` (its root), and produces a binary
|
||||
addressed as **`system/services/init`** — the repeated leaf resolves away:
|
||||
|
||||
| Source (root file) | Addressed as (module / binary / FHS path) |
|
||||
| Source (root file) | Addressed as (module / binary / hierarchy path) |
|
||||
|----------------------------------------|--------------------------------------------|
|
||||
| `system/services/init/init.zig` | `system/services/init` → `/system/services/init` |
|
||||
| `system/drivers/ps2-bus/ps2-bus.zig` | `system/drivers/ps2-bus` → `/system/drivers/ps2-bus` |
|
||||
|
||||
@@ -38,8 +38,11 @@ So the design is small:
|
||||
the existing VFS wire protocol, whose read/write have stream semantics.**
|
||||
|
||||
No device numbers, no `/dev` special casing, no new syscalls, no new protocol —
|
||||
a service mounts itself at a path (per the FSH, e.g. `/device/console`), clients
|
||||
open it with `runtime.fs` like any file, and `FileStatus.kind` says what it is.
|
||||
a service is reachable at a path, clients open it with `runtime.fs` like any
|
||||
file, and the node kind says what it is. (Since the protocol namespace landed
|
||||
in design, that path is `/protocol/console` — a protocol node, see
|
||||
[os-development/protocol-namespace.md](os-development/protocol-namespace.md) —
|
||||
rather than a mounted device file; the stream semantics below are unchanged.)
|
||||
|
||||
### Stream semantics (the actual contract change)
|
||||
|
||||
@@ -160,7 +163,12 @@ Phase 1 both list; neither track repeats them.
|
||||
- [zig-self-hosting.md](zig-self-hosting.md) — ditto ("stdio as fds").
|
||||
- [file-system-development/vfs-protocol.md](file-system-development/vfs-protocol.md) —
|
||||
the wire protocol this note extends.
|
||||
- [file-system-development/danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md)
|
||||
— where `/device/console` lives.
|
||||
- [file-system-development/file-system-hierarchy.md](file-system-development/file-system-hierarchy.md)
|
||||
— the tree the console surfaces in.
|
||||
- [os-development/protocol-namespace.md](os-development/protocol-namespace.md) —
|
||||
supersedes this note's device-node naming: the console lands as a protocol
|
||||
(`/protocol/console`, a protocol node), not a `/dev`-style device file. The
|
||||
stream semantics designed here (line discipline, cooked/raw modes) carry over
|
||||
unchanged.
|
||||
- [device-driver-development/input.md](device-driver-development/input.md) — the
|
||||
`InputEvent` stream the console cooks.
|
||||
|
||||
@@ -1,34 +1,70 @@
|
||||
# IPC: message-passing channels
|
||||
# IPC: the kernel-ipc transport
|
||||
|
||||
Inter-process communication is the **backbone of a microkernel**. Once drivers and
|
||||
services run isolated in their own address spaces ([vision](../vision.md)), they can't
|
||||
just call each other — a request becomes a **message**. In a microkernel, whatever
|
||||
just call each other — a request becomes bytes on a wire. In a microkernel, whatever
|
||||
was a function call across a monolithic kernel is IPC, so it's a first-class
|
||||
concern, not an afterthought.
|
||||
|
||||
There are two layers, built a milestone apart:
|
||||
This document describes **one transport** — the bottom layer (L0) of the
|
||||
communication stack defined in
|
||||
[communication.md](../os-development/communication.md), which owns the model
|
||||
and the vocabulary (*protocol*, *channel*, *packet*, *signal*, *endpoint*).
|
||||
kernel-ipc is the **first** transport, not the only possible one: in
|
||||
buffer-plus-doorbell terms it is a kernel-owned mailbox with the scheduler as
|
||||
the doorbell. Its distinguishing properties, which the layers above may rely
|
||||
on where they say so:
|
||||
|
||||
- **`system/kernel/ipc.zig`** — a bounded blocking channel between *kernel threads*,
|
||||
described below. The primitive, and where the blocking discipline was worked out.
|
||||
- **`system/kernel/ipc-synchronous.zig`** — synchronous call/reply between *processes*, across
|
||||
address spaces. What user-space servers and drivers actually talk over. It's the
|
||||
second half of this document.
|
||||
- **Rendezvous.** A call is a synchronous meeting, copied sender-page to
|
||||
receiver-page — natural backpressure, no queue to size.
|
||||
- **Capability carriage.** The *only* transport that can move a handle
|
||||
between processes. Channels are therefore always established over
|
||||
kernel-ipc, and it remains every channel's control path even when bulk
|
||||
data is negotiated onto a fatter transport (a shared-memory ring).
|
||||
- **Verified source.** Every delivery carries the kernel-stamped badge — the
|
||||
identity the channel layer attaches to received packets.
|
||||
- **Bounded packets.** 256 bytes call/reply, 64 pushed — the floor every
|
||||
protocol may assume on any transport.
|
||||
|
||||
## The channel
|
||||
Three properties keep the networking analogy honest — kernel-ipc is
|
||||
networking-*shaped*, not TCP:
|
||||
|
||||
The first form is a **bounded blocking channel** (`system/kernel/ipc.zig`): a fixed-size
|
||||
ring buffer of messages with a producer/consumer rendezvous, built on the
|
||||
scheduler's [wait queues](../os-development/scheduling.md).
|
||||
- **Channels over it are RPC-shaped, not streams.** Packets, call/reply,
|
||||
datagram pushes — closer to UDP plus RPC than to a byte stream. Ordering
|
||||
exists per exchange (a reply answers its call), not across a channel.
|
||||
- **Possession is the connection.** There is no handshake state in the
|
||||
kernel: holding the capability *is* having the channel. A provider's one
|
||||
endpoint terminates every client's channel at once, demultiplexed by badge
|
||||
— like every client sharing the server's listening socket, with
|
||||
per-connection state living in the provider, keyed by badge. A *private*
|
||||
channel (a dedicated endpoint pair) is built when wanted: that is exactly
|
||||
what `subscribe` does.
|
||||
- **Packets never fragment.** If it doesn't fit in a packet, it isn't a
|
||||
packet: bulk data lives in shared memory and a packet (or signal) is the
|
||||
doorbell. The display path already works this way.
|
||||
|
||||
The rest of this document is the implementation, bottom-up: the kernel-thread
|
||||
queue the blocking discipline was worked out on, then endpoints — this
|
||||
transport's termination points.
|
||||
|
||||
## The kernel-thread queue
|
||||
|
||||
The first form is a **bounded blocking queue** (`system/kernel/ipc.zig`): a
|
||||
fixed-size ring buffer of messages with a producer/consumer rendezvous, built
|
||||
on the scheduler's [wait queues](../os-development/scheduling.md). (Its type
|
||||
is still named `Channel(T, capacity)` — it predates the vocabulary above, and
|
||||
is a *queue between kernel threads in one address space*, not a channel in
|
||||
the model's sense; a rename can ride a later flag-day.)
|
||||
|
||||
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
|
||||
ring buffer, a count, and two wait queues:
|
||||
|
||||
- **`send(msg)`** — if the channel is full, block on the *not-full* queue; otherwise
|
||||
- **`send(msg)`** — if the queue is full, block on the *not-full* queue; otherwise
|
||||
write the message, bump the count, and wake a waiting receiver.
|
||||
- **`receive()`** — if the channel is empty, block on the *not-empty* queue; otherwise
|
||||
- **`receive()`** — if the queue is empty, block on the *not-empty* queue; otherwise
|
||||
take a message, drop the count, and wake a waiting sender.
|
||||
|
||||
Neither side busy-waits: a full channel parks the sender, an empty one parks the
|
||||
Neither side busy-waits: a full queue parks the sender, an empty one parks the
|
||||
receiver, and each operation wakes the other side when it makes progress possible.
|
||||
|
||||
Two details make it correct:
|
||||
@@ -45,47 +81,52 @@ Two details make it correct:
|
||||
CPU. `waitLocked` / `wakeLocked` are the variants that assume the caller already
|
||||
holds that critical section.
|
||||
|
||||
## Verifying it
|
||||
### Verifying it
|
||||
|
||||
The `ipc` test (see [testing.md](../testing.md)) runs a producer and a consumer passing
|
||||
**100 messages through a 4-slot channel**. The small buffer means the channel goes
|
||||
**100 messages through a 4-slot queue**. The small buffer means the queue goes
|
||||
full and empty over and over, so both the blocking-send and blocking-receive paths are
|
||||
exercised heavily. The messages arrive intact and in order (their sum is the
|
||||
expected `5050`), and neither task busy-waits — they block and wake each other.
|
||||
|
||||
## Endpoints: call/reply across address spaces
|
||||
## Endpoints: the termination points
|
||||
|
||||
A channel connects two kernel threads sharing one address space. Real servers are
|
||||
*processes*, so the payload has to cross an address-space boundary. That's
|
||||
A queue connects two kernel threads sharing one address space. Real providers are
|
||||
*processes*, so a packet has to cross an address-space boundary. That's
|
||||
`system/kernel/ipc-synchronous.zig`, and its shape is L4's: a synchronous **rendezvous** at an
|
||||
`Endpoint`, with the message copied directly from the sender's pages to the receiver's
|
||||
`Endpoint`, with the packet copied directly from the sender's pages to the receiver's
|
||||
(`copyAcross` walks both sets of page tables through the physmap — no CR3 switch, no
|
||||
bounce buffer).
|
||||
|
||||
Two syscalls carry it:
|
||||
Two syscalls carry the request/reply exchange:
|
||||
|
||||
- **`ipc_call(h, msg, reply)`** — copy `msg` to the server, block until it replies.
|
||||
- **`ipc_call(h, msg, reply)`** — copy the request packet to the provider, block
|
||||
until the reply packet comes back.
|
||||
- **`ipc_reply_wait(h, reply, recv)`** — reply to the client you're still holding (if
|
||||
any), then block for the next request. One syscall, because a server's steady state
|
||||
any), then block for the next request. One syscall, because a provider's steady state
|
||||
is *always* "finish the last one, wait for the next".
|
||||
|
||||
An endpoint is reached by **handle** — a small integer index into the process's handle
|
||||
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable. The
|
||||
bootstrap problem (how do you get the first handle?) is solved by a tiny name registry:
|
||||
a server calls `ipc_register(service_id, h)` under a well-known small integer, and a
|
||||
client calls `ipc_lookup(service_id)`.
|
||||
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable.
|
||||
The provider never learns the client's identity beyond the **badge** delivered
|
||||
alongside each packet: the caller's task id, stamped by the kernel —
|
||||
unforgeable source addressing, a property a network's source field lacks.
|
||||
|
||||
The server never learns the client's identity beyond a **badge**, delivered alongside
|
||||
the message: the caller's task id.
|
||||
The bootstrap problem — how a channel is first established — is the subject of
|
||||
[protocol-namespace.md](../os-development/protocol-namespace.md): a protocol is
|
||||
resolved by name and the channel arrives as a capability. (The mechanism this
|
||||
replaces, `ipc_register`/`ipc_lookup` under compile-time `ServiceId` integers,
|
||||
is retired by that design.)
|
||||
|
||||
### Interrupts are messages too
|
||||
### Interrupts are signals
|
||||
|
||||
`notifyFromIsr` posts an *asynchronous* notification to an endpoint — no payload, no
|
||||
`notifyFromIsr` posts an *asynchronous* signal to an endpoint — no payload, no
|
||||
reply owed — and wakes whoever is blocked in `reply_wait`. Its badge has the top bit
|
||||
set (`notify_badge_bit`), which is how a driver's single event loop distinguishes "a
|
||||
client wants something" from "the hardware wants something". Notifications sit in a
|
||||
client wants something" from "the hardware wants something". Signals sit in a
|
||||
small coalescing ring on the endpoint, so an interrupt taken while the driver was busy
|
||||
elsewhere is not lost.
|
||||
elsewhere is not lost — coalesced, never dropped, which is exactly a signal's
|
||||
contract (the *count* may collapse; the *fact* may not).
|
||||
|
||||
This is what makes a user-space driver possible at all, and it's the subject of
|
||||
[drivers.md](drivers.md).
|
||||
@@ -93,40 +134,44 @@ This is what makes a user-space driver possible at all, and it's the subject of
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Priority inheritance** through IPC — still open: a high-priority client
|
||||
blocked on a low-priority server suffers unbounded priority inversion.
|
||||
blocked on a low-priority provider suffers unbounded priority inversion.
|
||||
- **Handle transfer.** *Landed as cap-passing (M13)*: `ipc_call` and
|
||||
`ipc_reply_wait` carry an optional capability alongside the bytes (`send_cap`),
|
||||
copying an endpoint or shared-memory handle into the peer's table. First user:
|
||||
[input](input.md) subscribers register by handing over their own endpoint, and
|
||||
class drivers get a private channel to one device.
|
||||
copying an endpoint or shared-memory handle into the peer's table — the
|
||||
mechanism by which channels are established and private channels built. First
|
||||
user: [input](input.md) subscribers register by handing over their own
|
||||
endpoint, and class drivers get a private channel to one device.
|
||||
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
|
||||
shape (logging, notifications between servers). *Landed as `ipc_send`* — a
|
||||
non-blocking post to an endpoint's bounded payload queue, delivered through
|
||||
`reply_wait` as a buffered message (badge bit `notify_message_bit`). Built for, and
|
||||
first used by, the [input service](input.md)'s keyboard-event broadcast, where a
|
||||
synchronous push would let one dead subscriber hang the fan-out. A full queue drops
|
||||
the oldest (discrete messages, not a coalescing level like the notification ring).
|
||||
- **A bounded reply** — half landed. The copy is still 256 bytes
|
||||
(`MESSAGE_MAXIMUM`) under the big kernel lock, but bulk transfer got its shared
|
||||
shape (logging, event fan-out). *Landed as `ipc_send`* — a
|
||||
non-blocking post of an event packet (≤ 64 bytes) to an endpoint's bounded
|
||||
queue, delivered through `reply_wait` (badge bit `notify_message_bit`). Built
|
||||
for, and first used by, the [input service](input.md)'s keyboard-event
|
||||
broadcast, where a synchronous push would let one dead subscriber hang the
|
||||
fan-out. A full queue drops the oldest — event packets are droppable by
|
||||
design ([protocol-namespace.md](../os-development/protocol-namespace.md)'s
|
||||
wiring section states the rule).
|
||||
- **A bounded reply** — half landed. The copy is still one packet
|
||||
(256 bytes) under the big kernel lock, but bulk transfer got its shared
|
||||
pages: `shared_memory_create`/`map`/`physical`, the region handle delegated as
|
||||
a capability (above). virtio-gpu's scanout surface is the first user
|
||||
a capability (above) — the packets-never-fragment rule in practice.
|
||||
virtio-gpu's scanout surface is the first user
|
||||
([display-v2.md](display-v2.md)).
|
||||
|
||||
## Lifecycle conventions over IPC (M17)
|
||||
|
||||
Three conventions from [process-lifecycle.md](../os-development/process-lifecycle.md) ride the
|
||||
notification mechanism:
|
||||
signal mechanism:
|
||||
|
||||
- **Signals** arrive as notifications on the endpoint a process nominated with
|
||||
`signal_bind` (`process.bindSignals`): badge = the signal bit plus the
|
||||
coalesced pending mask (`process.signalsFrom` decodes). Statements,
|
||||
- **Process signals** arrive as endpoint signals on the endpoint a process
|
||||
nominated with `signal_bind` (`process.bindSignals`): badge = the signal bit
|
||||
plus the coalesced pending mask (`process.signalsFrom` decodes). Statements,
|
||||
never questions; no payload, no reply.
|
||||
- **One-shot timers** (`timer_bind`, `time.timerOnce`) land as a
|
||||
timer-bit notification — the timed wait: a service arms a deadline and keeps
|
||||
timer-bit signal — the timed wait: a service arms a deadline and keeps
|
||||
serving, instead of blocking in sleep.
|
||||
- **The universal ping**: a **zero-length request is the liveness probe**,
|
||||
answered with a zero-length reply by the service harness itself
|
||||
(`service.run`). No protocol's requests start at length zero, so the
|
||||
encoding cannot collide, and a wedged service simply fails to answer — which
|
||||
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
|
||||
protocol message.
|
||||
protocol packet.
|
||||
|
||||
@@ -1,128 +0,0 @@
|
||||
# DanOS Filesystem Hierarchy Standard (DFHS)
|
||||
|
||||
Most modern Unix and Unix-like operating systems follow the FHS. DanOS has its own FHS structure which extends the unix FHS. Root path resolution is provided by the kernel-resident VFS root (`fs_resolve`, `system/kernel/vfs.zig`); mounted filesystem servers serve the subtrees they own.
|
||||
|
||||
## Directory structure
|
||||
|
||||
| Path | Description |
|
||||
|------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| / | Primary hierarchy root and root directory of the entire file system hierarchy. |
|
||||
| /bin | Essential command binaries that need to be available in single-user mode, including to bring up the system or repair it, for all users (e.g., cat, ls, cp). |
|
||||
| /boot | Boot loader files (e.g., EFI, initial-ramdisk.img ). |
|
||||
| /dev | POSIX Device files (e.g., /dev/null, /dev/disk0, /dev/tty, /dev/random). |
|
||||
| /etc | Host-specific system-wide configuration files. |
|
||||
| /home | Users' home directories, containing saved files, personal settings, etc. |
|
||||
| /lib | Libraries essential for the binaries in /bin and /sbin. eg realtime, system, ipc etc. |
|
||||
| /sbin | Essential system binaries (e.g init) |
|
||||
| /srv | Site-specific data served by this system, such as data and scripts for web servers, data offered by FTP servers, and repositories for version control systems |
|
||||
| /system | DanOS operating system files (similar idea to C:\Windows). A true representation of danos — its layout mirrors the source tree, so `/system` is what danos *is*. |
|
||||
| /system/devices | danos virtual device tree e.g. similar to /sys on linux but with danos device tree conventions (the structures in the devices module) |
|
||||
| /system/drivers | driver binaries, one sub-project each (e.g. /system/drivers/pci-bus, /system/drivers/ps2-bus) |
|
||||
| /system/services | system-service binaries — init, the FAT server, and other user-mode servers (e.g. /system/services/init, /system/services/fat) |
|
||||
| /system/kernel | the kernel image |
|
||||
| /test | Test fixtures for the QEMU integration suite. Read-only and initrd-backed like /system, and its layout likewise mirrors the source tree (the repo's test/ directory). Present on development and test images; a volume without it still boots. |
|
||||
| /test/system/services | test-fixture binaries (e.g. /test/system/services/vfs-test, /test/system/services/thread-test) — the same path in the repo source tree and on the boot volume |
|
||||
| /tmp | Directory for temporary files (see also /var/tmp). Often not preserved between system reboots and may be severely size-restricted. |
|
||||
| /usr | Secondary hierarchy for read-only user data; contains the majority of (multi-)user utilities and applications. Should be shareable and read-only. |
|
||||
| /var | Variable files: files whose content is expected to continually change during normal operation of the system, such as logs, spool files, and temporary e-mail files. |
|
||||
|
||||
## File types
|
||||
|
||||
POSIX specifies the long format of the ls command to represent the Unix file type as the first letter for an entry.
|
||||
|
||||
| type | symbol | Description |
|
||||
|-------------------|--------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| regular | - | An ordinary file holding an uninterpreted byte stream. Reads and writes are positional, and the file grows on demand (e.g., a binary in /bin, a config file in /etc). |
|
||||
| directory | d | A container mapping names to other files. It may only be modified through directory operations, never written to directly. |
|
||||
| symbolic link | l | A file whose contents are a path that is resolved in its place. The target need not exist, and may cross mount points. |
|
||||
| FIFO special | p | A named pipe: an in-order byte stream between processes, where writers block until a reader opens the other end. |
|
||||
| block special | b | A device node addressed in fixed-size blocks with the kernel free to buffer and reorder access (e.g., /dev/disk0). |
|
||||
| character special | c | A device node addressed as an unbuffered byte stream, delivered to the driver in order (e.g., /dev/tty, /dev/null). |
|
||||
| socket | s | A named endpoint for bidirectional message-passing between processes, bound to a path rather than an address. |
|
||||
|
||||
## /dev
|
||||
|
||||
`/dev` holds the names through which processes reach devices. It is deliberately not
|
||||
the device tree: the tree — every node discovered by ACPI or PCI enumeration, with its
|
||||
resources and its parent — lives under [/system/devices](#directory-structure) and is
|
||||
addressed by device id. `/dev` is the much smaller set of devices that have a driver
|
||||
willing to serve them, addressed by name.
|
||||
|
||||
A device node is not a file the VFS can read. The bytes live in a driver process
|
||||
([drivers.md](../device-driver-development/drivers.md)), so opening a `/dev` name has to resolve to that driver's
|
||||
IPC endpoint, and subsequent reads and writes are calls against it. Resolve-to-endpoint
|
||||
is exactly what the kernel's `fs_resolve` already does for any mounted backend, and
|
||||
`FileStatus.kind` is the field that marks a device node; **what is not implemented today
|
||||
is `/dev` itself** — no service mounts it. (The flat eight-node ramfs this section once
|
||||
described is retired: the kernel-resident VFS root in `system/kernel/vfs.zig` serves a
|
||||
read-only initrd mount per top-level tree — `/system`, and `/test` on images that carry
|
||||
the fixtures — with real directories and node kinds, and filesystem
|
||||
backends such as the FAT server mount the rest.) The three sections below describe the
|
||||
intended shape, and are honest about which parts the kernel can already support.
|
||||
|
||||
### Character devices
|
||||
|
||||
A character device is a byte stream with no addressable position: bytes are delivered
|
||||
to the driver in the order written, and a read consumes what is there. Terminals,
|
||||
serial lines, keyboards and mice are all of this shape. These are the natural first
|
||||
device nodes in danos, because a character driver needs nothing the kernel doesn't
|
||||
already provide — it claims its device, maps its registers with `mmio_map`, and blocks
|
||||
on `replyWait` for either an interrupt or a client request. `system/drivers/ps2-bus/ps2-bus.zig`
|
||||
is already that program, minus the file-node client half.
|
||||
|
||||
The obstacle was never the file type; it is which hardware a ring-3 driver can reach.
|
||||
Direct `in`/`out` from user space is still a #GP (no TSS I/O bitmap, IOPL never raised),
|
||||
but a driver no longer needs it: **`io_read`/`io_write`** grant port access the same way
|
||||
`mmio_map` grants memory — gated by `device_claim` and the device's discovered `io_port`
|
||||
resource. So the 16550 UART at `0x3F8` and the PS/2 controller at `0x60`/`0x64` (and thus
|
||||
`/dev/ttyS0` and a keyboard node) are now writable as ordinary ring-3 drivers; the
|
||||
low-rate legacy hardware that needs port I/O is fine with a syscall per access. A
|
||||
memory-mapped device such as the framebuffer, needing no port I/O at all, remains the
|
||||
easiest first entry.
|
||||
|
||||
### Block devices
|
||||
|
||||
A block device is addressed in fixed-size blocks and, unlike a character device, the
|
||||
layer above is free to buffer, reorder, coalesce and retry requests against it. Disks
|
||||
and other persistent storage are the whole population of this class.
|
||||
|
||||
A block driver is now **writable, but not yet memory-safe.** Every storage controller
|
||||
worth naming is a bus master: it is programmed by handing it the physical address of a
|
||||
descriptor ring and left to read and write memory on its own. That ring is exactly what
|
||||
**`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its
|
||||
physical address disclosed — and **`/lib/device/mmio`**'s barriers order the descriptor writes
|
||||
against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver
|
||||
can be written today (the M14/M15 work in [driver-model.md](../device-driver-development/driver-model.md); the earlier
|
||||
"cannot host a block driver at all" is no longer true).
|
||||
|
||||
What is *not* yet true is that it is safe. A device programmed with an arbitrary physical
|
||||
address writes to arbitrary physical memory, and page tables do not sit between a device
|
||||
and RAM — an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
|
||||
are programmed, so granting a DMA-capable device to a driver process is still equivalent
|
||||
to granting ring 0. Until per-device domains confine a driver's DMA to the buffers it
|
||||
`dma_alloc`'d, a block driver works but forfeits the isolation that motivates user-space
|
||||
drivers — enforcement is the next step, and lands with that first driver. A ramdisk over
|
||||
the initial ramdisk remains the one block-shaped thing that needs no driver process at all.
|
||||
|
||||
### Pseudo-devices
|
||||
|
||||
A pseudo-device has the interface of a device and no hardware behind it: `/dev/null`
|
||||
discarding writes and reading as end-of-file, `/dev/zero` reading as an endless run of
|
||||
zero bytes, `/dev/full` failing writes with `ENOSPC`, `/dev/random` and `/dev/urandom`
|
||||
yielding unpredictable bytes.
|
||||
|
||||
These are the only `/dev` entries danos can implement immediately, and they are the
|
||||
sensible place to start, because they are exactly the entries that need no driver
|
||||
process, no `device_claim`, no MMIO grant and no interrupt. A future pseudo-device
|
||||
service would answer them out of its own address space — `null` and `zero` are a few
|
||||
lines each in its `read` and `write` handlers — and mount itself at `/dev` the way the
|
||||
FAT server mounts `/mnt/usb`. The two pieces of structure every later device node
|
||||
depends on (and that the flat ramfs of the time lacked) exist now: directories, so that
|
||||
`/dev/null` is a path rather than a name; and a populated `FileStatus.kind`, so that a
|
||||
caller can tell a character device from a regular file.
|
||||
|
||||
`/dev/random` is the one that is not free. It needs an entropy source, and the honest
|
||||
options on this kernel are `RDRAND`/`RDSEED` where CPUID advertises them, and the HPET
|
||||
counter's low bits as a poor fallback. Neither is a seeded CSPRNG, and a `/dev/random`
|
||||
that is merely unpredictable-looking is worse than none — nothing should be keyed from
|
||||
it until it is a real one.
|
||||
@@ -0,0 +1,90 @@
|
||||
# The danos file-system hierarchy
|
||||
|
||||
danos is not unix, and its tree does not follow the unix FHS. Paths are the
|
||||
system's universal namespace — files, the device inventory, and protocol
|
||||
endpoints all live in one tree — but what a path *yields* differs by subtree:
|
||||
bytes, facts, or a connection. Root path resolution is provided by the
|
||||
kernel-resident VFS root (`fs_resolve`, `system/kernel/vfs.zig`); mounted
|
||||
backends serve the subtrees they own.
|
||||
|
||||
Naming follows the codebase conventions: kebab-case, full words, no
|
||||
abbreviations. Every top-level name says what its subtree *is*.
|
||||
|
||||
## The tree
|
||||
|
||||
| Path | What it is |
|
||||
|-------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `/` | The root of the one namespace. |
|
||||
| `/applications` | Installed applications, one directory per application — the directory is the identity, the same rule as source sub-projects. *(Planned; empty today.)* |
|
||||
| `/protocol` | The contract namespace: one protocol node per contract, grouped into directories by domain (`/protocol/display`, `/protocol/networking/ip`). Synthetic — no bytes; opening a name yields a connection to the current provider. See [protocol-namespace.md](../os-development/protocol-namespace.md). |
|
||||
| `/system` | The operating system — what danos *is*. Its program subtrees mirror the source tree exactly. |
|
||||
| `/system/kernel` | The kernel image. |
|
||||
| `/system/drivers` | Driver binaries, one per sub-project (`/system/drivers/pci-bus`, `/system/drivers/ps2-bus`). |
|
||||
| `/system/services` | System-service binaries (`/system/services/init`, `/system/services/fat`). |
|
||||
| `/system/devices` | The device inventory: every node hardware discovery found, with its resources and parent — the structures of the devices module, as a browsable virtual tree. Informational only; you *read about* hardware here and *talk to* it through `/protocol`. *(Planned; served by device-manager.)* |
|
||||
| `/system/configuration` | Machine configuration (`init.csv`, `devices.csv`). Writable, served from the boot volume. |
|
||||
| `/system/logs` | Per-boot logs: `/system/logs/<boot-stamp>/<binary-path>.log`. Writable, served from the boot volume. |
|
||||
| `/test` | Test fixtures for the QEMU integration suite. Read-only and initrd-backed like the program subtrees of `/system`, mirroring the repo's `test/` directory. Present on development and test images; a volume without it still boots. |
|
||||
| `/volumes` | Attached storage volumes, one directory per volume (`/volumes/usb`). A volume's own tree appears beneath its name. |
|
||||
|
||||
Read-only and writable halves of `/system`: the program subtrees (`kernel`,
|
||||
`drivers`, `services`) and the future `devices` are immutable at runtime —
|
||||
initrd-backed or synthetic — while `configuration` and `logs` are mutable
|
||||
machine state served by the boot-volume FAT backend. The kernel's
|
||||
reserved-prefix rule (no mount may shadow `/system`, `/test`, or `/protocol`)
|
||||
needs a carve-out for exactly these two writable subtrees; that lands with the
|
||||
path migration below.
|
||||
|
||||
Deliberately not defined yet: a temporary-files location and per-application
|
||||
mutable storage. Both belong to the `/applications` design and will be
|
||||
specified there, not guessed at here.
|
||||
|
||||
## Node kinds
|
||||
|
||||
What a path resolves to. These fill `FileStatus.kind` and
|
||||
`DirectoryEntry.kind` in the [vfs protocol](vfs-protocol.md)
|
||||
(`library/protocol/vfs/vfs-protocol.zig`); enum values are append-only.
|
||||
|
||||
| Kind | Meaning |
|
||||
|--------------------|-------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `regular` | An ordinary file: an uninterpreted byte stream, positional reads and writes, grows on demand. |
|
||||
| `directory` | A container mapping names to nodes; modified only through directory operations. |
|
||||
| `character_device` | A node whose read/write have **stream semantics**: unseekable, reads block until bytes exist, size is meaningless. The console and every tty-shaped node ([character-devices-and-tty.md](../character-devices-and-tty.md)); what a POSIX layer's `isatty` detects. |
|
||||
| `block_device` | A node addressed in fixed-size sectors — a raw volume. Reserved: recognized, nothing serves one yet. |
|
||||
| `symbolic_link` | Reserved: a recognized value, not implemented by any backend. |
|
||||
| `fifo` | Reserved for the future pipe object (wanted by the POSIX compatibility layer); not implemented. |
|
||||
| `protocol` | A node naming a contract: `open` yields an IPC connection (an endpoint capability) instead of a file id — the kind of every leaf under `/protocol`. *(Being added; see protocol-namespace.md.)* |
|
||||
|
||||
Note the layering: `protocol` says what *opening the name* does (you get a
|
||||
conversation); `character_device`/`block_device` say what *read and write
|
||||
mean* on a node a provider serves you. The two compose — `/protocol/console`
|
||||
is a protocol node in the registry, and the node opened over that connection
|
||||
reports `character_device`, which is what gives it stream semantics. Only
|
||||
`socket` is retired (its value stays reserved for wire stability): a named
|
||||
rendezvous point is exactly what a protocol node is.
|
||||
|
||||
## What is deliberately absent
|
||||
|
||||
There is no `/bin`, `/boot`, `/dev`, `/etc`, `/home`, `/lib`, `/mnt`, `/sbin`,
|
||||
`/srv`, `/tmp`, `/usr`, or `/var`. These encode unix history — the
|
||||
binary/library split of small disks, configuration-as-scattered-text, devices
|
||||
as magic files — that danos does not carry. A POSIX compatibility layer (the
|
||||
Python track's mini-libc) may *present* whichever of these its programs
|
||||
expect, mapped onto the real tree; the tree itself stays danos-native.
|
||||
|
||||
## Migration
|
||||
|
||||
The tree above is the specification; some code still writes the unix paths it
|
||||
replaced. The flag-day converting them:
|
||||
|
||||
| Today (in code) | Becomes | Where |
|
||||
|------------------------------------------|-------------------------------------------|-----------------------------------------------------------------|
|
||||
| `/etc/init.csv` | `/system/configuration/init.csv` | `system/services/init/init.zig` |
|
||||
| `/etc/devices.csv` | `/system/configuration/devices.csv` | `system/services/device-manager/device-manager.zig` |
|
||||
| `/var/log/...` | `/system/logs/...` | `system/services/logger/logger.zig`, the FAT server's `/var` mount |
|
||||
| `/mnt/usb` | `/volumes/usb` | `system/services/fat/fat.zig`, the fat/vfs tests |
|
||||
| `ServiceId` lookup | resolve + open under `/protocol` | every service and client; [protocol-namespace.md](../os-development/protocol-namespace.md) |
|
||||
|
||||
The boot-image builder and the on-volume directory layout move in the same
|
||||
change, so a freshly written image and the paths the services expect never
|
||||
disagree.
|
||||
@@ -149,8 +149,8 @@ Bitwise OR in `Request.flags`, meaningful for `open` only:
|
||||
|
||||
## NodeKind
|
||||
|
||||
Aligned to the FSH file-type table
|
||||
(docs/danos-file-system-hierarchy-FSH.md):
|
||||
Aligned to the node-kind table in the file-system hierarchy
|
||||
(docs/file-system-development/file-system-hierarchy.md):
|
||||
|
||||
| value | kind |
|
||||
|------:|------|
|
||||
@@ -163,7 +163,12 @@ Aligned to the FSH file-type table
|
||||
| 6 | socket |
|
||||
|
||||
Clients should map unknown values to *regular* rather than reject — the
|
||||
table can grow.
|
||||
table can grow. Kind 6 (`socket`) keeps its wire value but is retired from
|
||||
the design — a named rendezvous point is exactly what a `protocol` node is,
|
||||
planned as value 7 with the protocol namespace
|
||||
(docs/os-development/protocol-namespace.md). `character_device` (stream
|
||||
semantics — the tty/console shape) and `block_device` (raw sector-addressed
|
||||
volumes, reserved) remain part of the design.
|
||||
|
||||
## Lifetimes and trust
|
||||
|
||||
|
||||
@@ -0,0 +1,120 @@
|
||||
# Communication: the four layers
|
||||
|
||||
*Design, agreed 2026-07-31. The model document — the vocabulary and layering
|
||||
every other communication document speaks.*
|
||||
|
||||
danos separates **what is said** from **how the bytes move**, so that the
|
||||
mechanism is replaceable. The shape is a network stack's, cut into four
|
||||
layers; a program only ever touches the top two.
|
||||
|
||||
```
|
||||
L3 namespace /protocol/... names establishment points protocol-namespace.md
|
||||
L2 protocol the language: packet schemas, verbs, targets the envelope, library/protocol/*
|
||||
L1 channel two ends exchanging packets and signals the client library's Channel
|
||||
L0 transport a buffer + a doorbell: moves the bytes ipc.md (kernel-ipc), later shm-ring, …
|
||||
```
|
||||
|
||||
## Vocabulary
|
||||
|
||||
| Term | Meaning |
|
||||
|---|---|
|
||||
| **protocol** | The language: which packets exist, what their fields mean, which verbs a provider answers. Defined transport-independently in a `library/protocol/*` module. |
|
||||
| **channel** | An open conversation between two processes, speaking one protocol. Established by opening a `/protocol/...` name; both ends can send and receive. |
|
||||
| **packet** | The unit a protocol transmits: a bounded, atomic header+payload. Never fragmented — if it doesn't fit, it isn't a packet; bulk data rides shared memory with a packet as the doorbell. |
|
||||
| **signal** | A payload-less poke below the packet layer: "something happened, come look." Coalescing — the count may collapse, the fact may not. |
|
||||
| **transport** | What moves the bytes of one channel: a buffer plus a doorbell. Chosen (and upgradable) at establishment, invisible above L1. |
|
||||
| **endpoint** | A termination point where a transport delivers. The kernel-ipc transport's endpoint is its kernel mailbox object. |
|
||||
|
||||
## Addressing: parties by channel, objects by target
|
||||
|
||||
There are no network-style addresses in a packet. The two questions addresses
|
||||
answer are answered at different layers:
|
||||
|
||||
- **Who am I talking to?** The **channel**, decided once at establishment.
|
||||
Opening `/protocol/input` yields a channel; every packet sent on it goes to
|
||||
the peer. Nothing to route per-packet — like TCP, where no HTTP request
|
||||
carries the server's IP.
|
||||
- **Who sent this?** Attached to every received packet **by the channel
|
||||
layer**, from identity the transport can verify — under kernel-ipc, the
|
||||
kernel-stamped badge. The sender never writes a source field, which is what
|
||||
makes source unforgeable (the property a network's spoofable source header
|
||||
lacks).
|
||||
- **Which of your things?** The packet's **`target`** field: *object*
|
||||
addressing within the already-chosen peer — the vfs protocol's node id, the
|
||||
display protocol's layer id, a block volume. `target = 0` addresses the
|
||||
provider itself; a protocol without objects never uses it.
|
||||
|
||||
`target` is how instance multiplicity stays out of the namespace. Ten USB
|
||||
sticks and the namespace still holds exactly one name, `/protocol/block`: a
|
||||
channel to the provider, `enumerate` lists the current volumes as targets, a
|
||||
`targets_changed` signal announces hotplug, and a read names its volume in
|
||||
`target`. The unix `/dev/sda`,`/dev/sdb` problem is dissolved, not renamed.
|
||||
|
||||
If a future transport genuinely routes between machines, *it* carries real
|
||||
source/destination addressing internally at L0 — the way IP runs under TCP —
|
||||
and none of it surfaces into the packet header. Protocols stay ignorant of
|
||||
distance.
|
||||
|
||||
## The transport (L0): a buffer and a doorbell
|
||||
|
||||
Strip any transport to its skeleton and the same two parts remain:
|
||||
|
||||
| Transport | Buffer | Doorbell | Status |
|
||||
|---|---|---|---|
|
||||
| **kernel-ipc** | kernel-owned mailbox (the `Endpoint`) | the scheduler (rendezvous wake) | the first transport — [ipc.md](../device-driver-development/ipc.md) |
|
||||
| **shm-ring** | user-owned shared-memory ring | a signal | exists ad hoc (display bulk); to be formalized — the unlock for the 256-byte ceiling |
|
||||
| network | NIC queue | an interrupt | someday, when danos networks |
|
||||
|
||||
Transports differ in their **properties**, which the channel layer exposes and
|
||||
the protocol layer may depend on:
|
||||
|
||||
- **packet ceiling** — kernel-ipc: 256 bytes request/reply, 64 pushed. An
|
||||
shm-ring's ceiling is its slot size. Kernel-ipc's 256 is the *floor* every
|
||||
protocol may assume everywhere.
|
||||
- **synchrony** — kernel-ipc's call is a rendezvous: natural backpressure, no
|
||||
queue to size. An asynchronous transport buffers, so a channel over one
|
||||
needs explicit flow control. Backpressure is a *transport property*, not a
|
||||
channel guarantee — protocols that rely on it say so.
|
||||
- **droppability** — pushed event packets may drop when a ring fills;
|
||||
request/reply may not.
|
||||
- **capability carriage** — **only kernel-ipc can move a capability.**
|
||||
Handles are kernel objects; a user-space ring cannot transfer one. So
|
||||
kernel-ipc is always the *establishment and control* transport — channels
|
||||
are born on it, capabilities ride it — even when a channel's data is
|
||||
negotiated onto something fatter.
|
||||
|
||||
That negotiation is the upgrade path: a channel starts on kernel-ipc; the
|
||||
protocol's handshake may then delegate a shared-memory region (as a
|
||||
capability, over kernel-ipc) and move its bulk traffic there. The display
|
||||
path already does exactly this by hand; formalizing it in the channel layer
|
||||
makes it every protocol's option.
|
||||
|
||||
## The channel (L1)
|
||||
|
||||
A channel has two ends, and **the ends are peers**: each may send packets,
|
||||
each may receive, each may signal. Request/reply is a *pattern* over the
|
||||
channel — a send with a correlated receive, which the kernel-ipc transport
|
||||
happens to accelerate as a single rendezvous — not the definition of it. The
|
||||
event stream (subscribe, then pushes) and the change signal (poke, then
|
||||
re-read) are the other two patterns; all three are catalogued in
|
||||
[protocol-namespace.md](protocol-namespace.md)'s wiring section.
|
||||
|
||||
The channel layer's obligations: deliver packets whole, attach the verified
|
||||
source to every receive, expose the transport's properties, and hide the
|
||||
transport's mechanics. The client library's `Channel` type is this layer made
|
||||
concrete — a program holds channels that speak protocols and never touches a
|
||||
raw handle.
|
||||
|
||||
## The protocol (L2) and the namespace (L3)
|
||||
|
||||
A protocol defines its packets through the envelope — every packet begins
|
||||
`{operation, target}`, reserved verbs (`describe`, `enumerate`, `subscribe`,
|
||||
`unsubscribe`) mean the same thing in every protocol, and `Define` checks
|
||||
every packet against the transport floor at compile time. The full treatment,
|
||||
including how names are granted, resolved, and restricted per process, is
|
||||
[protocol-namespace.md](protocol-namespace.md).
|
||||
|
||||
Establishment points are named by contract — `/protocol/display`, never
|
||||
`/protocol/ipc-1` — because the name must outlive the mechanism: a
|
||||
transport named in the namespace could never be swapped, which would defeat
|
||||
this document's premise.
|
||||
@@ -0,0 +1,472 @@
|
||||
# The protocol namespace
|
||||
|
||||
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. Not yet implemented —
|
||||
the migration plan at the end is the work list.*
|
||||
|
||||
How a program finds, connects to, and is restricted from the things it talks to.
|
||||
Three ideas, kept deliberately separate:
|
||||
|
||||
1. **Naming** — a path under `/protocol` names a *contract*, not a service.
|
||||
2. **Access** — resolving that path yields an endpoint *capability*; what a process
|
||||
cannot resolve, it cannot reach.
|
||||
3. **Transport** — unchanged: packets over channels, moved by whichever
|
||||
transport the channel rides (kernel-ipc first).
|
||||
This document is layers **L3** (the namespace) and **L2** (the protocol
|
||||
and its envelope) of the communication stack;
|
||||
[communication.md](communication.md) owns the model and the vocabulary
|
||||
(*protocol* the language, *channel* the conversation, *packet* the
|
||||
transmitted unit, *signal* the payload-less poke, *transport* the
|
||||
replaceable mechanism), and
|
||||
[ipc.md](../device-driver-development/ipc.md) is the first transport.
|
||||
|
||||
## Why ServiceId has to go
|
||||
|
||||
Today a service calls `ipc_register(service_id, endpoint)` and a client calls
|
||||
`ipc_lookup(service_id)`, where `ServiceId` is a compile-time enum in `abi.zig`
|
||||
backed by a flat 16-slot table in the kernel. Three defects, in rising order:
|
||||
|
||||
- **Static.** The id space is baked into the ABI at compile time. A third-party
|
||||
program can never introduce a service; the one place danos is *less* dynamic
|
||||
than its own design.
|
||||
- **Ungated.** `ipc_register` is callable by any process and *replaces* an
|
||||
existing registration. Any process can hijack `.fat` or `.display` and
|
||||
impersonate it. `ipc_lookup` is equally ambient.
|
||||
- **Unrestrictable.** Because lookup is a syscall available to everyone, there is
|
||||
no point at which "this process may not talk to the display" can be enforced.
|
||||
Any future file-access restriction would be bypassable by speaking to the FAT
|
||||
server directly.
|
||||
|
||||
## Naming: contracts, not services
|
||||
|
||||
`/protocol/<name>` names a protocol — the contract a conversation follows — and
|
||||
resolving it connects you to whatever process currently provides that contract.
|
||||
The client never cared *which* binary answers; it cares that its messages are
|
||||
understood. Naming the contract makes that explicit, and buys:
|
||||
|
||||
- **Swappable providers.** Replace the display server; `/protocol/display`
|
||||
routes to the new one; clients notice nothing.
|
||||
- **Test fakes.** Spawn a program whose namespace wires `/protocol/display` to a
|
||||
mock. The name promises the protocol; the mock speaks it.
|
||||
- **One vocabulary.** The names mirror `library/protocol/`: a program imports
|
||||
the `display-protocol` module, then opens `/protocol/display`. What you
|
||||
compiled against and what you ask the namespace for are the same word.
|
||||
|
||||
A leaf names one contract — kebab-case, full words, matching the
|
||||
`library/protocol/` module that defines its wire format — and related
|
||||
contracts group into directories: `/protocol/networking/ip`,
|
||||
`/protocol/networking/bluetooth`. Directories organize *contracts only*;
|
||||
they never encode addressing (see below), so a directory appears because a
|
||||
domain has several contracts, never because hardware multiplied. The module
|
||||
tree mirrors the namespace (`library/protocol/networking/ip` ↔
|
||||
`/protocol/networking/ip`), and registrar grants scope naturally to subtrees
|
||||
— an application installed at `/applications/foo` can be granted
|
||||
`/protocol/applications/foo/...` and nothing above it. `/protocol` is
|
||||
top level, beside `/system` and `/applications`, because the boundary it names
|
||||
is spoken on both sides: applications talk to protocols as much as the OS does
|
||||
(see [file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md)).
|
||||
|
||||
**Addressing lives inside the protocol, never in the path.** Which volume, which
|
||||
layer, which input device — that is a destination field in the messages, the way
|
||||
TCP carries a destination address, and the way danos protocols already work (the
|
||||
display protocol multiplexes layer ids; the vfs protocol addresses node ids).
|
||||
The namespace answers exactly one question — *may this process speak this
|
||||
protocol at all* — so `/protocol/block` is one name no matter how many disks are
|
||||
attached. The source address is never in the message either: it is the IPC
|
||||
badge, stamped by the kernel per message, unforgeable — a property TCP's source
|
||||
address does not have.
|
||||
|
||||
`/system/devices` (the device inventory) stays purely informational: facts for
|
||||
diagnosis, never a routing mechanism. Unix conflated the two in `/dev`; danos
|
||||
does not. You *read about* hardware in `/system/devices`; you *talk to* it
|
||||
through `/protocol`.
|
||||
|
||||
## Resolution: a protocol node in the VFS
|
||||
|
||||
The kernel VFS router already does the hard part: `fs_resolve` matches a mount
|
||||
prefix and installs the backend's endpoint capability in the caller's handle
|
||||
table. The registry is just a backend mounted at `/protocol` — ring 3, like FAT.
|
||||
Connecting is a normal vfs-protocol `open` with one twist in the reply:
|
||||
|
||||
```
|
||||
client kernel router registry backend
|
||||
│ fs_resolve("/protocol/display") │
|
||||
│──────────────────────────▶│ prefix match: /protocol │
|
||||
│◀── registry endpoint ─────│ (capability installed) │
|
||||
│ vfs open("display") ──────────────────────────────────────▶│
|
||||
│◀───────────────── Reply + capability = provider endpoint ──│
|
||||
│ ipc_call(provider, display-protocol messages...) │
|
||||
```
|
||||
|
||||
Both capability moves use machinery the kernel already has: request-direction
|
||||
and reply-direction `send_cap` on `call`/`replyWait`. The vfs protocol needs two
|
||||
additions, both append-only:
|
||||
|
||||
- `NodeKind.protocol` — a node that names a contract; its `open` establishes
|
||||
a **channel** (delivered as an endpoint capability) instead of returning a
|
||||
file id. The node is the protocol, the channel is the conversation, and the
|
||||
addressing inside the packets decides where within the provider each one
|
||||
lands. `readdir` over `/protocol` lists protocol nodes like any others, so
|
||||
the tree stays browsable for diagnosis.
|
||||
- The convention that an `open` reply may carry a capability. File backends
|
||||
(FAT) never use it; synthetic backends (the registry, later the device
|
||||
inventory) do.
|
||||
|
||||
The path lookup happens once, at connect time. The hot path — `ipc_call` on the
|
||||
cached endpoint — is untouched. A provider crash turns the cached endpoint dead
|
||||
(`-EPEER`), and the client's recovery is to re-resolve: the restart story falls
|
||||
out of the naming layer for free.
|
||||
|
||||
## Registration: the registrar, held by init
|
||||
|
||||
The registry backend is **init**. It is already PID 1, already spawns every
|
||||
service from its manifest, and already holds the supervision link to each — it
|
||||
is the process that *knows* which binary is which. (If init grows
|
||||
uncomfortable, the same design lifts into a dedicated registry service that
|
||||
init spawns first and delegates to; nothing below changes.)
|
||||
|
||||
- **Binding.** A service creates its endpoint and sends the registry a `bind`
|
||||
request with the protocol name as payload and the endpoint attached as the
|
||||
call's capability.
|
||||
- **Authorization.** Init's manifest gains a column: the protocols each spawned
|
||||
binary may bind. A `bind` from any process not granted that name is refused
|
||||
(`-EPERM`) — the badge identifies the caller, the supervision records map
|
||||
badge to binary. This is the registrar authority; it never leaves init.
|
||||
- **Collision is an error.** A name already bound refuses a second bind — never
|
||||
last-writer-wins. When a provider dies, init (its supervisor) unbinds its
|
||||
names; the restarted instance binds again.
|
||||
- **Provenance.** The registry records name → task id → binary path, so a
|
||||
diagnostic listing answers "who serves this?" at a glance:
|
||||
|
||||
```
|
||||
/protocol/display pid 12 /system/services/display
|
||||
/protocol/input pid 7 /system/services/input
|
||||
```
|
||||
|
||||
`ipc_register` and `ipc_lookup` retire; the `ServiceId` enum leaves `abi.zig`.
|
||||
The kernel keeps one residual rule: `/protocol` becomes a reserved prefix like
|
||||
`/system` — `fs_mount` refuses to shadow it, and init's boot-time mount is the
|
||||
only one it will ever hold. (Full gating of `fs_mount` is a separate item on
|
||||
the security track; the reserved prefix closes the hole for this namespace
|
||||
without waiting for it.)
|
||||
|
||||
## Restriction: per-process namespaces, not ACLs
|
||||
|
||||
danos has no users and no principals, deliberately. Restriction is therefore
|
||||
**delegation**: what a process may open is decided by whoever spawned it, and
|
||||
enforcement is absence — a protocol you cannot resolve does not exist for you.
|
||||
"Permission denied" and "not found" are the same answer, which is the same
|
||||
discipline the device layer already follows: the claim is the capability; here,
|
||||
the resolvable name is the capability.
|
||||
|
||||
Two stages, deliberately ordered so the useful half lands first:
|
||||
|
||||
**Stage one — the registry filters by badge.** Init is both the spawner and the
|
||||
registry, so its manifest already knows which binary may *open* which protocols
|
||||
(a second manifest column, beside the bind grants). An `open` from a process
|
||||
whose binary is not granted that protocol is refused. No new kernel mechanism
|
||||
at all; the display driver's view can be narrowed to nothing, a future
|
||||
downloaded application's to `display` and `input`, today.
|
||||
|
||||
**Stage two — spawn passes the namespace.** `spawn` gains an initial
|
||||
capability: the child's connection to *its* registry view, chosen by the
|
||||
spawner. A newly spawned process starts with an empty handle table and this one
|
||||
handle — its world is whatever its parent wired in. This removes the last
|
||||
ambient reach (`fs_resolve` finding `/protocol` globally), lets any supervisor
|
||||
— not just init — narrow or fake a child's view (an application launcher
|
||||
granting an app only what its manifest declares; a test harness substituting
|
||||
every provider), and composes down the supervision tree. Stage one's manifest
|
||||
column becomes the *content* of the view init builds, so nothing is thrown
|
||||
away.
|
||||
|
||||
### A worked example: the microphone prompt
|
||||
|
||||
The scenario stage two exists for: an application opens
|
||||
`/protocol/audio-input`, and the user should be asked. The supervisor is an
|
||||
ordinary user process — an application launcher — and the flow needs no new
|
||||
security concepts:
|
||||
|
||||
1. The launcher spawned the app with a namespace channel that terminates at
|
||||
**the launcher itself**. The app's whole world is a conversation with its
|
||||
supervisor.
|
||||
2. The app's `open("audio-input")` packet lands in the launcher,
|
||||
badge-stamped. The launcher spawned the app, so badge → binary path
|
||||
(`/applications/foo`) is its own supervision record — "remember my choice"
|
||||
needs no identity system.
|
||||
3. Grant unknown → the launcher parks the request and shows a prompt (it is a
|
||||
user process with display access; init never does UI). Blocking an open on
|
||||
a human is architecturally fine: opens are connect-time, never hot-path.
|
||||
4. **Yes** → the launcher opens `/protocol/audio-input` in *its own*
|
||||
namespace and attaches the resulting channel to the parked reply. The app
|
||||
cannot tell a prompt happened — a consented open is indistinguishable from
|
||||
a direct one, merely slower.
|
||||
5. **No** → refuse the open, indistinguishable from "no such protocol" — or
|
||||
hand the app a **fake**: a silence-generating provider. The test-fake
|
||||
mechanism doubles as a privacy feature.
|
||||
|
||||
The capability discipline holds throughout: the launcher can only grant what
|
||||
it holds — if init never gave the launcher `audio-input`, no prompt can
|
||||
conjure it. Consent is delegation flowing down the supervision tree, never a
|
||||
global ACL edit. And the provider still sees the app's badge on every packet,
|
||||
so a coarser second check at the audio service remains possible.
|
||||
|
||||
Two mechanical requirements this scenario pins on stage two:
|
||||
|
||||
- **Parked replies.** A prompt takes seconds, and the service loop holds one
|
||||
outstanding reply today — the launcher must park request A, keep serving B
|
||||
and C, and reply to A later (by badge). The kernel already tracks owed
|
||||
replies (that is how death delivers `-EPEER`); multiple parked replies is
|
||||
the extension, in the harness and, if needed, the kernel.
|
||||
- **Granted channels are dedicated, hence revocable.** Once the app holds a
|
||||
channel capability, nobody reaches into its handle table — so a
|
||||
prompt-granted channel must be one that can be *killed*: a dedicated
|
||||
endpoint pair (or per-client session at the provider) whose death turns
|
||||
the app's capability into `-EPEER`. Revoking microphone access is then
|
||||
killing that channel, using machinery that already exists.
|
||||
|
||||
One adjacent problem, named and deferred: **trusted UI**. The prompt is only
|
||||
meaningful if the app cannot draw a convincing fake or overlay the real one —
|
||||
a display-layer question (a reserved surface for the supervisor chain), owned
|
||||
by the display track, not this one.
|
||||
|
||||
Fine-grained restriction *within* a protocol (this process may use volume A but
|
||||
not volume B) is not the namespace's job. The capability-shaped answer, when it
|
||||
is needed: the supervisor pre-opens a connection scoped to one target and passes
|
||||
that connection to the child, which never opens `/protocol/block` at all.
|
||||
Delegation again, not ACLs.
|
||||
|
||||
## The envelope: one addressing scheme for every protocol
|
||||
|
||||
Every protocol module today hand-rolls its `Request`/`Reply` with an
|
||||
`operation` first field. That convention becomes a library, so addressing is
|
||||
uniform and the rules are enforced by construction rather than by review. New
|
||||
module: **`library/protocol/envelope`** (the one protocol-layer module that is
|
||||
not itself a protocol).
|
||||
|
||||
```zig
|
||||
/// Every packet a danos protocol transmits begins with this header.
|
||||
pub const Header = extern struct {
|
||||
operation: u32, // the verb; values 0..15 are reserved universal verbs
|
||||
_padding: u32 = 0,
|
||||
/// Object addressing, never party addressing: which of the peer's
|
||||
/// objects this packet operates on — a volume, layer, node, device.
|
||||
/// 0 addresses the provider itself. Parties are addressed by the
|
||||
/// channel; the protocol defines target's meaning; the field's place
|
||||
/// and width are universal.
|
||||
target: u64 = 0,
|
||||
};
|
||||
|
||||
/// Reserved verbs, answered by every provider.
|
||||
pub const operation_describe: u32 = 0; // -> protocol name, version, target kinds
|
||||
pub const operation_enumerate: u32 = 1; // -> the current targets, one per reply page
|
||||
pub const operation_subscribe: u32 = 2; // capability = the subscriber's endpoint
|
||||
pub const operation_unsubscribe: u32 = 3;
|
||||
pub const first_protocol_operation: u32 = 16;
|
||||
|
||||
/// Every reply begins with this.
|
||||
pub const Status = extern struct {
|
||||
status: i32, // 0 or a negative errno
|
||||
_padding: u32 = 0,
|
||||
len: u32 = 0, // payload bytes following the header
|
||||
_padding2: u32 = 0,
|
||||
};
|
||||
```
|
||||
|
||||
A protocol is then *defined through* the envelope, not beside it:
|
||||
|
||||
```zig
|
||||
pub const Protocol = envelope.Define(.{
|
||||
.name = "display",
|
||||
.version = 1,
|
||||
.operations = &.{
|
||||
.{ .name = "configure_layer", .request = ConfigureLayer, .reply = void },
|
||||
.{ .name = "blit", .request = Blit, .reply = void },
|
||||
...
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
`Define` is comptime and is where the enforcement lives:
|
||||
|
||||
- verbs are numbered automatically from `first_protocol_operation`, so no
|
||||
protocol can collide with the reserved range;
|
||||
- every packet is size-checked at compile time against the kernel-ipc floor
|
||||
— `packet_maximum` (256) for request/reply, `post_maximum` (64) for event
|
||||
packets. Ceilings are transport properties
|
||||
([communication.md](communication.md)); the floor is what every protocol
|
||||
may assume on any transport. The errors that today surface as runtime
|
||||
truncation become compile errors, and packets-never-fragment is enforced
|
||||
at the source;
|
||||
- the generated type carries encode/decode helpers and a provider-side dispatch
|
||||
table, so a provider answers `describe` automatically and unknown operations
|
||||
with `-ENOSYS` uniformly;
|
||||
- the service harness (`library/kernel/service.zig`) accepts the generated
|
||||
dispatch type, which is what makes the envelope *enforced*: a protocol that
|
||||
bypasses `Define` does not plug into the harness.
|
||||
|
||||
Universal conventions that ride on the reserved verbs:
|
||||
|
||||
- **`describe`** is the version handshake. Version lives in the handshake, not
|
||||
in every message — the 256-byte budget is too small to spend per call.
|
||||
- **`enumerate`** is how multi-target protocols expose their targets, and the
|
||||
standard `targets_changed` notification (a notify bit) tells subscribers to
|
||||
re-enumerate — arrival and removal of volumes, layers, devices all take the
|
||||
same shape. Hotplug fits the notification ring far better than a filesystem
|
||||
tree ever did.
|
||||
- **Source is the badge.** No protocol defines a "sender" field; the kernel's
|
||||
per-message badge is the only source identity, and providers key per-client
|
||||
state on it.
|
||||
|
||||
### Paths resolve once; integers do the work
|
||||
|
||||
A rule the envelope makes official: **a path appears in a conversation at most
|
||||
once — at resolve or open — and everything after it addresses integers.** The
|
||||
namespace resolves `/protocol/display` to an endpoint; a backend's `open`
|
||||
resolves a path payload to a node id; from then on every packet carries the
|
||||
integer in `target`. Integers compare in one instruction and fit the fixed
|
||||
header, and the 256-byte message budget never re-carries path strings on the
|
||||
hot path. This is already the system's shape — vfs node ids, display layer ids
|
||||
— and the envelope pins it as the required shape for every protocol.
|
||||
|
||||
Two integer identities, not to be confused:
|
||||
|
||||
- **An open handle** — what vfs `open` returns today: transient, meaningful
|
||||
only within one client's session with one provider, swept when the client
|
||||
exits. Cheap, and all a protocol usually needs. Handles must be **scoped per
|
||||
client** — validated against the badge, or drawn from a per-client id
|
||||
namespace. (Today the FAT server's node ids are guessable small integers
|
||||
honoured across clients; that hole closes with this rule.)
|
||||
- **A persistent node identity** — a unix inode number, stable across opens
|
||||
and renames. danos deliberately does not promise this, because FAT cannot
|
||||
deliver it: a FAT file's identity is its directory entry, and rename or
|
||||
truncation moves every candidate anchor. If a future filesystem or a cache
|
||||
layer needs stable identity, that is the backend's promise to make, never
|
||||
the protocol's assumption.
|
||||
|
||||
The five existing protocol modules (`vfs`, `display`, `input`, `power`,
|
||||
`block`, plus `scanout`, `usb-transfer`, `device-manager`) rebase onto the
|
||||
envelope during the migration flag-day. `input-protocol`'s subscribe/publish
|
||||
split and `vfs-protocol`'s node addressing both map cleanly (`node` and layer
|
||||
ids become `target`).
|
||||
|
||||
## Wiring: how conversations flow
|
||||
|
||||
The patterns below are channel-layer (L1) shapes; the delivery mechanics are
|
||||
the kernel-ipc transport's, described here because it is the transport every
|
||||
channel starts on. Kernel-ipc provides exactly three delivery shapes, and
|
||||
every one is unicast. An endpoint is a mailbox owned by one process — its
|
||||
creator receives; anyone holding its capability sends into it. That direction
|
||||
never reverses:
|
||||
|
||||
1. **Synchronous call** — request/reply. The kernel parks the caller and
|
||||
`replyWait` delivers the reply straight back, so the provider answers
|
||||
without holding any capability to the client. Badge-stamped, blocking, and
|
||||
the *only* shape that carries capabilities (in the request, and in the
|
||||
reply — which is how a reverse path is bootstrapped).
|
||||
2. **Asynchronous send** — an event packet pushed into the receiver's post
|
||||
ring, at most `post_maximum` (64) bytes, no reply owed, never blocks the
|
||||
sender. Strictly one-way: to be pushed to, you must first hand the pusher
|
||||
your endpoint.
|
||||
3. **Signals** — payload-less notification bits, below the packet layer,
|
||||
coalescing: "something changed, come look."
|
||||
|
||||
A bidirectional link is therefore always **a pair of endpoints**, one per
|
||||
direction, each delivered by cap-passing. Three conversation patterns are
|
||||
built from these, and the envelope names all three:
|
||||
|
||||
- **Request/response** — the synchronous call. The default, and the only
|
||||
place capabilities move.
|
||||
- **Event stream** — `subscribe` (a synchronous call whose attached
|
||||
capability is the subscriber's own endpoint), after which the provider
|
||||
pushes events asynchronously; `unsubscribe` or subscriber exit ends it.
|
||||
Listened-to, not blocked-on.
|
||||
- **Change signal** — a signal plus re-read: `targets_changed` →
|
||||
`enumerate`. For state whose truth lives with the provider.
|
||||
|
||||
**Broadcast is a provider pattern, never a kernel primitive.** The kernel
|
||||
does not know subscriber sets — a service does. The input service is the
|
||||
model: sources *publish* (a unicast call to the service), the service
|
||||
*broadcasts* (a fan-out loop of asynchronous sends over its subscriber list,
|
||||
so one dead subscriber can never stall the rest). One fan-out point per event
|
||||
domain, owned by the service that defines the event.
|
||||
|
||||
The harness owns the machinery: the subscriber table, the dead-subscriber
|
||||
sweep (via process-exit notifications), and the fan-out loop — all written by
|
||||
hand in `input.zig` today, lifted into the service harness so every protocol
|
||||
gets identical semantics. `Define` declares a protocol's events (`.events`),
|
||||
and each event type is checked against `post_maximum` at compile time,
|
||||
generalizing the assert `input-protocol` already carries.
|
||||
|
||||
**Event packets are droppable.** A slow subscriber's ring fills, and the
|
||||
provider must not block on it — so an event stream is a hint or a coalescing
|
||||
signal, never a ledger. Anything that must not be lost is either re-readable
|
||||
state (the change-signal pattern) or bulk data in shared memory with a
|
||||
packet as the doorbell, which is how the display path already works — the
|
||||
packets-never-fragment rule and this one are the same rule seen from two
|
||||
sides.
|
||||
|
||||
**Source direction (open point).** Today event sources are *clients*: an
|
||||
input driver resolves `/protocol/input` and delivers each event as a
|
||||
synchronous `publish` call — one capability, obtained by resolution, covers
|
||||
everything, and the badge tells the service exactly who each event came from.
|
||||
The inversion — the service subscribing to each driver — would require every
|
||||
driver to be individually discoverable and its endpoint ferried to the
|
||||
service, machinery whose payoff (the service choosing its sources) the
|
||||
namespace already provides more cheaply: only a process granted open on
|
||||
`/protocol/input` can publish into it. Sources stay clients for now;
|
||||
revisited at restriction stage two, when a supervisor can wire capabilities
|
||||
at spawn time.
|
||||
|
||||
## What this deliberately does not solve
|
||||
|
||||
The wider security track, for which this namespace is the foundation, not the
|
||||
whole:
|
||||
|
||||
- **File access restriction** — the point of the exercise. The same stage-two
|
||||
namespace mechanism extends from protocol names to file paths: the spawner
|
||||
decides which subtrees resolve. Designed separately once this lands.
|
||||
- `fs_mount` gating beyond the reserved prefixes; `system_spawn` gating;
|
||||
`klog_read` being world-readable; backends checking the badge on per-node
|
||||
operations (the FAT server honours node ids across clients today).
|
||||
- Kernel hardening items already noted in-tree: SMEP/SMAP and SYSRET
|
||||
canonical-RIP, now designed in [smep-smap.md](smep-smap.md).
|
||||
- Pipes/FIFOs for the POSIX layer — a byte-stream object *beside* message IPC,
|
||||
wanted by the Python track, unrelated to naming.
|
||||
- **Trusted UI** — a permission prompt an application cannot fake or overlay
|
||||
(see the microphone example). A display-track concern: the supervisor chain
|
||||
needs a reserved surface.
|
||||
|
||||
## Migration plan
|
||||
|
||||
Flag-day per phase, in the style of the DMA-capability conversion — no
|
||||
dual-stack periods, the QEMU suite green at each phase boundary.
|
||||
|
||||
**P1 — mechanics, no behavior change.** The `envelope` module with its comptime
|
||||
`Define`, unit tests; `NodeKind.protocol` and the open-reply-capability
|
||||
convention in `vfs-protocol`; existing protocols untouched.
|
||||
|
||||
**P2 — the registry.** Init serves `/protocol` (bind with manifest
|
||||
authorization, collision refusal, unbind on provider death, provenance);
|
||||
kernel reserves the `/protocol` prefix; every service converts from
|
||||
`ipc_register` to `bind`, every client from `ipc_lookup` to resolve-and-open;
|
||||
`ServiceId`, `ipc_register`, `ipc_lookup` deleted. Tests: unauthorized bind
|
||||
refused, collision refused, provider restart re-binds and a client re-resolves.
|
||||
|
||||
**P3 — restriction, stage one.** The open-grant column in init's manifest;
|
||||
registry refuses ungranted opens. Test: a fixture process denied a protocol its
|
||||
neighbour is granted.
|
||||
|
||||
**P4 — protocol rebase.** Existing protocol modules re-expressed through
|
||||
`Define`; providers move onto the generated dispatch; `describe`/`enumerate`
|
||||
answered everywhere; the conformance test fixture exercises the reserved verbs
|
||||
against every registered provider.
|
||||
|
||||
**P5 — restriction, stage two.** Spawn's initial capability; namespace views
|
||||
built by the spawner; ambient resolution of `/protocol` retired. Includes the
|
||||
two requirements the microphone example pins: **parked replies** (a
|
||||
supervisor parks an open, keeps serving, replies later by badge) and
|
||||
**dedicated, killable granted channels** (revocation = channel death →
|
||||
`-EPEER`). Scoped separately — it touches `spawn`, the loader contract, and
|
||||
every supervisor — and lands together with the file-path half of namespacing.
|
||||
|
||||
The unix-path migration ([file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md#migration))
|
||||
is independent of P1–P5 and can land before or after.
|
||||
@@ -0,0 +1,131 @@
|
||||
# SMEP and SMAP — supervisor-mode hardening
|
||||
|
||||
*Design, 2026-07-31. Not yet implemented. Companion to
|
||||
[protocol-namespace.md](protocol-namespace.md) on the security track — this is
|
||||
the hardware half; that is the namespace half.*
|
||||
|
||||
Two CR4 bits that make the CPU refuse the two things a kernel should never do
|
||||
with user memory:
|
||||
|
||||
- **SMEP** (Supervisor Mode Execution Prevention, CR4 bit 20): instruction
|
||||
fetch in ring 0 from a page whose U/S bit says *user* → #PF. Kills the
|
||||
classic ret2usr exploit shape — a kernel bug that redirects control flow
|
||||
can no longer land in attacker-prepared user code.
|
||||
- **SMAP** (Supervisor Mode Access Prevention, CR4 bit 21): data read/write
|
||||
in ring 0 to a user page → #PF, unless `EFLAGS.AC` is set. `stac`/`clac`
|
||||
open and close deliberate access windows; danos's design needs no windows
|
||||
at all (below).
|
||||
|
||||
Detection is CPUID leaf 7, subleaf 0, EBX bit 7 (SMEP) and bit 20 (SMAP).
|
||||
Both bits are per-core state: the BSP and every AP must set them.
|
||||
|
||||
## Why, in danos terms
|
||||
|
||||
Every syscall argument is an attacker-controlled integer, and several take
|
||||
pointers. A kernel bug that dereferences a crafted pointer reads, writes, or
|
||||
executes memory of the attacker's choosing — the exact bug class the
|
||||
isolation tracks exist to prevent. SMEP/SMAP turn that class from "silent
|
||||
compromise" into "immediate, attributable #PF with a kernel RIP in the log."
|
||||
|
||||
The second benefit matters as much as the first: **SMAP is a permanent
|
||||
tripwire.** Once it is on, any *future* syscall that touches user memory
|
||||
directly — instead of going through the checked copy layer — faults the
|
||||
first time the QEMU suite runs it. The discipline stops depending on review.
|
||||
|
||||
## Where danos already stands
|
||||
|
||||
The design is closer than it looks, because the IPC layer was built right:
|
||||
|
||||
- **The copy layer is already SMAP-proof.** `copyAcross` and `copyFromUser`
|
||||
(`system/kernel/ipc-synchronous.zig:305,333`) never dereference a user
|
||||
virtual address: they walk the page tables and move bytes through the
|
||||
physmap — kernel mappings throughout. SMAP cannot object.
|
||||
- **Syscall entry already clears AC.** `SFMASK = 0x4_0700` clears IF, TF,
|
||||
DF, **AC** on every `syscall`
|
||||
(`system/kernel/architecture/x86_64/per-cpu.zig:76`). The syscall path is
|
||||
SMAP-clean from day one.
|
||||
- **The interrupt path is not.** Hardware does *not* clear AC on IDT
|
||||
delivery, and ring 3 can set AC with `popfq` — so a hostile process could
|
||||
take an interrupt with AC=1 and have the handler run with SMAP suspended.
|
||||
`isr_common` (`system/kernel/architecture/x86_64/isr.s:366`) needs a
|
||||
`clac` beside its `swapgs`.
|
||||
- **CR4 today:** the BSP inherits firmware CR4 (no kernel write anywhere);
|
||||
APs set PAE/OSFXSR/OSXMMEXCPT in `trampoline.s:62-68`. Neither path sets
|
||||
SMEP/SMAP yet, and both must.
|
||||
- **The stragglers.** A handful of syscalls still dereference user pointers
|
||||
raw after a bounds check — every one is a SMAP #PF waiting to happen, and
|
||||
every one is *already* a latent kernel fault today (an unmapped-but-in-
|
||||
range user page oopses the kernel instead of failing the call). The audit
|
||||
list, from the 2026-07-31 survey of `system/kernel/process.zig`:
|
||||
|
||||
| Syscall | Raw access |
|
||||
|---|---|
|
||||
| `system_spawn` | name + argument blob (`:972`, `:980`) |
|
||||
| `fs_resolve` | path in, result out (`:1780`, `:1797`) |
|
||||
| `fs_mount` / `fs_unmount` | prefix + rewrite strings (`:1864`) |
|
||||
| `fs_node` | read buffer out (`:1820`) |
|
||||
| `debug_write` | message bytes (`:1700`) |
|
||||
| `klog_read` | log bytes out (`:1740`) |
|
||||
| `process_enumerate` | descriptor array out (`:1132`) |
|
||||
| `device_enumerate` | descriptor array out (`:388`) |
|
||||
|
||||
(Some paths already do it right — the futex word and the device-register
|
||||
descriptor go through `copyFromUser` (`:1087`, `:924`). The write
|
||||
direction has no helper yet.)
|
||||
|
||||
- **One known gap inside the copy layer itself:** the walk checks presence,
|
||||
not the leaf U/S and writable bits (`ipc-synchronous.zig:20-22` flags
|
||||
this). Today that is nearly moot — the user half contains only mappings
|
||||
the kernel itself created for that process — but it must close before
|
||||
shared or copy-on-write mappings exist, and closing it is part of making
|
||||
the copy layer the single trusted door.
|
||||
|
||||
## The plan
|
||||
|
||||
**H1 — copy discipline (the real work).** A `user-memory` kernel module:
|
||||
`copyFromUser` / `copyToUser` (the missing write direction) via the physmap
|
||||
walk, with U/S and writable leaf checks closing the in-tree TODO. Convert
|
||||
the eight stragglers. This fixes the latent unmapped-page kernel fault on
|
||||
its own — it is worth doing even if SMEP/SMAP never shipped. QEMU suite
|
||||
green; no behavior change visible to correct programs.
|
||||
|
||||
**H2 — SMEP.** A leaf-7 feature probe (the kernel has per-leaf `cpuid`
|
||||
helpers in `apic.zig` to generalize); set CR4.SMEP during per-CPU bring-up
|
||||
on BSP and APs — prefer the Zig-side per-CPU init over the trampoline
|
||||
assembly, so one code path covers every core and the trampoline stays
|
||||
minimal. Audit first that ring 0 never executes user-mapped pages: kernel
|
||||
text lives in the kernel half, `jump_to_user` is kernel code, and the AP
|
||||
trampoline page is kernel-mapped — expected clean, verify before flipping.
|
||||
|
||||
**H3 — SMAP.** Add `clac` at `isr_common` entry. `clac` is #UD on CPUs
|
||||
without SMAP, so the instruction is a 3-byte NOP in the image, patched to
|
||||
`clac` at boot when CPUID advertises SMAP (one-time patch beats a
|
||||
conditional branch in the hottest path in the kernel). Then set CR4.SMAP in
|
||||
the same per-CPU init. From this point the whole QEMU suite doubles as the
|
||||
enforcement test: any missed raw dereference is a vector-14 with a kernel
|
||||
RIP and a user CR2 — loud and attributable.
|
||||
|
||||
**H4 — keep it honest.** A line in the coding standards: kernel code
|
||||
touches user memory only through `user-memory`; there is no `stac` anywhere
|
||||
in the tree, and a PR that adds one is wrong by definition. SMAP enforces
|
||||
the rule mechanically at test time.
|
||||
|
||||
Feature-gating follows the timekeeping rule (work on any VM, real Intel,
|
||||
real AMD): both bits are probed, absence is logged and tolerated — like the
|
||||
IOMMU's fail-open, the machine still boots, just unhardened. QEMU: TCG
|
||||
implements both; KVM inherits the host (Intel Ivy Bridge+ for SMEP,
|
||||
Broadwell+ for SMAP; AMD Zen+ for both). The test images should run with
|
||||
`-cpu max` so the suite always exercises the enabled paths.
|
||||
|
||||
## Adjacent, deliberately separate
|
||||
|
||||
- **SYSRET canonical-RIP hardening** (`isr.s:192-194` documents it): a
|
||||
non-canonical return RIP makes `sysretq` #GP *in ring 0* on Intel. Same
|
||||
hardening bucket, independent fix (validate RCX before `sysretq`, fall
|
||||
back to `iretq`), should ride the same branch as H2/H3 but is not
|
||||
SMEP/SMAP.
|
||||
- **KPTI / Meltdown-class leaks are out of scope.** SMEP/SMAP police
|
||||
architectural accesses, not speculative ones. danos runs one kernel
|
||||
mapping in every address space and accepts that on affected hardware;
|
||||
revisit only if the threat model ever includes hostile native code on
|
||||
shared machines.
|
||||
@@ -11,7 +11,7 @@ sequential pass and hands the bytes to the kernel unmodified.
|
||||
|
||||
The capsule is a *performance artifact*, not a source of truth. The boot
|
||||
volume's `/system` and `/test` file trees remain the canonical layout (see
|
||||
[danos-file-system-hierarchy-FSH.md](../file-system-development/danos-file-system-hierarchy-FSH.md));
|
||||
[file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md));
|
||||
the capsule is a pre-baked snapshot of the same binaries, derived from the same
|
||||
build graph, so the running system is identical whether the loader read the
|
||||
capsule or walked the tree.
|
||||
@@ -36,11 +36,11 @@ so it need be no fancier. Little-endian throughout:
|
||||
|
||||
```
|
||||
Header magic: u32 = "DNR2" (0x32524E44), count: u32
|
||||
Entry × count name: [64]u8 (NUL-padded FHS path), offset: u64, len: u64
|
||||
Entry × count name: [64]u8 (NUL-padded hierarchy path), offset: u64, len: u64
|
||||
blobs... each entry's file bytes, at its offset within the image
|
||||
```
|
||||
|
||||
- **Names are full FHS paths** (`/system/services/init`), not basenames — that
|
||||
- **Names are full hierarchy paths** (`/system/services/init`), not basenames — that
|
||||
is what "v2" means. The 64-byte capacity matches `abi.maximum_process_name`,
|
||||
so a task named after its binary path is never truncated. Paths longer than
|
||||
63 bytes are a build error (`pack-system-image.py` rejects them).
|
||||
@@ -54,14 +54,14 @@ blobs... each entry's file bytes, at its offset within the image
|
||||
|
||||
## How it is built
|
||||
|
||||
`build.zig` maintains one `bundled` list — every user binary and its FHS home.
|
||||
`build.zig` maintains one `bundled` list — every user binary and its hierarchy home.
|
||||
Three artifacts are derived from that same list, in the same build graph, so
|
||||
they cannot drift apart:
|
||||
|
||||
1. **The tree**: each binary installed at its FHS path (`zig-out/system/...`
|
||||
1. **The tree**: each binary installed at its hierarchy path (`zig-out/system/...`
|
||||
and `zig-out/test/...`, mirrored onto the FAT boot volume by
|
||||
`tools/make-fat-image.py`).
|
||||
2. **The manifest** (`system/manifest`): the FHS path of every bundled binary,
|
||||
2. **The manifest** (`system/manifest`): the hierarchy path of every bundled binary,
|
||||
one per line — the loader's per-file fallback input.
|
||||
3. **The capsule**: `tools/pack-system-image.py` packs the same binaries into
|
||||
the v2 container, installed at `zig-out/boot/system.img` and placed on the
|
||||
@@ -103,7 +103,7 @@ the kernel (`kernel.zig`) then publishes the same bytes twice, to two
|
||||
consumers:
|
||||
|
||||
- **The process layer** (`process.zig`): `system_spawn` looks binaries up in
|
||||
the ramdisk via `Reader.find` — exact FHS path, or unique basename for
|
||||
the ramdisk via `Reader.find` — exact hierarchy path, or unique basename for
|
||||
pre-path callers — and loads them as fresh ring-3 processes. The stored path
|
||||
becomes the task's name.
|
||||
- **The VFS root** (`vfs.zig`, `setInitialRamdisk`): the image is mounted as
|
||||
@@ -111,7 +111,7 @@ consumers:
|
||||
paths, so `/system` and, when the fixtures are bundled, `/test`. Directory
|
||||
nodes are derived from the entry paths (the unique parents), so the trees
|
||||
are listable and their files readable over the normal VFS protocol — the
|
||||
FHS boot tree every process sees comes straight out of the capsule bytes.
|
||||
boot tree every process sees comes straight out of the capsule bytes.
|
||||
|
||||
The image is never copied after the handoff and never mutated: the initrd is
|
||||
immutable, which is what makes the VFS's node serving lock-free.
|
||||
|
||||
@@ -82,7 +82,7 @@ return real answers.
|
||||
- `--disable-shared`; static `Modules/Setup`: `posix errno _io _codecs _weakref
|
||||
time math _stat _collections itertools _functools _locale _sre` plus what the
|
||||
interpreter core insists on; threadless build (WASI precedent).
|
||||
- `Lib/` on the FAT image under the FSH (e.g. `/system/python/lib`);
|
||||
- `Lib/` on the FAT image under the hierarchy (e.g. `/system/python/lib`);
|
||||
`PYTHONHOME` set accordingly; `.pyc` written with **checked-hash
|
||||
invalidation** (FAT's 2-second mtime granularity makes mtime-based validation
|
||||
lie during fast edit-run cycles).
|
||||
|
||||
@@ -257,5 +257,5 @@ in practice cares.
|
||||
the spawn-argv work extends.
|
||||
- [device-driver-development/ipc.md](device-driver-development/ipc.md) — the IPC
|
||||
surface Python services speak.
|
||||
- [file-system-development/danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md)
|
||||
- [file-system-development/file-system-hierarchy.md](file-system-development/file-system-hierarchy.md)
|
||||
— where `Lib/` and `site-packages` land on the image.
|
||||
|
||||
@@ -347,7 +347,7 @@ Two current decisions fall out of this roadmap:
|
||||
- [syscall.md](os-development/syscall.md) — the kernel↔runtime ABI `runtime.os` is built on.
|
||||
- [sysv.md](os-development/sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs.
|
||||
- [ipc.md](device-driver-development/ipc.md) — the IPC the VFS/FAT operations travel over.
|
||||
- [danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md) — the
|
||||
- [file-system-hierarchy.md](file-system-development/file-system-hierarchy.md) — the
|
||||
filesystem layout the file surface serves.
|
||||
- [coding-standards.md](coding-standards.md) — danos naming (why the compat spellings
|
||||
are confined, and now retired).
|
||||
|
||||
@@ -33,8 +33,8 @@ pub const Operation = enum(u32) {
|
||||
rename, // rename(old\0new payload) -> status
|
||||
};
|
||||
|
||||
/// The type of a filesystem node, aligned to the FSH file-type table
|
||||
/// (docs/danos-file-system-hierarchy-FSH.md). Fills `FileStatus.kind` and
|
||||
/// The type of a filesystem node, aligned to the node-kind table
|
||||
/// (docs/file-system-development/file-system-hierarchy.md). Fills `FileStatus.kind` and
|
||||
/// `DirectoryEntry.kind`; `regular = 0` keeps the historical hardcoded value.
|
||||
pub const NodeKind = enum(u32) {
|
||||
regular = 0,
|
||||
|
||||
Reference in New Issue
Block a user