docs: the communication stack — /protocol namespace, layered model, non-unix hierarchy, SMEP/SMAP plan
The security-track design set. communication.md is the model: four layers (namespace / protocol / channel / transport), packets and signals, parties addressed by the channel and objects by target, transports as replaceable buffer+doorbell mechanisms. protocol-namespace.md is L3+L2: /protocol names contracts, resolution establishes a channel, init is the registrar, restriction is per-process namespace delegation (with the microphone-prompt worked example), the envelope is the universal packet header, migration P1-P5 retires ServiceId. file-system-hierarchy.md replaces the unix-FHS spec with the danos-native tree (/applications, /protocol, /system, /volumes) and its migration table. ipc.md is re-cut as the kernel-ipc transport document. smep-smap.md designs the kernel hardening: copy discipline for the eight raw user-pointer syscalls, then SMEP, then SMAP as a permanent tripwire.
This commit is contained in:
@@ -1,34 +1,70 @@
|
||||
# IPC: message-passing channels
|
||||
# IPC: the kernel-ipc transport
|
||||
|
||||
Inter-process communication is the **backbone of a microkernel**. Once drivers and
|
||||
services run isolated in their own address spaces ([vision](../vision.md)), they can't
|
||||
just call each other — a request becomes a **message**. In a microkernel, whatever
|
||||
just call each other — a request becomes bytes on a wire. In a microkernel, whatever
|
||||
was a function call across a monolithic kernel is IPC, so it's a first-class
|
||||
concern, not an afterthought.
|
||||
|
||||
There are two layers, built a milestone apart:
|
||||
This document describes **one transport** — the bottom layer (L0) of the
|
||||
communication stack defined in
|
||||
[communication.md](../os-development/communication.md), which owns the model
|
||||
and the vocabulary (*protocol*, *channel*, *packet*, *signal*, *endpoint*).
|
||||
kernel-ipc is the **first** transport, not the only possible one: in
|
||||
buffer-plus-doorbell terms it is a kernel-owned mailbox with the scheduler as
|
||||
the doorbell. Its distinguishing properties, which the layers above may rely
|
||||
on where they say so:
|
||||
|
||||
- **`system/kernel/ipc.zig`** — a bounded blocking channel between *kernel threads*,
|
||||
described below. The primitive, and where the blocking discipline was worked out.
|
||||
- **`system/kernel/ipc-synchronous.zig`** — synchronous call/reply between *processes*, across
|
||||
address spaces. What user-space servers and drivers actually talk over. It's the
|
||||
second half of this document.
|
||||
- **Rendezvous.** A call is a synchronous meeting, copied sender-page to
|
||||
receiver-page — natural backpressure, no queue to size.
|
||||
- **Capability carriage.** The *only* transport that can move a handle
|
||||
between processes. Channels are therefore always established over
|
||||
kernel-ipc, and it remains every channel's control path even when bulk
|
||||
data is negotiated onto a fatter transport (a shared-memory ring).
|
||||
- **Verified source.** Every delivery carries the kernel-stamped badge — the
|
||||
identity the channel layer attaches to received packets.
|
||||
- **Bounded packets.** 256 bytes call/reply, 64 pushed — the floor every
|
||||
protocol may assume on any transport.
|
||||
|
||||
## The channel
|
||||
Three properties keep the networking analogy honest — kernel-ipc is
|
||||
networking-*shaped*, not TCP:
|
||||
|
||||
The first form is a **bounded blocking channel** (`system/kernel/ipc.zig`): a fixed-size
|
||||
ring buffer of messages with a producer/consumer rendezvous, built on the
|
||||
scheduler's [wait queues](../os-development/scheduling.md).
|
||||
- **Channels over it are RPC-shaped, not streams.** Packets, call/reply,
|
||||
datagram pushes — closer to UDP plus RPC than to a byte stream. Ordering
|
||||
exists per exchange (a reply answers its call), not across a channel.
|
||||
- **Possession is the connection.** There is no handshake state in the
|
||||
kernel: holding the capability *is* having the channel. A provider's one
|
||||
endpoint terminates every client's channel at once, demultiplexed by badge
|
||||
— like every client sharing the server's listening socket, with
|
||||
per-connection state living in the provider, keyed by badge. A *private*
|
||||
channel (a dedicated endpoint pair) is built when wanted: that is exactly
|
||||
what `subscribe` does.
|
||||
- **Packets never fragment.** If it doesn't fit in a packet, it isn't a
|
||||
packet: bulk data lives in shared memory and a packet (or signal) is the
|
||||
doorbell. The display path already works this way.
|
||||
|
||||
The rest of this document is the implementation, bottom-up: the kernel-thread
|
||||
queue the blocking discipline was worked out on, then endpoints — this
|
||||
transport's termination points.
|
||||
|
||||
## The kernel-thread queue
|
||||
|
||||
The first form is a **bounded blocking queue** (`system/kernel/ipc.zig`): a
|
||||
fixed-size ring buffer of messages with a producer/consumer rendezvous, built
|
||||
on the scheduler's [wait queues](../os-development/scheduling.md). (Its type
|
||||
is still named `Channel(T, capacity)` — it predates the vocabulary above, and
|
||||
is a *queue between kernel threads in one address space*, not a channel in
|
||||
the model's sense; a rename can ride a later flag-day.)
|
||||
|
||||
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
|
||||
ring buffer, a count, and two wait queues:
|
||||
|
||||
- **`send(msg)`** — if the channel is full, block on the *not-full* queue; otherwise
|
||||
- **`send(msg)`** — if the queue is full, block on the *not-full* queue; otherwise
|
||||
write the message, bump the count, and wake a waiting receiver.
|
||||
- **`receive()`** — if the channel is empty, block on the *not-empty* queue; otherwise
|
||||
- **`receive()`** — if the queue is empty, block on the *not-empty* queue; otherwise
|
||||
take a message, drop the count, and wake a waiting sender.
|
||||
|
||||
Neither side busy-waits: a full channel parks the sender, an empty one parks the
|
||||
Neither side busy-waits: a full queue parks the sender, an empty one parks the
|
||||
receiver, and each operation wakes the other side when it makes progress possible.
|
||||
|
||||
Two details make it correct:
|
||||
@@ -45,47 +81,52 @@ Two details make it correct:
|
||||
CPU. `waitLocked` / `wakeLocked` are the variants that assume the caller already
|
||||
holds that critical section.
|
||||
|
||||
## Verifying it
|
||||
### Verifying it
|
||||
|
||||
The `ipc` test (see [testing.md](../testing.md)) runs a producer and a consumer passing
|
||||
**100 messages through a 4-slot channel**. The small buffer means the channel goes
|
||||
**100 messages through a 4-slot queue**. The small buffer means the queue goes
|
||||
full and empty over and over, so both the blocking-send and blocking-receive paths are
|
||||
exercised heavily. The messages arrive intact and in order (their sum is the
|
||||
expected `5050`), and neither task busy-waits — they block and wake each other.
|
||||
|
||||
## Endpoints: call/reply across address spaces
|
||||
## Endpoints: the termination points
|
||||
|
||||
A channel connects two kernel threads sharing one address space. Real servers are
|
||||
*processes*, so the payload has to cross an address-space boundary. That's
|
||||
A queue connects two kernel threads sharing one address space. Real providers are
|
||||
*processes*, so a packet has to cross an address-space boundary. That's
|
||||
`system/kernel/ipc-synchronous.zig`, and its shape is L4's: a synchronous **rendezvous** at an
|
||||
`Endpoint`, with the message copied directly from the sender's pages to the receiver's
|
||||
`Endpoint`, with the packet copied directly from the sender's pages to the receiver's
|
||||
(`copyAcross` walks both sets of page tables through the physmap — no CR3 switch, no
|
||||
bounce buffer).
|
||||
|
||||
Two syscalls carry it:
|
||||
Two syscalls carry the request/reply exchange:
|
||||
|
||||
- **`ipc_call(h, msg, reply)`** — copy `msg` to the server, block until it replies.
|
||||
- **`ipc_call(h, msg, reply)`** — copy the request packet to the provider, block
|
||||
until the reply packet comes back.
|
||||
- **`ipc_reply_wait(h, reply, recv)`** — reply to the client you're still holding (if
|
||||
any), then block for the next request. One syscall, because a server's steady state
|
||||
any), then block for the next request. One syscall, because a provider's steady state
|
||||
is *always* "finish the last one, wait for the next".
|
||||
|
||||
An endpoint is reached by **handle** — a small integer index into the process's handle
|
||||
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable. The
|
||||
bootstrap problem (how do you get the first handle?) is solved by a tiny name registry:
|
||||
a server calls `ipc_register(service_id, h)` under a well-known small integer, and a
|
||||
client calls `ipc_lookup(service_id)`.
|
||||
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable.
|
||||
The provider never learns the client's identity beyond the **badge** delivered
|
||||
alongside each packet: the caller's task id, stamped by the kernel —
|
||||
unforgeable source addressing, a property a network's source field lacks.
|
||||
|
||||
The server never learns the client's identity beyond a **badge**, delivered alongside
|
||||
the message: the caller's task id.
|
||||
The bootstrap problem — how a channel is first established — is the subject of
|
||||
[protocol-namespace.md](../os-development/protocol-namespace.md): a protocol is
|
||||
resolved by name and the channel arrives as a capability. (The mechanism this
|
||||
replaces, `ipc_register`/`ipc_lookup` under compile-time `ServiceId` integers,
|
||||
is retired by that design.)
|
||||
|
||||
### Interrupts are messages too
|
||||
### Interrupts are signals
|
||||
|
||||
`notifyFromIsr` posts an *asynchronous* notification to an endpoint — no payload, no
|
||||
`notifyFromIsr` posts an *asynchronous* signal to an endpoint — no payload, no
|
||||
reply owed — and wakes whoever is blocked in `reply_wait`. Its badge has the top bit
|
||||
set (`notify_badge_bit`), which is how a driver's single event loop distinguishes "a
|
||||
client wants something" from "the hardware wants something". Notifications sit in a
|
||||
client wants something" from "the hardware wants something". Signals sit in a
|
||||
small coalescing ring on the endpoint, so an interrupt taken while the driver was busy
|
||||
elsewhere is not lost.
|
||||
elsewhere is not lost — coalesced, never dropped, which is exactly a signal's
|
||||
contract (the *count* may collapse; the *fact* may not).
|
||||
|
||||
This is what makes a user-space driver possible at all, and it's the subject of
|
||||
[drivers.md](drivers.md).
|
||||
@@ -93,40 +134,44 @@ This is what makes a user-space driver possible at all, and it's the subject of
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Priority inheritance** through IPC — still open: a high-priority client
|
||||
blocked on a low-priority server suffers unbounded priority inversion.
|
||||
blocked on a low-priority provider suffers unbounded priority inversion.
|
||||
- **Handle transfer.** *Landed as cap-passing (M13)*: `ipc_call` and
|
||||
`ipc_reply_wait` carry an optional capability alongside the bytes (`send_cap`),
|
||||
copying an endpoint or shared-memory handle into the peer's table. First user:
|
||||
[input](input.md) subscribers register by handing over their own endpoint, and
|
||||
class drivers get a private channel to one device.
|
||||
copying an endpoint or shared-memory handle into the peer's table — the
|
||||
mechanism by which channels are established and private channels built. First
|
||||
user: [input](input.md) subscribers register by handing over their own
|
||||
endpoint, and class drivers get a private channel to one device.
|
||||
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
|
||||
shape (logging, notifications between servers). *Landed as `ipc_send`* — a
|
||||
non-blocking post to an endpoint's bounded payload queue, delivered through
|
||||
`reply_wait` as a buffered message (badge bit `notify_message_bit`). Built for, and
|
||||
first used by, the [input service](input.md)'s keyboard-event broadcast, where a
|
||||
synchronous push would let one dead subscriber hang the fan-out. A full queue drops
|
||||
the oldest (discrete messages, not a coalescing level like the notification ring).
|
||||
- **A bounded reply** — half landed. The copy is still 256 bytes
|
||||
(`MESSAGE_MAXIMUM`) under the big kernel lock, but bulk transfer got its shared
|
||||
shape (logging, event fan-out). *Landed as `ipc_send`* — a
|
||||
non-blocking post of an event packet (≤ 64 bytes) to an endpoint's bounded
|
||||
queue, delivered through `reply_wait` (badge bit `notify_message_bit`). Built
|
||||
for, and first used by, the [input service](input.md)'s keyboard-event
|
||||
broadcast, where a synchronous push would let one dead subscriber hang the
|
||||
fan-out. A full queue drops the oldest — event packets are droppable by
|
||||
design ([protocol-namespace.md](../os-development/protocol-namespace.md)'s
|
||||
wiring section states the rule).
|
||||
- **A bounded reply** — half landed. The copy is still one packet
|
||||
(256 bytes) under the big kernel lock, but bulk transfer got its shared
|
||||
pages: `shared_memory_create`/`map`/`physical`, the region handle delegated as
|
||||
a capability (above). virtio-gpu's scanout surface is the first user
|
||||
a capability (above) — the packets-never-fragment rule in practice.
|
||||
virtio-gpu's scanout surface is the first user
|
||||
([display-v2.md](display-v2.md)).
|
||||
|
||||
## Lifecycle conventions over IPC (M17)
|
||||
|
||||
Three conventions from [process-lifecycle.md](../os-development/process-lifecycle.md) ride the
|
||||
notification mechanism:
|
||||
signal mechanism:
|
||||
|
||||
- **Signals** arrive as notifications on the endpoint a process nominated with
|
||||
`signal_bind` (`process.bindSignals`): badge = the signal bit plus the
|
||||
coalesced pending mask (`process.signalsFrom` decodes). Statements,
|
||||
- **Process signals** arrive as endpoint signals on the endpoint a process
|
||||
nominated with `signal_bind` (`process.bindSignals`): badge = the signal bit
|
||||
plus the coalesced pending mask (`process.signalsFrom` decodes). Statements,
|
||||
never questions; no payload, no reply.
|
||||
- **One-shot timers** (`timer_bind`, `time.timerOnce`) land as a
|
||||
timer-bit notification — the timed wait: a service arms a deadline and keeps
|
||||
timer-bit signal — the timed wait: a service arms a deadline and keeps
|
||||
serving, instead of blocking in sleep.
|
||||
- **The universal ping**: a **zero-length request is the liveness probe**,
|
||||
answered with a zero-length reply by the service harness itself
|
||||
(`service.run`). No protocol's requests start at length zero, so the
|
||||
encoding cannot collide, and a wedged service simply fails to answer — which
|
||||
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
|
||||
protocol message.
|
||||
protocol packet.
|
||||
|
||||
Reference in New Issue
Block a user