Files
danos/docs/device-driver-development/ipc.md
T
Daniel Samson 1379b699f3 init: /protocol replaces the ServiceId registry
A protocol is reached by name now, not by a compile-time integer. Init is
PID 1 and already knows which binary it started, so init serves /protocol
as a vfs backend: bind claims a contract with the provider's endpoint
attached, open answers with that endpoint as the reply's capability, and
readdir lists what is bound with the task and binary behind it. The kernel
reserves the prefix — nothing may mount over it, under it, or unmount it —
and ServiceId, ipc_register and ipc_lookup are gone, their syscall numbers
left vacant.

A bind is authorized by who the caller *is*: the kernel-stamped binary
together with the supervising task's identity, matched against
/system/configuration/protocol.csv. Identity, not spelling — spawn is
ungated, so an attacker can run any bundled binary, and a name-only rule
would have let it launder grants through an init of its own making. A name
a live process holds is refused to everyone else; a dead one's is released.

Three review rounds against a hostile ring-3 process found what 108 green
tests could not, because the suite contains no attacker. Publishing init's
supervision endpoint as the registry put PID 1's mailbox in every process's
hands, where two forged bytes reached the shutdown path: privileged traffic
is now believed only from the task that holds the contract it speaks for.
A capability arriving on a request outlived every path that ignored it,
one handle per call until the table was full — in init, and in the harness
ten services share — so the arriving capability is owned by the turn and
released unless a handler says otherwise. And the kernel let anyone holding
an endpoint handle aim signals, timers, exit notices and interrupts at it:
binding now requires having created it.

Suite 108/108. The new protocol-registry case asserts eleven properties,
each one an attack that must fail.
2026-08-01 02:39:07 +01:00

200 lines
12 KiB
Markdown

# IPC: the kernel-ipc transport
Inter-process communication is the **backbone of a microkernel**. Once drivers and
services run isolated in their own address spaces ([vision](../vision.md)), they can't
just call each other — a request becomes bytes on a wire. In a microkernel, whatever
was a function call across a monolithic kernel is IPC, so it's a first-class
concern, not an afterthought.
This document describes **one transport** — the bottom layer (L0) of the
communication stack defined in
[communication.md](../os-development/communication.md), which owns the model
and the vocabulary (*protocol*, *channel*, *packet*, *signal*, *endpoint*).
kernel-ipc is the **first** transport, not the only possible one: in
buffer-plus-doorbell terms it is a kernel-owned mailbox with the scheduler as
the doorbell. Its distinguishing properties, which the layers above may rely
on where they say so:
- **Rendezvous.** A call is a synchronous meeting, copied sender-page to
receiver-page — natural backpressure, no queue to size.
- **Capability carriage.** The *only* transport that can move a handle
between processes. Channels are therefore always established over
kernel-ipc, and it remains every channel's control path even when bulk
data is negotiated onto a fatter transport (a shared-memory ring).
- **Verified source.** Every delivery carries the kernel-stamped badge — the
identity the channel layer attaches to received packets.
- **Bounded packets.** 256 bytes call/reply, 64 pushed — the floor every
protocol may assume on any transport.
Three properties keep the networking analogy honest — kernel-ipc is
networking-*shaped*, not TCP:
- **Channels over it are RPC-shaped, not streams.** Packets, call/reply,
datagram pushes — closer to UDP plus RPC than to a byte stream. Ordering
exists per exchange (a reply answers its call), not across a channel.
- **Possession is the connection.** There is no handshake state in the
kernel: holding the capability *is* having the channel. A provider's one
endpoint terminates every client's channel at once, demultiplexed by badge
— like every client sharing the server's listening socket, with
per-connection state living in the provider, keyed by badge. A *private*
channel (a dedicated endpoint pair) is built when wanted: that is exactly
what `subscribe` does.
- **Packets never fragment.** If it doesn't fit in a packet, it isn't a
packet: bulk data lives in shared memory and a packet (or signal) is the
doorbell. The display path already works this way.
The rest of this document is the implementation, bottom-up: the kernel-thread
queue the blocking discipline was worked out on, then endpoints — this
transport's termination points.
## The kernel-thread queue
The first form is a **bounded blocking queue** (`system/kernel/ipc.zig`): a
fixed-size ring buffer of messages with a producer/consumer rendezvous, built
on the scheduler's [wait queues](../os-development/scheduling.md). (Its type
is still named `Channel(T, capacity)` — it predates the vocabulary above, and
is a *queue between kernel threads in one address space*, not a channel in
the model's sense; a rename can ride a later flag-day.)
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
ring buffer, a count, and two wait queues:
- **`send(msg)`** — if the queue is full, block on the *not-full* queue; otherwise
write the message, bump the count, and wake a waiting receiver.
- **`receive()`** — if the queue is empty, block on the *not-empty* queue; otherwise
take a message, drop the count, and wake a waiting sender.
Neither side busy-waits: a full queue parks the sender, an empty one parks the
receiver, and each operation wakes the other side when it makes progress possible.
Two details make it correct:
- **Recheck in a loop.** A woken task re-tests the condition (`while (full) wait`)
rather than assuming the slot is still available — another waiter may have taken
it first. This is the standard guard against spurious or racing wakeups.
- **One critical section.** `send`/`receive` run under the [big kernel
lock](../os-development/smp.md) (`sync.enter` / `sync.leave`), which disables interrupts on this
core *and* takes the kernel's one spinlock — since SMP, the interrupt flag alone
is not atomicity, because `cli` on one core does nothing to another. So checking
the condition and committing the block/enqueue happen atomically both with respect
to the timer preempting mid-operation and to the other side running on another
CPU. `waitLocked` / `wakeLocked` are the variants that assume the caller already
holds that critical section.
### Verifying it
The `ipc` test (see [testing.md](../testing.md)) runs a producer and a consumer passing
**100 messages through a 4-slot queue**. The small buffer means the queue goes
full and empty over and over, so both the blocking-send and blocking-receive paths are
exercised heavily. The messages arrive intact and in order (their sum is the
expected `5050`), and neither task busy-waits — they block and wake each other.
## Endpoints: the termination points
A queue connects two kernel threads sharing one address space. Real providers are
*processes*, so a packet has to cross an address-space boundary. That's
`system/kernel/ipc-synchronous.zig`, and its shape is L4's: a synchronous **rendezvous** at an
`Endpoint`, with the packet copied directly from the sender's pages to the receiver's
(`copyAcross` walks both sets of page tables through the physmap — no CR3 switch, no
bounce buffer).
Two syscalls carry the request/reply exchange:
- **`ipc_call(h, msg, reply)`** — copy the request packet to the provider, block
until the reply packet comes back.
- **`ipc_reply_wait(h, reply, recv)`** — reply to the client you're still holding (if
any), then block for the next request. One syscall, because a provider's steady state
is *always* "finish the last one, wait for the next".
An endpoint is reached by **handle** — a small integer index into the process's handle
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable.
The provider never learns the client's identity beyond the **badge** delivered
alongside each packet: the caller's task id, stamped by the kernel —
unforgeable source addressing, a property a network's source field lacks.
The bootstrap problem — how a channel is first established — is the subject of
[protocol-namespace.md](../os-development/protocol-namespace.md): a protocol is
resolved by name and the channel arrives as a capability. (The mechanism it
replaced — `ipc_register`/`ipc_lookup` under compile-time `ServiceId` integers —
is gone: both syscalls and the enum were deleted when the registry landed, and
their syscall numbers are left vacant.)
### Interrupts are signals
`notifyFromIsr` posts an *asynchronous* signal to an endpoint — no payload, no
reply owed — and wakes whoever is blocked in `reply_wait`. Its badge has the top bit
set (`notify_badge_bit`), which is how a driver's single event loop distinguishes "a
client wants something" from "the hardware wants something". Signals sit in a
small coalescing ring on the endpoint, so an interrupt taken while the driver was busy
elsewhere is not lost — coalesced, never dropped, which is exactly a signal's
contract (the *count* may collapse; the *fact* may not).
This is what makes a user-space driver possible at all, and it's the subject of
[drivers.md](drivers.md).
## What's next (partly done since)
- **Priority inheritance** through IPC — still open: a high-priority client
blocked on a low-priority provider suffers unbounded priority inversion.
- **Handle transfer.** *Landed as cap-passing (M13)*: `ipc_call` and
`ipc_reply_wait` carry an optional capability alongside the bytes (`send_cap`),
copying an endpoint or shared-memory handle into the peer's table — the
mechanism by which channels are established and private channels built. First
user: [input](input.md) subscribers register by handing over their own
endpoint, and class drivers get a private channel to one device.
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
shape (logging, event fan-out). *Landed as `ipc_send`* — a
non-blocking post of an event packet (≤ 64 bytes) to an endpoint's bounded
queue, delivered through `reply_wait` (badge bit `notify_message_bit`). Built
for, and first used by, the [input service](input.md)'s keyboard-event
broadcast, where a synchronous push would let one dead subscriber hang the
fan-out. A full queue drops the oldest — event packets are droppable by
design ([protocol-namespace.md](../os-development/protocol-namespace.md)'s
wiring section states the rule).
- **A bounded reply** — half landed. The copy is still one packet
(256 bytes) under the big kernel lock, but bulk transfer got its shared
pages: `shared_memory_create`/`map`/`physical`, the region handle delegated as
a capability (above) — the packets-never-fragment rule in practice.
virtio-gpu's scanout surface is the first user
([display-v2.md](display-v2.md)).
## Lifecycle conventions over IPC (M17)
Three conventions from [process-lifecycle.md](../os-development/process-lifecycle.md) ride the
signal mechanism:
- **Process signals** arrive as endpoint signals on the endpoint a process
nominated with `signal_bind` (`process.bindSignals`): badge = the signal bit
plus the coalesced pending mask (`process.signalsFrom` decodes). Statements,
never questions; no payload, no reply.
- **One-shot timers** (`timer_bind`, `time.timerOnce`) land as a
timer-bit signal — the timed wait: a service arms a deadline and keeps
serving, instead of blocking in sleep.
- **Kernel notifications go only to your own endpoint.** `signal_bind`,
`timer_bind`, `process_subscribe`, `irq_bind`, `msi_bind`, and spawn's exit
endpoint all *nominate where the kernel will speak*, and all of them refuse an
endpoint the caller did not create (`-EPERM`; the check is `ipc.ownedBy`,
normalized to the process, so any thread may nominate an endpoint a sibling
created). Holding a handle is not enough, because holding a handle is cheap:
`fs_resolve` installs a mounted backend's capability in *any* caller's table,
so every process holds a handle to PID 1's mailbox. Without the rule, "bind
init's endpoint, then signal yourself" is a genuine, kernel-stamped `terminate`
badge in PID 1's queue — a shutdown a receiver has no way to disbelieve — and
timers, which carry no identity at all, multiply any loop that re-arms on its
own landing.
- **A capability that arrives belongs to the turn.** The kernel installs a sent
capability in the receiver's table whatever the message's length or kind, so a
receive loop must dispose of one on *every* path — the ping, the notification,
the malformed request. The service harness (`service.run`) and PID 1 both hold
it in an `ipc.Arrival`, released by a `defer`, and a handler that means to keep
it says `take()`: forgetting closes, keeping is explicit. The reverse
arrangement leaks a handle-table slot per request, and thirty-two unauthorized
zero-length pings then end a service's ability to accept any capability —
no subscribe, no shared-memory handover — for the rest of the boot.
- **The universal ping**: a **zero-length request is the liveness probe**,
answered with a zero-length reply by the service harness itself
(`service.run`). No protocol's requests start at length zero, so the
encoding cannot collide, and a wedged service simply fails to answer — which
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
protocol packet.