Files
danos/docs/device-driver-development/ipc.md
T
Daniel Samson 1379b699f3 init: /protocol replaces the ServiceId registry
A protocol is reached by name now, not by a compile-time integer. Init is
PID 1 and already knows which binary it started, so init serves /protocol
as a vfs backend: bind claims a contract with the provider's endpoint
attached, open answers with that endpoint as the reply's capability, and
readdir lists what is bound with the task and binary behind it. The kernel
reserves the prefix — nothing may mount over it, under it, or unmount it —
and ServiceId, ipc_register and ipc_lookup are gone, their syscall numbers
left vacant.

A bind is authorized by who the caller *is*: the kernel-stamped binary
together with the supervising task's identity, matched against
/system/configuration/protocol.csv. Identity, not spelling — spawn is
ungated, so an attacker can run any bundled binary, and a name-only rule
would have let it launder grants through an init of its own making. A name
a live process holds is refused to everyone else; a dead one's is released.

Three review rounds against a hostile ring-3 process found what 108 green
tests could not, because the suite contains no attacker. Publishing init's
supervision endpoint as the registry put PID 1's mailbox in every process's
hands, where two forged bytes reached the shutdown path: privileged traffic
is now believed only from the task that holds the contract it speaks for.
A capability arriving on a request outlived every path that ignored it,
one handle per call until the table was full — in init, and in the harness
ten services share — so the arriving capability is owned by the turn and
released unless a handler says otherwise. And the kernel let anyone holding
an endpoint handle aim signals, timers, exit notices and interrupts at it:
binding now requires having created it.

Suite 108/108. The new protocol-registry case asserts eleven properties,
each one an attack that must fail.
2026-08-01 02:39:07 +01:00

12 KiB

IPC: the kernel-ipc transport

Inter-process communication is the backbone of a microkernel. Once drivers and services run isolated in their own address spaces (vision), they can't just call each other — a request becomes bytes on a wire. In a microkernel, whatever was a function call across a monolithic kernel is IPC, so it's a first-class concern, not an afterthought.

This document describes one transport — the bottom layer (L0) of the communication stack defined in communication.md, which owns the model and the vocabulary (protocol, channel, packet, signal, endpoint). kernel-ipc is the first transport, not the only possible one: in buffer-plus-doorbell terms it is a kernel-owned mailbox with the scheduler as the doorbell. Its distinguishing properties, which the layers above may rely on where they say so:

  • Rendezvous. A call is a synchronous meeting, copied sender-page to receiver-page — natural backpressure, no queue to size.
  • Capability carriage. The only transport that can move a handle between processes. Channels are therefore always established over kernel-ipc, and it remains every channel's control path even when bulk data is negotiated onto a fatter transport (a shared-memory ring).
  • Verified source. Every delivery carries the kernel-stamped badge — the identity the channel layer attaches to received packets.
  • Bounded packets. 256 bytes call/reply, 64 pushed — the floor every protocol may assume on any transport.

Three properties keep the networking analogy honest — kernel-ipc is networking-shaped, not TCP:

  • Channels over it are RPC-shaped, not streams. Packets, call/reply, datagram pushes — closer to UDP plus RPC than to a byte stream. Ordering exists per exchange (a reply answers its call), not across a channel.
  • Possession is the connection. There is no handshake state in the kernel: holding the capability is having the channel. A provider's one endpoint terminates every client's channel at once, demultiplexed by badge — like every client sharing the server's listening socket, with per-connection state living in the provider, keyed by badge. A private channel (a dedicated endpoint pair) is built when wanted: that is exactly what subscribe does.
  • Packets never fragment. If it doesn't fit in a packet, it isn't a packet: bulk data lives in shared memory and a packet (or signal) is the doorbell. The display path already works this way.

The rest of this document is the implementation, bottom-up: the kernel-thread queue the blocking discipline was worked out on, then endpoints — this transport's termination points.

The kernel-thread queue

The first form is a bounded blocking queue (system/kernel/ipc.zig): a fixed-size ring buffer of messages with a producer/consumer rendezvous, built on the scheduler's wait queues. (Its type is still named Channel(T, capacity) — it predates the vocabulary above, and is a queue between kernel threads in one address space, not a channel in the model's sense; a rename can ride a later flag-day.)

Channel(T, capacity) is generic over the message type and buffer size. It holds a ring buffer, a count, and two wait queues:

  • send(msg) — if the queue is full, block on the not-full queue; otherwise write the message, bump the count, and wake a waiting receiver.
  • receive() — if the queue is empty, block on the not-empty queue; otherwise take a message, drop the count, and wake a waiting sender.

Neither side busy-waits: a full queue parks the sender, an empty one parks the receiver, and each operation wakes the other side when it makes progress possible.

Two details make it correct:

  • Recheck in a loop. A woken task re-tests the condition (while (full) wait) rather than assuming the slot is still available — another waiter may have taken it first. This is the standard guard against spurious or racing wakeups.
  • One critical section. send/receive run under the big kernel lock (sync.enter / sync.leave), which disables interrupts on this core and takes the kernel's one spinlock — since SMP, the interrupt flag alone is not atomicity, because cli on one core does nothing to another. So checking the condition and committing the block/enqueue happen atomically both with respect to the timer preempting mid-operation and to the other side running on another CPU. waitLocked / wakeLocked are the variants that assume the caller already holds that critical section.

Verifying it

The ipc test (see testing.md) runs a producer and a consumer passing 100 messages through a 4-slot queue. The small buffer means the queue goes full and empty over and over, so both the blocking-send and blocking-receive paths are exercised heavily. The messages arrive intact and in order (their sum is the expected 5050), and neither task busy-waits — they block and wake each other.

Endpoints: the termination points

A queue connects two kernel threads sharing one address space. Real providers are processes, so a packet has to cross an address-space boundary. That's system/kernel/ipc-synchronous.zig, and its shape is L4's: a synchronous rendezvous at an Endpoint, with the packet copied directly from the sender's pages to the receiver's (copyAcross walks both sets of page tables through the physmap — no CR3 switch, no bounce buffer).

Two syscalls carry the request/reply exchange:

  • ipc_call(h, msg, reply) — copy the request packet to the provider, block until the reply packet comes back.
  • ipc_reply_wait(h, reply, recv) — reply to the client you're still holding (if any), then block for the next request. One syscall, because a provider's steady state is always "finish the last one, wait for the next".

An endpoint is reached by handle — a small integer index into the process's handle table (Task.handles), exactly like a file descriptor, and just as unforgeable. The provider never learns the client's identity beyond the badge delivered alongside each packet: the caller's task id, stamped by the kernel — unforgeable source addressing, a property a network's source field lacks.

The bootstrap problem — how a channel is first established — is the subject of protocol-namespace.md: a protocol is resolved by name and the channel arrives as a capability. (The mechanism it replaced — ipc_register/ipc_lookup under compile-time ServiceId integers — is gone: both syscalls and the enum were deleted when the registry landed, and their syscall numbers are left vacant.)

Interrupts are signals

notifyFromIsr posts an asynchronous signal to an endpoint — no payload, no reply owed — and wakes whoever is blocked in reply_wait. Its badge has the top bit set (notify_badge_bit), which is how a driver's single event loop distinguishes "a client wants something" from "the hardware wants something". Signals sit in a small coalescing ring on the endpoint, so an interrupt taken while the driver was busy elsewhere is not lost — coalesced, never dropped, which is exactly a signal's contract (the count may collapse; the fact may not).

This is what makes a user-space driver possible at all, and it's the subject of drivers.md.

What's next (partly done since)

  • Priority inheritance through IPC — still open: a high-priority client blocked on a low-priority provider suffers unbounded priority inversion.
  • Handle transfer. Landed as cap-passing (M13): ipc_call and ipc_reply_wait carry an optional capability alongside the bytes (send_cap), copying an endpoint or shared-memory handle into the peer's table — the mechanism by which channels are established and private channels built. First user: input subscribers register by handing over their own endpoint, and class drivers get a private channel to one device.
  • Asynchronous / buffered send for the cases where a rendezvous is the wrong shape (logging, event fan-out). Landed as ipc_send — a non-blocking post of an event packet (≤ 64 bytes) to an endpoint's bounded queue, delivered through reply_wait (badge bit notify_message_bit). Built for, and first used by, the input service's keyboard-event broadcast, where a synchronous push would let one dead subscriber hang the fan-out. A full queue drops the oldest — event packets are droppable by design (protocol-namespace.md's wiring section states the rule).
  • A bounded reply — half landed. The copy is still one packet (256 bytes) under the big kernel lock, but bulk transfer got its shared pages: shared_memory_create/map/physical, the region handle delegated as a capability (above) — the packets-never-fragment rule in practice. virtio-gpu's scanout surface is the first user (display-v2.md).

Lifecycle conventions over IPC (M17)

Three conventions from process-lifecycle.md ride the signal mechanism:

  • Process signals arrive as endpoint signals on the endpoint a process nominated with signal_bind (process.bindSignals): badge = the signal bit plus the coalesced pending mask (process.signalsFrom decodes). Statements, never questions; no payload, no reply.
  • One-shot timers (timer_bind, time.timerOnce) land as a timer-bit signal — the timed wait: a service arms a deadline and keeps serving, instead of blocking in sleep.
  • Kernel notifications go only to your own endpoint. signal_bind, timer_bind, process_subscribe, irq_bind, msi_bind, and spawn's exit endpoint all nominate where the kernel will speak, and all of them refuse an endpoint the caller did not create (-EPERM; the check is ipc.ownedBy, normalized to the process, so any thread may nominate an endpoint a sibling created). Holding a handle is not enough, because holding a handle is cheap: fs_resolve installs a mounted backend's capability in any caller's table, so every process holds a handle to PID 1's mailbox. Without the rule, "bind init's endpoint, then signal yourself" is a genuine, kernel-stamped terminate badge in PID 1's queue — a shutdown a receiver has no way to disbelieve — and timers, which carry no identity at all, multiply any loop that re-arms on its own landing.
  • A capability that arrives belongs to the turn. The kernel installs a sent capability in the receiver's table whatever the message's length or kind, so a receive loop must dispose of one on every path — the ping, the notification, the malformed request. The service harness (service.run) and PID 1 both hold it in an ipc.Arrival, released by a defer, and a handler that means to keep it says take(): forgetting closes, keeping is explicit. The reverse arrangement leaks a handle-table slot per request, and thirty-two unauthorized zero-length pings then end a service's ability to accept any capability — no subscribe, no shared-memory handover — for the rest of the boot.
  • The universal ping: a zero-length request is the liveness probe, answered with a zero-length reply by the service harness itself (service.run). No protocol's requests start at length zero, so the encoding cannot collide, and a wedged service simply fails to answer — which is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service protocol packet.