docs: the communication stack — /protocol namespace, layered model, non-unix hierarchy, SMEP/SMAP plan

The security-track design set. communication.md is the model: four layers
(namespace / protocol / channel / transport), packets and signals, parties
addressed by the channel and objects by target, transports as replaceable
buffer+doorbell mechanisms. protocol-namespace.md is L3+L2: /protocol names
contracts, resolution establishes a channel, init is the registrar,
restriction is per-process namespace delegation (with the microphone-prompt
worked example), the envelope is the universal packet header, migration
P1-P5 retires ServiceId. file-system-hierarchy.md replaces the unix-FHS
spec with the danos-native tree (/applications, /protocol, /system,
/volumes) and its migration table. ipc.md is re-cut as the kernel-ipc
transport document. smep-smap.md designs the kernel hardening: copy
discipline for the eight raw user-pointer syscalls, then SMEP, then SMAP
as a permanent tripwire.
This commit is contained in:
Daniel Samson
2026-07-31 18:11:28 +01:00
parent de9870175f
commit c96ef87714
14 changed files with 946 additions and 203 deletions
+120
View File
@@ -0,0 +1,120 @@
# Communication: the four layers
*Design, agreed 2026-07-31. The model document — the vocabulary and layering
every other communication document speaks.*
danos separates **what is said** from **how the bytes move**, so that the
mechanism is replaceable. The shape is a network stack's, cut into four
layers; a program only ever touches the top two.
```
L3 namespace /protocol/... names establishment points protocol-namespace.md
L2 protocol the language: packet schemas, verbs, targets the envelope, library/protocol/*
L1 channel two ends exchanging packets and signals the client library's Channel
L0 transport a buffer + a doorbell: moves the bytes ipc.md (kernel-ipc), later shm-ring, …
```
## Vocabulary
| Term | Meaning |
|---|---|
| **protocol** | The language: which packets exist, what their fields mean, which verbs a provider answers. Defined transport-independently in a `library/protocol/*` module. |
| **channel** | An open conversation between two processes, speaking one protocol. Established by opening a `/protocol/...` name; both ends can send and receive. |
| **packet** | The unit a protocol transmits: a bounded, atomic header+payload. Never fragmented — if it doesn't fit, it isn't a packet; bulk data rides shared memory with a packet as the doorbell. |
| **signal** | A payload-less poke below the packet layer: "something happened, come look." Coalescing — the count may collapse, the fact may not. |
| **transport** | What moves the bytes of one channel: a buffer plus a doorbell. Chosen (and upgradable) at establishment, invisible above L1. |
| **endpoint** | A termination point where a transport delivers. The kernel-ipc transport's endpoint is its kernel mailbox object. |
## Addressing: parties by channel, objects by target
There are no network-style addresses in a packet. The two questions addresses
answer are answered at different layers:
- **Who am I talking to?** The **channel**, decided once at establishment.
Opening `/protocol/input` yields a channel; every packet sent on it goes to
the peer. Nothing to route per-packet — like TCP, where no HTTP request
carries the server's IP.
- **Who sent this?** Attached to every received packet **by the channel
layer**, from identity the transport can verify — under kernel-ipc, the
kernel-stamped badge. The sender never writes a source field, which is what
makes source unforgeable (the property a network's spoofable source header
lacks).
- **Which of your things?** The packet's **`target`** field: *object*
addressing within the already-chosen peer — the vfs protocol's node id, the
display protocol's layer id, a block volume. `target = 0` addresses the
provider itself; a protocol without objects never uses it.
`target` is how instance multiplicity stays out of the namespace. Ten USB
sticks and the namespace still holds exactly one name, `/protocol/block`: a
channel to the provider, `enumerate` lists the current volumes as targets, a
`targets_changed` signal announces hotplug, and a read names its volume in
`target`. The unix `/dev/sda`,`/dev/sdb` problem is dissolved, not renamed.
If a future transport genuinely routes between machines, *it* carries real
source/destination addressing internally at L0 — the way IP runs under TCP —
and none of it surfaces into the packet header. Protocols stay ignorant of
distance.
## The transport (L0): a buffer and a doorbell
Strip any transport to its skeleton and the same two parts remain:
| Transport | Buffer | Doorbell | Status |
|---|---|---|---|
| **kernel-ipc** | kernel-owned mailbox (the `Endpoint`) | the scheduler (rendezvous wake) | the first transport — [ipc.md](../device-driver-development/ipc.md) |
| **shm-ring** | user-owned shared-memory ring | a signal | exists ad hoc (display bulk); to be formalized — the unlock for the 256-byte ceiling |
| network | NIC queue | an interrupt | someday, when danos networks |
Transports differ in their **properties**, which the channel layer exposes and
the protocol layer may depend on:
- **packet ceiling** — kernel-ipc: 256 bytes request/reply, 64 pushed. An
shm-ring's ceiling is its slot size. Kernel-ipc's 256 is the *floor* every
protocol may assume everywhere.
- **synchrony** — kernel-ipc's call is a rendezvous: natural backpressure, no
queue to size. An asynchronous transport buffers, so a channel over one
needs explicit flow control. Backpressure is a *transport property*, not a
channel guarantee — protocols that rely on it say so.
- **droppability** — pushed event packets may drop when a ring fills;
request/reply may not.
- **capability carriage** — **only kernel-ipc can move a capability.**
Handles are kernel objects; a user-space ring cannot transfer one. So
kernel-ipc is always the *establishment and control* transport — channels
are born on it, capabilities ride it — even when a channel's data is
negotiated onto something fatter.
That negotiation is the upgrade path: a channel starts on kernel-ipc; the
protocol's handshake may then delegate a shared-memory region (as a
capability, over kernel-ipc) and move its bulk traffic there. The display
path already does exactly this by hand; formalizing it in the channel layer
makes it every protocol's option.
## The channel (L1)
A channel has two ends, and **the ends are peers**: each may send packets,
each may receive, each may signal. Request/reply is a *pattern* over the
channel — a send with a correlated receive, which the kernel-ipc transport
happens to accelerate as a single rendezvous — not the definition of it. The
event stream (subscribe, then pushes) and the change signal (poke, then
re-read) are the other two patterns; all three are catalogued in
[protocol-namespace.md](protocol-namespace.md)'s wiring section.
The channel layer's obligations: deliver packets whole, attach the verified
source to every receive, expose the transport's properties, and hide the
transport's mechanics. The client library's `Channel` type is this layer made
concrete — a program holds channels that speak protocols and never touches a
raw handle.
## The protocol (L2) and the namespace (L3)
A protocol defines its packets through the envelope — every packet begins
`{operation, target}`, reserved verbs (`describe`, `enumerate`, `subscribe`,
`unsubscribe`) mean the same thing in every protocol, and `Define` checks
every packet against the transport floor at compile time. The full treatment,
including how names are granted, resolved, and restricted per process, is
[protocol-namespace.md](protocol-namespace.md).
Establishment points are named by contract — `/protocol/display`, never
`/protocol/ipc-1` — because the name must outlive the mechanism: a
transport named in the namespace could never be swapped, which would defeat
this document's premise.
+472
View File
@@ -0,0 +1,472 @@
# The protocol namespace
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. Not yet implemented —
the migration plan at the end is the work list.*
How a program finds, connects to, and is restricted from the things it talks to.
Three ideas, kept deliberately separate:
1. **Naming** — a path under `/protocol` names a *contract*, not a service.
2. **Access** — resolving that path yields an endpoint *capability*; what a process
cannot resolve, it cannot reach.
3. **Transport** — unchanged: packets over channels, moved by whichever
transport the channel rides (kernel-ipc first).
This document is layers **L3** (the namespace) and **L2** (the protocol
and its envelope) of the communication stack;
[communication.md](communication.md) owns the model and the vocabulary
(*protocol* the language, *channel* the conversation, *packet* the
transmitted unit, *signal* the payload-less poke, *transport* the
replaceable mechanism), and
[ipc.md](../device-driver-development/ipc.md) is the first transport.
## Why ServiceId has to go
Today a service calls `ipc_register(service_id, endpoint)` and a client calls
`ipc_lookup(service_id)`, where `ServiceId` is a compile-time enum in `abi.zig`
backed by a flat 16-slot table in the kernel. Three defects, in rising order:
- **Static.** The id space is baked into the ABI at compile time. A third-party
program can never introduce a service; the one place danos is *less* dynamic
than its own design.
- **Ungated.** `ipc_register` is callable by any process and *replaces* an
existing registration. Any process can hijack `.fat` or `.display` and
impersonate it. `ipc_lookup` is equally ambient.
- **Unrestrictable.** Because lookup is a syscall available to everyone, there is
no point at which "this process may not talk to the display" can be enforced.
Any future file-access restriction would be bypassable by speaking to the FAT
server directly.
## Naming: contracts, not services
`/protocol/<name>` names a protocol — the contract a conversation follows — and
resolving it connects you to whatever process currently provides that contract.
The client never cared *which* binary answers; it cares that its messages are
understood. Naming the contract makes that explicit, and buys:
- **Swappable providers.** Replace the display server; `/protocol/display`
routes to the new one; clients notice nothing.
- **Test fakes.** Spawn a program whose namespace wires `/protocol/display` to a
mock. The name promises the protocol; the mock speaks it.
- **One vocabulary.** The names mirror `library/protocol/`: a program imports
the `display-protocol` module, then opens `/protocol/display`. What you
compiled against and what you ask the namespace for are the same word.
A leaf names one contract — kebab-case, full words, matching the
`library/protocol/` module that defines its wire format — and related
contracts group into directories: `/protocol/networking/ip`,
`/protocol/networking/bluetooth`. Directories organize *contracts only*;
they never encode addressing (see below), so a directory appears because a
domain has several contracts, never because hardware multiplied. The module
tree mirrors the namespace (`library/protocol/networking/ip` ↔
`/protocol/networking/ip`), and registrar grants scope naturally to subtrees
— an application installed at `/applications/foo` can be granted
`/protocol/applications/foo/...` and nothing above it. `/protocol` is
top level, beside `/system` and `/applications`, because the boundary it names
is spoken on both sides: applications talk to protocols as much as the OS does
(see [file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md)).
**Addressing lives inside the protocol, never in the path.** Which volume, which
layer, which input device — that is a destination field in the messages, the way
TCP carries a destination address, and the way danos protocols already work (the
display protocol multiplexes layer ids; the vfs protocol addresses node ids).
The namespace answers exactly one question — *may this process speak this
protocol at all* — so `/protocol/block` is one name no matter how many disks are
attached. The source address is never in the message either: it is the IPC
badge, stamped by the kernel per message, unforgeable — a property TCP's source
address does not have.
`/system/devices` (the device inventory) stays purely informational: facts for
diagnosis, never a routing mechanism. Unix conflated the two in `/dev`; danos
does not. You *read about* hardware in `/system/devices`; you *talk to* it
through `/protocol`.
## Resolution: a protocol node in the VFS
The kernel VFS router already does the hard part: `fs_resolve` matches a mount
prefix and installs the backend's endpoint capability in the caller's handle
table. The registry is just a backend mounted at `/protocol` — ring 3, like FAT.
Connecting is a normal vfs-protocol `open` with one twist in the reply:
```
client kernel router registry backend
│ fs_resolve("/protocol/display") │
│──────────────────────────▶│ prefix match: /protocol │
│◀── registry endpoint ─────│ (capability installed) │
│ vfs open("display") ──────────────────────────────────────▶│
│◀───────────────── Reply + capability = provider endpoint ──│
│ ipc_call(provider, display-protocol messages...) │
```
Both capability moves use machinery the kernel already has: request-direction
and reply-direction `send_cap` on `call`/`replyWait`. The vfs protocol needs two
additions, both append-only:
- `NodeKind.protocol` — a node that names a contract; its `open` establishes
a **channel** (delivered as an endpoint capability) instead of returning a
file id. The node is the protocol, the channel is the conversation, and the
addressing inside the packets decides where within the provider each one
lands. `readdir` over `/protocol` lists protocol nodes like any others, so
the tree stays browsable for diagnosis.
- The convention that an `open` reply may carry a capability. File backends
(FAT) never use it; synthetic backends (the registry, later the device
inventory) do.
The path lookup happens once, at connect time. The hot path — `ipc_call` on the
cached endpoint — is untouched. A provider crash turns the cached endpoint dead
(`-EPEER`), and the client's recovery is to re-resolve: the restart story falls
out of the naming layer for free.
## Registration: the registrar, held by init
The registry backend is **init**. It is already PID 1, already spawns every
service from its manifest, and already holds the supervision link to each — it
is the process that *knows* which binary is which. (If init grows
uncomfortable, the same design lifts into a dedicated registry service that
init spawns first and delegates to; nothing below changes.)
- **Binding.** A service creates its endpoint and sends the registry a `bind`
request with the protocol name as payload and the endpoint attached as the
call's capability.
- **Authorization.** Init's manifest gains a column: the protocols each spawned
binary may bind. A `bind` from any process not granted that name is refused
(`-EPERM`) — the badge identifies the caller, the supervision records map
badge to binary. This is the registrar authority; it never leaves init.
- **Collision is an error.** A name already bound refuses a second bind — never
last-writer-wins. When a provider dies, init (its supervisor) unbinds its
names; the restarted instance binds again.
- **Provenance.** The registry records name → task id → binary path, so a
diagnostic listing answers "who serves this?" at a glance:
```
/protocol/display pid 12 /system/services/display
/protocol/input pid 7 /system/services/input
```
`ipc_register` and `ipc_lookup` retire; the `ServiceId` enum leaves `abi.zig`.
The kernel keeps one residual rule: `/protocol` becomes a reserved prefix like
`/system` — `fs_mount` refuses to shadow it, and init's boot-time mount is the
only one it will ever hold. (Full gating of `fs_mount` is a separate item on
the security track; the reserved prefix closes the hole for this namespace
without waiting for it.)
## Restriction: per-process namespaces, not ACLs
danos has no users and no principals, deliberately. Restriction is therefore
**delegation**: what a process may open is decided by whoever spawned it, and
enforcement is absence — a protocol you cannot resolve does not exist for you.
"Permission denied" and "not found" are the same answer, which is the same
discipline the device layer already follows: the claim is the capability; here,
the resolvable name is the capability.
Two stages, deliberately ordered so the useful half lands first:
**Stage one — the registry filters by badge.** Init is both the spawner and the
registry, so its manifest already knows which binary may *open* which protocols
(a second manifest column, beside the bind grants). An `open` from a process
whose binary is not granted that protocol is refused. No new kernel mechanism
at all; the display driver's view can be narrowed to nothing, a future
downloaded application's to `display` and `input`, today.
**Stage two — spawn passes the namespace.** `spawn` gains an initial
capability: the child's connection to *its* registry view, chosen by the
spawner. A newly spawned process starts with an empty handle table and this one
handle — its world is whatever its parent wired in. This removes the last
ambient reach (`fs_resolve` finding `/protocol` globally), lets any supervisor
— not just init — narrow or fake a child's view (an application launcher
granting an app only what its manifest declares; a test harness substituting
every provider), and composes down the supervision tree. Stage one's manifest
column becomes the *content* of the view init builds, so nothing is thrown
away.
### A worked example: the microphone prompt
The scenario stage two exists for: an application opens
`/protocol/audio-input`, and the user should be asked. The supervisor is an
ordinary user process — an application launcher — and the flow needs no new
security concepts:
1. The launcher spawned the app with a namespace channel that terminates at
**the launcher itself**. The app's whole world is a conversation with its
supervisor.
2. The app's `open("audio-input")` packet lands in the launcher,
badge-stamped. The launcher spawned the app, so badge → binary path
(`/applications/foo`) is its own supervision record — "remember my choice"
needs no identity system.
3. Grant unknown → the launcher parks the request and shows a prompt (it is a
user process with display access; init never does UI). Blocking an open on
a human is architecturally fine: opens are connect-time, never hot-path.
4. **Yes** → the launcher opens `/protocol/audio-input` in *its own*
namespace and attaches the resulting channel to the parked reply. The app
cannot tell a prompt happened — a consented open is indistinguishable from
a direct one, merely slower.
5. **No** → refuse the open, indistinguishable from "no such protocol" — or
hand the app a **fake**: a silence-generating provider. The test-fake
mechanism doubles as a privacy feature.
The capability discipline holds throughout: the launcher can only grant what
it holds — if init never gave the launcher `audio-input`, no prompt can
conjure it. Consent is delegation flowing down the supervision tree, never a
global ACL edit. And the provider still sees the app's badge on every packet,
so a coarser second check at the audio service remains possible.
Two mechanical requirements this scenario pins on stage two:
- **Parked replies.** A prompt takes seconds, and the service loop holds one
outstanding reply today — the launcher must park request A, keep serving B
and C, and reply to A later (by badge). The kernel already tracks owed
replies (that is how death delivers `-EPEER`); multiple parked replies is
the extension, in the harness and, if needed, the kernel.
- **Granted channels are dedicated, hence revocable.** Once the app holds a
channel capability, nobody reaches into its handle table — so a
prompt-granted channel must be one that can be *killed*: a dedicated
endpoint pair (or per-client session at the provider) whose death turns
the app's capability into `-EPEER`. Revoking microphone access is then
killing that channel, using machinery that already exists.
One adjacent problem, named and deferred: **trusted UI**. The prompt is only
meaningful if the app cannot draw a convincing fake or overlay the real one —
a display-layer question (a reserved surface for the supervisor chain), owned
by the display track, not this one.
Fine-grained restriction *within* a protocol (this process may use volume A but
not volume B) is not the namespace's job. The capability-shaped answer, when it
is needed: the supervisor pre-opens a connection scoped to one target and passes
that connection to the child, which never opens `/protocol/block` at all.
Delegation again, not ACLs.
## The envelope: one addressing scheme for every protocol
Every protocol module today hand-rolls its `Request`/`Reply` with an
`operation` first field. That convention becomes a library, so addressing is
uniform and the rules are enforced by construction rather than by review. New
module: **`library/protocol/envelope`** (the one protocol-layer module that is
not itself a protocol).
```zig
/// Every packet a danos protocol transmits begins with this header.
pub const Header = extern struct {
operation: u32, // the verb; values 0..15 are reserved universal verbs
_padding: u32 = 0,
/// Object addressing, never party addressing: which of the peer's
/// objects this packet operates on — a volume, layer, node, device.
/// 0 addresses the provider itself. Parties are addressed by the
/// channel; the protocol defines target's meaning; the field's place
/// and width are universal.
target: u64 = 0,
};
/// Reserved verbs, answered by every provider.
pub const operation_describe: u32 = 0; // -> protocol name, version, target kinds
pub const operation_enumerate: u32 = 1; // -> the current targets, one per reply page
pub const operation_subscribe: u32 = 2; // capability = the subscriber's endpoint
pub const operation_unsubscribe: u32 = 3;
pub const first_protocol_operation: u32 = 16;
/// Every reply begins with this.
pub const Status = extern struct {
status: i32, // 0 or a negative errno
_padding: u32 = 0,
len: u32 = 0, // payload bytes following the header
_padding2: u32 = 0,
};
```
A protocol is then *defined through* the envelope, not beside it:
```zig
pub const Protocol = envelope.Define(.{
.name = "display",
.version = 1,
.operations = &.{
.{ .name = "configure_layer", .request = ConfigureLayer, .reply = void },
.{ .name = "blit", .request = Blit, .reply = void },
...
},
});
```
`Define` is comptime and is where the enforcement lives:
- verbs are numbered automatically from `first_protocol_operation`, so no
protocol can collide with the reserved range;
- every packet is size-checked at compile time against the kernel-ipc floor
— `packet_maximum` (256) for request/reply, `post_maximum` (64) for event
packets. Ceilings are transport properties
([communication.md](communication.md)); the floor is what every protocol
may assume on any transport. The errors that today surface as runtime
truncation become compile errors, and packets-never-fragment is enforced
at the source;
- the generated type carries encode/decode helpers and a provider-side dispatch
table, so a provider answers `describe` automatically and unknown operations
with `-ENOSYS` uniformly;
- the service harness (`library/kernel/service.zig`) accepts the generated
dispatch type, which is what makes the envelope *enforced*: a protocol that
bypasses `Define` does not plug into the harness.
Universal conventions that ride on the reserved verbs:
- **`describe`** is the version handshake. Version lives in the handshake, not
in every message — the 256-byte budget is too small to spend per call.
- **`enumerate`** is how multi-target protocols expose their targets, and the
standard `targets_changed` notification (a notify bit) tells subscribers to
re-enumerate — arrival and removal of volumes, layers, devices all take the
same shape. Hotplug fits the notification ring far better than a filesystem
tree ever did.
- **Source is the badge.** No protocol defines a "sender" field; the kernel's
per-message badge is the only source identity, and providers key per-client
state on it.
### Paths resolve once; integers do the work
A rule the envelope makes official: **a path appears in a conversation at most
once — at resolve or open — and everything after it addresses integers.** The
namespace resolves `/protocol/display` to an endpoint; a backend's `open`
resolves a path payload to a node id; from then on every packet carries the
integer in `target`. Integers compare in one instruction and fit the fixed
header, and the 256-byte message budget never re-carries path strings on the
hot path. This is already the system's shape — vfs node ids, display layer ids
— and the envelope pins it as the required shape for every protocol.
Two integer identities, not to be confused:
- **An open handle** — what vfs `open` returns today: transient, meaningful
only within one client's session with one provider, swept when the client
exits. Cheap, and all a protocol usually needs. Handles must be **scoped per
client** — validated against the badge, or drawn from a per-client id
namespace. (Today the FAT server's node ids are guessable small integers
honoured across clients; that hole closes with this rule.)
- **A persistent node identity** — a unix inode number, stable across opens
and renames. danos deliberately does not promise this, because FAT cannot
deliver it: a FAT file's identity is its directory entry, and rename or
truncation moves every candidate anchor. If a future filesystem or a cache
layer needs stable identity, that is the backend's promise to make, never
the protocol's assumption.
The five existing protocol modules (`vfs`, `display`, `input`, `power`,
`block`, plus `scanout`, `usb-transfer`, `device-manager`) rebase onto the
envelope during the migration flag-day. `input-protocol`'s subscribe/publish
split and `vfs-protocol`'s node addressing both map cleanly (`node` and layer
ids become `target`).
## Wiring: how conversations flow
The patterns below are channel-layer (L1) shapes; the delivery mechanics are
the kernel-ipc transport's, described here because it is the transport every
channel starts on. Kernel-ipc provides exactly three delivery shapes, and
every one is unicast. An endpoint is a mailbox owned by one process — its
creator receives; anyone holding its capability sends into it. That direction
never reverses:
1. **Synchronous call** — request/reply. The kernel parks the caller and
`replyWait` delivers the reply straight back, so the provider answers
without holding any capability to the client. Badge-stamped, blocking, and
the *only* shape that carries capabilities (in the request, and in the
reply — which is how a reverse path is bootstrapped).
2. **Asynchronous send** — an event packet pushed into the receiver's post
ring, at most `post_maximum` (64) bytes, no reply owed, never blocks the
sender. Strictly one-way: to be pushed to, you must first hand the pusher
your endpoint.
3. **Signals** — payload-less notification bits, below the packet layer,
coalescing: "something changed, come look."
A bidirectional link is therefore always **a pair of endpoints**, one per
direction, each delivered by cap-passing. Three conversation patterns are
built from these, and the envelope names all three:
- **Request/response** — the synchronous call. The default, and the only
place capabilities move.
- **Event stream** — `subscribe` (a synchronous call whose attached
capability is the subscriber's own endpoint), after which the provider
pushes events asynchronously; `unsubscribe` or subscriber exit ends it.
Listened-to, not blocked-on.
- **Change signal** — a signal plus re-read: `targets_changed` →
`enumerate`. For state whose truth lives with the provider.
**Broadcast is a provider pattern, never a kernel primitive.** The kernel
does not know subscriber sets — a service does. The input service is the
model: sources *publish* (a unicast call to the service), the service
*broadcasts* (a fan-out loop of asynchronous sends over its subscriber list,
so one dead subscriber can never stall the rest). One fan-out point per event
domain, owned by the service that defines the event.
The harness owns the machinery: the subscriber table, the dead-subscriber
sweep (via process-exit notifications), and the fan-out loop — all written by
hand in `input.zig` today, lifted into the service harness so every protocol
gets identical semantics. `Define` declares a protocol's events (`.events`),
and each event type is checked against `post_maximum` at compile time,
generalizing the assert `input-protocol` already carries.
**Event packets are droppable.** A slow subscriber's ring fills, and the
provider must not block on it — so an event stream is a hint or a coalescing
signal, never a ledger. Anything that must not be lost is either re-readable
state (the change-signal pattern) or bulk data in shared memory with a
packet as the doorbell, which is how the display path already works — the
packets-never-fragment rule and this one are the same rule seen from two
sides.
**Source direction (open point).** Today event sources are *clients*: an
input driver resolves `/protocol/input` and delivers each event as a
synchronous `publish` call — one capability, obtained by resolution, covers
everything, and the badge tells the service exactly who each event came from.
The inversion — the service subscribing to each driver — would require every
driver to be individually discoverable and its endpoint ferried to the
service, machinery whose payoff (the service choosing its sources) the
namespace already provides more cheaply: only a process granted open on
`/protocol/input` can publish into it. Sources stay clients for now;
revisited at restriction stage two, when a supervisor can wire capabilities
at spawn time.
## What this deliberately does not solve
The wider security track, for which this namespace is the foundation, not the
whole:
- **File access restriction** — the point of the exercise. The same stage-two
namespace mechanism extends from protocol names to file paths: the spawner
decides which subtrees resolve. Designed separately once this lands.
- `fs_mount` gating beyond the reserved prefixes; `system_spawn` gating;
`klog_read` being world-readable; backends checking the badge on per-node
operations (the FAT server honours node ids across clients today).
- Kernel hardening items already noted in-tree: SMEP/SMAP and SYSRET
canonical-RIP, now designed in [smep-smap.md](smep-smap.md).
- Pipes/FIFOs for the POSIX layer — a byte-stream object *beside* message IPC,
wanted by the Python track, unrelated to naming.
- **Trusted UI** — a permission prompt an application cannot fake or overlay
(see the microphone example). A display-track concern: the supervisor chain
needs a reserved surface.
## Migration plan
Flag-day per phase, in the style of the DMA-capability conversion — no
dual-stack periods, the QEMU suite green at each phase boundary.
**P1 — mechanics, no behavior change.** The `envelope` module with its comptime
`Define`, unit tests; `NodeKind.protocol` and the open-reply-capability
convention in `vfs-protocol`; existing protocols untouched.
**P2 — the registry.** Init serves `/protocol` (bind with manifest
authorization, collision refusal, unbind on provider death, provenance);
kernel reserves the `/protocol` prefix; every service converts from
`ipc_register` to `bind`, every client from `ipc_lookup` to resolve-and-open;
`ServiceId`, `ipc_register`, `ipc_lookup` deleted. Tests: unauthorized bind
refused, collision refused, provider restart re-binds and a client re-resolves.
**P3 — restriction, stage one.** The open-grant column in init's manifest;
registry refuses ungranted opens. Test: a fixture process denied a protocol its
neighbour is granted.
**P4 — protocol rebase.** Existing protocol modules re-expressed through
`Define`; providers move onto the generated dispatch; `describe`/`enumerate`
answered everywhere; the conformance test fixture exercises the reserved verbs
against every registered provider.
**P5 — restriction, stage two.** Spawn's initial capability; namespace views
built by the spawner; ambient resolution of `/protocol` retired. Includes the
two requirements the microphone example pins: **parked replies** (a
supervisor parks an open, keeps serving, replies later by badge) and
**dedicated, killable granted channels** (revocation = channel death →
`-EPEER`). Scoped separately — it touches `spawn`, the loader contract, and
every supervisor — and lands together with the file-path half of namespacing.
The unix-path migration ([file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md#migration))
is independent of P1–P5 and can land before or after.
+131
View File
@@ -0,0 +1,131 @@
# SMEP and SMAP — supervisor-mode hardening
*Design, 2026-07-31. Not yet implemented. Companion to
[protocol-namespace.md](protocol-namespace.md) on the security track — this is
the hardware half; that is the namespace half.*
Two CR4 bits that make the CPU refuse the two things a kernel should never do
with user memory:
- **SMEP** (Supervisor Mode Execution Prevention, CR4 bit 20): instruction
fetch in ring 0 from a page whose U/S bit says *user* → #PF. Kills the
classic ret2usr exploit shape — a kernel bug that redirects control flow
can no longer land in attacker-prepared user code.
- **SMAP** (Supervisor Mode Access Prevention, CR4 bit 21): data read/write
in ring 0 to a user page → #PF, unless `EFLAGS.AC` is set. `stac`/`clac`
open and close deliberate access windows; danos's design needs no windows
at all (below).
Detection is CPUID leaf 7, subleaf 0, EBX bit 7 (SMEP) and bit 20 (SMAP).
Both bits are per-core state: the BSP and every AP must set them.
## Why, in danos terms
Every syscall argument is an attacker-controlled integer, and several take
pointers. A kernel bug that dereferences a crafted pointer reads, writes, or
executes memory of the attacker's choosing — the exact bug class the
isolation tracks exist to prevent. SMEP/SMAP turn that class from "silent
compromise" into "immediate, attributable #PF with a kernel RIP in the log."
The second benefit matters as much as the first: **SMAP is a permanent
tripwire.** Once it is on, any *future* syscall that touches user memory
directly — instead of going through the checked copy layer — faults the
first time the QEMU suite runs it. The discipline stops depending on review.
## Where danos already stands
The design is closer than it looks, because the IPC layer was built right:
- **The copy layer is already SMAP-proof.** `copyAcross` and `copyFromUser`
(`system/kernel/ipc-synchronous.zig:305,333`) never dereference a user
virtual address: they walk the page tables and move bytes through the
physmap — kernel mappings throughout. SMAP cannot object.
- **Syscall entry already clears AC.** `SFMASK = 0x4_0700` clears IF, TF,
DF, **AC** on every `syscall`
(`system/kernel/architecture/x86_64/per-cpu.zig:76`). The syscall path is
SMAP-clean from day one.
- **The interrupt path is not.** Hardware does *not* clear AC on IDT
delivery, and ring 3 can set AC with `popfq` — so a hostile process could
take an interrupt with AC=1 and have the handler run with SMAP suspended.
`isr_common` (`system/kernel/architecture/x86_64/isr.s:366`) needs a
`clac` beside its `swapgs`.
- **CR4 today:** the BSP inherits firmware CR4 (no kernel write anywhere);
APs set PAE/OSFXSR/OSXMMEXCPT in `trampoline.s:62-68`. Neither path sets
SMEP/SMAP yet, and both must.
- **The stragglers.** A handful of syscalls still dereference user pointers
raw after a bounds check — every one is a SMAP #PF waiting to happen, and
every one is *already* a latent kernel fault today (an unmapped-but-in-
range user page oopses the kernel instead of failing the call). The audit
list, from the 2026-07-31 survey of `system/kernel/process.zig`:
| Syscall | Raw access |
|---|---|
| `system_spawn` | name + argument blob (`:972`, `:980`) |
| `fs_resolve` | path in, result out (`:1780`, `:1797`) |
| `fs_mount` / `fs_unmount` | prefix + rewrite strings (`:1864`) |
| `fs_node` | read buffer out (`:1820`) |
| `debug_write` | message bytes (`:1700`) |
| `klog_read` | log bytes out (`:1740`) |
| `process_enumerate` | descriptor array out (`:1132`) |
| `device_enumerate` | descriptor array out (`:388`) |
(Some paths already do it right — the futex word and the device-register
descriptor go through `copyFromUser` (`:1087`, `:924`). The write
direction has no helper yet.)
- **One known gap inside the copy layer itself:** the walk checks presence,
not the leaf U/S and writable bits (`ipc-synchronous.zig:20-22` flags
this). Today that is nearly moot — the user half contains only mappings
the kernel itself created for that process — but it must close before
shared or copy-on-write mappings exist, and closing it is part of making
the copy layer the single trusted door.
## The plan
**H1 — copy discipline (the real work).** A `user-memory` kernel module:
`copyFromUser` / `copyToUser` (the missing write direction) via the physmap
walk, with U/S and writable leaf checks closing the in-tree TODO. Convert
the eight stragglers. This fixes the latent unmapped-page kernel fault on
its own — it is worth doing even if SMEP/SMAP never shipped. QEMU suite
green; no behavior change visible to correct programs.
**H2 — SMEP.** A leaf-7 feature probe (the kernel has per-leaf `cpuid`
helpers in `apic.zig` to generalize); set CR4.SMEP during per-CPU bring-up
on BSP and APs — prefer the Zig-side per-CPU init over the trampoline
assembly, so one code path covers every core and the trampoline stays
minimal. Audit first that ring 0 never executes user-mapped pages: kernel
text lives in the kernel half, `jump_to_user` is kernel code, and the AP
trampoline page is kernel-mapped — expected clean, verify before flipping.
**H3 — SMAP.** Add `clac` at `isr_common` entry. `clac` is #UD on CPUs
without SMAP, so the instruction is a 3-byte NOP in the image, patched to
`clac` at boot when CPUID advertises SMAP (one-time patch beats a
conditional branch in the hottest path in the kernel). Then set CR4.SMAP in
the same per-CPU init. From this point the whole QEMU suite doubles as the
enforcement test: any missed raw dereference is a vector-14 with a kernel
RIP and a user CR2 — loud and attributable.
**H4 — keep it honest.** A line in the coding standards: kernel code
touches user memory only through `user-memory`; there is no `stac` anywhere
in the tree, and a PR that adds one is wrong by definition. SMAP enforces
the rule mechanically at test time.
Feature-gating follows the timekeeping rule (work on any VM, real Intel,
real AMD): both bits are probed, absence is logged and tolerated — like the
IOMMU's fail-open, the machine still boots, just unhardened. QEMU: TCG
implements both; KVM inherits the host (Intel Ivy Bridge+ for SMEP,
Broadwell+ for SMAP; AMD Zen+ for both). The test images should run with
`-cpu max` so the suite always exercises the enabled paths.
## Adjacent, deliberately separate
- **SYSRET canonical-RIP hardening** (`isr.s:192-194` documents it): a
non-canonical return RIP makes `sysretq` #GP *in ring 0* on Intel. Same
hardening bucket, independent fix (validate RCX before `sysretq`, fall
back to `iretq`), should ride the same branch as H2/H3 but is not
SMEP/SMAP.
- **KPTI / Meltdown-class leaks are out of scope.** SMEP/SMAP police
architectural accesses, not speculative ones. danos runs one kernel
mapping in every address space and accepts that on affected hardware;
revisit only if the threat model ever includes hostile native code on
shared machines.
+8 -8
View File
@@ -11,7 +11,7 @@ sequential pass and hands the bytes to the kernel unmodified.
The capsule is a *performance artifact*, not a source of truth. The boot
volume's `/system` and `/test` file trees remain the canonical layout (see
[danos-file-system-hierarchy-FSH.md](../file-system-development/danos-file-system-hierarchy-FSH.md));
[file-system-hierarchy.md](../file-system-development/file-system-hierarchy.md));
the capsule is a pre-baked snapshot of the same binaries, derived from the same
build graph, so the running system is identical whether the loader read the
capsule or walked the tree.
@@ -36,11 +36,11 @@ so it need be no fancier. Little-endian throughout:
```
Header magic: u32 = "DNR2" (0x32524E44), count: u32
Entry × count name: [64]u8 (NUL-padded FHS path), offset: u64, len: u64
Entry × count name: [64]u8 (NUL-padded hierarchy path), offset: u64, len: u64
blobs... each entry's file bytes, at its offset within the image
```
- **Names are full FHS paths** (`/system/services/init`), not basenames — that
- **Names are full hierarchy paths** (`/system/services/init`), not basenames — that
is what "v2" means. The 64-byte capacity matches `abi.maximum_process_name`,
so a task named after its binary path is never truncated. Paths longer than
63 bytes are a build error (`pack-system-image.py` rejects them).
@@ -54,14 +54,14 @@ blobs... each entry's file bytes, at its offset within the image
## How it is built
`build.zig` maintains one `bundled` list — every user binary and its FHS home.
`build.zig` maintains one `bundled` list — every user binary and its hierarchy home.
Three artifacts are derived from that same list, in the same build graph, so
they cannot drift apart:
1. **The tree**: each binary installed at its FHS path (`zig-out/system/...`
1. **The tree**: each binary installed at its hierarchy path (`zig-out/system/...`
and `zig-out/test/...`, mirrored onto the FAT boot volume by
`tools/make-fat-image.py`).
2. **The manifest** (`system/manifest`): the FHS path of every bundled binary,
2. **The manifest** (`system/manifest`): the hierarchy path of every bundled binary,
one per line — the loader's per-file fallback input.
3. **The capsule**: `tools/pack-system-image.py` packs the same binaries into
the v2 container, installed at `zig-out/boot/system.img` and placed on the
@@ -103,7 +103,7 @@ the kernel (`kernel.zig`) then publishes the same bytes twice, to two
consumers:
- **The process layer** (`process.zig`): `system_spawn` looks binaries up in
the ramdisk via `Reader.find` — exact FHS path, or unique basename for
the ramdisk via `Reader.find` — exact hierarchy path, or unique basename for
pre-path callers — and loads them as fresh ring-3 processes. The stored path
becomes the task's name.
- **The VFS root** (`vfs.zig`, `setInitialRamdisk`): the image is mounted as
@@ -111,7 +111,7 @@ consumers:
paths, so `/system` and, when the fixtures are bundled, `/test`. Directory
nodes are derived from the entry paths (the unique parents), so the trees
are listable and their files readable over the normal VFS protocol — the
FHS boot tree every process sees comes straight out of the capsule bytes.
boot tree every process sees comes straight out of the capsule bytes.
The image is never copied after the handoff and never mutated: the initrd is
immutable, which is what makes the VFS's node serving lock-free.