merge: security track group 2 — the protocol namespace replaces ServiceId

# Conflicts:
#	docs/security-track-plan.md
This commit is contained in:
Daniel Samson
2026-08-01 04:45:14 +01:00
73 changed files with 3925 additions and 364 deletions
+2 -2
View File
@@ -87,13 +87,13 @@ rest of the system hasn't had to face:
│ (ResourceKind.memory = [base, height*pitch], write-combining hint,
│ plus DisplayInfo{width, height, pitch, format, refresh_hz})
▼
display service (system/services/display/, ServiceId.display) ← the compositor
display service (system/services/display/, /protocol/display) ← the compositor
│ device.claim(display node) → mmio_map(WRITE-COMBINING) = FRONT buffer (the LFB)
│ mmap(cacheable) a BACK buffer of the same geometry
│ owns: an ordered LAYER STACK + a per-frame DAMAGE tracker (rect list or tile grid)
│ loop: composite dirty layers → back buffer → present dirty rects → front
│ backend is an INTERNAL interface: {gop-fb} at boot; {virtio-gpu} on hot-attach (v2)
▼ reached by name (ipc_lookup); clients drive it over the display protocol
▼ reached by name (open /protocol/display); clients drive it over the display protocol
┌────────────────────────────────────┬──────────────────────────────────────┐
drawing clients (v1) surface clients (deferred)
display commands: display surfaces:
+25 -3
View File
@@ -114,9 +114,10 @@ unforgeable source addressing, a property a network's source field lacks.
The bootstrap problem — how a channel is first established — is the subject of
[protocol-namespace.md](../os-development/protocol-namespace.md): a protocol is
resolved by name and the channel arrives as a capability. (The mechanism this
replaces, `ipc_register`/`ipc_lookup` under compile-time `ServiceId` integers,
is retired by that design.)
resolved by name and the channel arrives as a capability. (The mechanism it
replaced — `ipc_register`/`ipc_lookup` under compile-time `ServiceId` integers —
is gone: both syscalls and the enum were deleted when the registry landed, and
their syscall numbers are left vacant.)
### Interrupts are signals
@@ -169,6 +170,27 @@ signal mechanism:
- **One-shot timers** (`timer_bind`, `time.timerOnce`) land as a
timer-bit signal — the timed wait: a service arms a deadline and keeps
serving, instead of blocking in sleep.
- **Kernel notifications go only to your own endpoint.** `signal_bind`,
`timer_bind`, `process_subscribe`, `irq_bind`, `msi_bind`, and spawn's exit
endpoint all *nominate where the kernel will speak*, and all of them refuse an
endpoint the caller did not create (`-EPERM`; the check is `ipc.ownedBy`,
normalized to the process, so any thread may nominate an endpoint a sibling
created). Holding a handle is not enough, because holding a handle is cheap:
`fs_resolve` installs a mounted backend's capability in *any* caller's table,
so every process holds a handle to PID 1's mailbox. Without the rule, "bind
init's endpoint, then signal yourself" is a genuine, kernel-stamped `terminate`
badge in PID 1's queue — a shutdown a receiver has no way to disbelieve — and
timers, which carry no identity at all, multiply any loop that re-arms on its
own landing.
- **A capability that arrives belongs to the turn.** The kernel installs a sent
capability in the receiver's table whatever the message's length or kind, so a
receive loop must dispose of one on *every* path — the ping, the notification,
the malformed request. The service harness (`service.run`) and PID 1 both hold
it in an `ipc.Arrival`, released by a `defer`, and a handler that means to keep
it says `take()`: forgetting closes, keeping is explicit. The reverse
arrangement leaks a handle-table slot per request, and thirty-two unauthorized
zero-length pings then end a service's ability to accept any capability —
no subscribe, no shared-memory handover — for the rest of the boot.
- **The universal ping**: a **zero-length request is the liveness probe**,
answered with a zero-length reply by the service harness itself
(`service.run`). No protocol's requests start at length zero, so the
+14 -3
View File
@@ -161,15 +161,26 @@ Aligned to the node-kind table in the file-system hierarchy
| 4 | symbolic link |
| 5 | fifo |
| 6 | socket |
| 7 | protocol |
Clients should map unknown values to *regular* rather than reject — the
table can grow. Kind 6 (`socket`) keeps its wire value but is retired from
the design — a named rendezvous point is exactly what a `protocol` node is,
planned as value 7 with the protocol namespace
landed as value 7 with the protocol namespace
(docs/os-development/protocol-namespace.md). `character_device` (stream
semantics — the tty/console shape) and `block_device` (raw sector-addressed
volumes, reserved) remain part of the design.
## An open reply may carry a capability
`open` rides `ipc_call`, whose reply direction can hand back an endpoint
capability alongside the `Reply` header. A file backend never uses it — FAT
answers with a node id and nothing else — but a **synthetic** backend does:
opening a `protocol` node returns the provider's endpoint, and possession of
that endpoint *is* the channel. The convention is per-backend, not
per-operation, so a client that opens an ordinary file simply receives no
capability, exactly as before.
## Lifetimes and trust
Open-node ids live in the backend. A client that dies without closing leaks
@@ -188,8 +199,8 @@ What a non-Zig implementation may rely on, and what it must not:
- Operation values, flag bits, `NodeKind` values, and struct layouts are
**append-only and frozen once shipped**. The unit test in
`library/protocol/vfs/vfs-protocol.zig` pins a sample of them (the `DirectoryEntry`
size, `NodeKind` 0–1, `Operation` values 0, 4 and 5); this page is the
full record of the frozen values.
size, `NodeKind` 0–1 and 6–7, `Operation` values 0, 4 and 5); this page is
the full record of the frozen values.
- The 256-byte message ceiling is a property of the current IPC transport,
not a promise; clients should read `maximum_payload`-shaped limits from the
reply lengths they actually get (loop-until-done), not hard-code 224.
+2 -2
View File
@@ -231,8 +231,8 @@ Two consequences of neutrality bind on later work:
- **Cross-firmware surfaces are named by domain, not firmware.** System power is
a [`power`](power.md) protocol, not an "ACPI events" protocol: on x86 the acpi
service registers it, on ARM a PSCI/mailbox service registers the same
`ServiceId.power`, and subscribers never learn the difference.
service binds it, on ARM a PSCI/mailbox service binds the same
`/protocol/power`, and subscribers never learn the difference.
- **Identity must widen before the fdt service exists.** `DeviceDescriptor`'s
8-byte `hid` holds an EISA id but cannot hold an FDT `compatible` string
(`"brcm,bcm2835-aux-uart"`); the identity field grows before the ARM path can
+3 -3
View File
@@ -15,9 +15,9 @@ Where the events come from is firmware-specific — on x86 they ride the ACPI SC
([acpi.md](acpi.md)); on a Raspberry Pi they would come from PSCI or a mailbox.
What subscribers want is not: *the lid closed* means the same thing regardless of
who noticed. So the surface is **domain-named**. There is a `power-protocol`
module and a well-known `ServiceId.power = 5`; on x86 the **acpi service**
registers it, and on ARM a PSCI/mailbox service will register the *same* id.
Subscribers call `ipc.lookup(.power)` and never learn which firmware they
module and a contract named `/protocol/power`; on x86 the **acpi service**
binds it, and on ARM a PSCI/mailbox service will bind the *same* name.
Subscribers open `/protocol/power` and never learn which firmware they
are on — the neutrality the whole [discovery](discovery.md) migration exists to
preserve, carried one layer up into a running-system surface.
+4 -2
View File
@@ -1,7 +1,9 @@
# The protocol namespace
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. Not yet implemented —
the migration plan at the end is the work list.*
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P3 of the
migration plan at the end have landed (the envelope, the registry and the
`ServiceId` flag-day, and restriction stage one); P4 and P5 are the remaining
work list.*
How a program finds, connects to, and is restricted from the things it talks to.
Three ideas, kept deliberately separate:
+1 -1
View File
@@ -27,7 +27,7 @@ packages whose build.zig calls `build_support.userBinary` (with `.threaded =
true` where a binary spawns threads) and get packed into the initial-ramdisk;
new syscalls extend [abi.zig](../../system/abi.zig) `SystemCall` + a
`library/kernel` wrapper; test services live beside the code they exercise and
register a `ServiceId` if they must be looked up.
bind a `/protocol/test/...` name if they must be reachable.
## How to verify along the way
+8 -6
View File
@@ -270,13 +270,15 @@ stays single-threaded and lean.
- *Handles do not cross threads.* The handle table lives on the `Task`
([scheduler.zig](../../system/kernel/scheduler.zig)), so a handle number is meaningful
only to the thread that created it — thread A's endpoint handle `3` is not thread B's.
A thread that needs to reach an endpoint another thread owns looks it up
(`ipc.lookup(service)`) to install its **own** handle to the same underlying endpoint.
This is how the display's mouse-listener thread reaches the compositor loop's endpoint
to poke it awake (docs/display.md).
A thread that needs to reach an endpoint another thread owns opens the name
(`channel.openEndpoint("display")`) to install its **own** handle to the same
underlying endpoint — an ordinary client open, with no special mechanism for the
fact that the provider happens to be this process. This is how the display's
mouse-listener thread reaches the compositor loop's endpoint to poke it awake
(docs/display.md).
- *IPC syscalls that touch shared kernel state now serialize under the big kernel lock.*
`create_ipc_endpoint`/`ipc_register`/`ipc_lookup` allocate from the kernel heap and
mutate the global service registry, endpoint refcounts, and handle tables. Those paths
`create_ipc_endpoint` allocates from the kernel heap and
mutates endpoint refcounts and handle tables. Those paths
were unlocked because a single-threaded process could not race itself; a multi-threaded
one can, from two cores at once. They now take `sync.enter()` like `call`/`reply_wait`/
`send` already did — the kernel heap has no lock of its own (heap.zig: "every kernel
+3 -3
View File
@@ -137,15 +137,15 @@ Grouped as `abi.zig` groups them:
| process | `danos_exit`, `danos_yield`, `danos_sleep`, `danos_spawn`, `danos_process_enumerate`, `danos_process_kill`, `danos_process_exit_reason`, `danos_process_subscribe`, `danos_process_signal`, `danos_signal_bind` |
| threads | `danos_thread_spawn`, `danos_thread_exit`, `danos_current_core`, `danos_futex_wait`, `danos_futex_wake`, `danos_thread_self`, `danos_thread_join`, `danos_set_thread_pointer` |
| memory | `danos_mmap`, `danos_munmap`, `danos_dma_alloc`, `danos_dma_free`, `danos_shared_memory_create`, `danos_shared_memory_map`, `danos_shared_memory_physical` |
| ipc | `danos_endpoint_create`, `danos_ipc_register`, `danos_ipc_lookup`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` |
| ipc | `danos_endpoint_create`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` (naming is not a syscall: a provider binds its contract at the registry and a client resolves `/protocol/<name>` — see [protocol-namespace.md](protocol-namespace.md)) |
| devices | `danos_device_enumerate`, `danos_device_claim`, `danos_device_register`, `danos_mmio_map`, `danos_irq_bind`, `danos_irq_ack`, `danos_msi_bind`, `danos_io_read`, `danos_io_write` |
| time | `danos_clock`, `danos_wall_clock`, `danos_timer_bind` |
| diagnostics | `danos_debug_write` (leveled, kernel-stamped records), `danos_klog_read`, `danos_klog_status` |
| filesystem naming | `danos_fs_resolve`, `danos_fs_node`, `danos_fs_mount`, `danos_fs_unmount` (naming only — file DATA still crosses the vfs-protocol IPC, see below) |
The constants that ride alongside the calls — mmap protection bits, DMA
flags, notification badge bits, `ExitReason`, `Signal`, well-known service
ids, `page_size`, the IPC message maximum — move to the public header too:
flags, notification badge bits, `ExitReason`, `Signal`, `page_size`, the IPC
message maximum — move to the public header too:
they are wire values a Rust program needs verbatim. What stays private in
`abi.zig` is exactly the thing the vDSO exists to hide: the `SystemCall`
numbers and the trap convention.
+5 -5
View File
@@ -36,10 +36,10 @@ plain `main` checkout always tells the truth about where the work is.**
| | |
|---|---|
| Working on | **P3** — green at 109/109, adversarial review running, not yet committed |
| Branch carrying it | `feat/security-group-2` (pushed to origin) |
| On `main` | through the group 1 merge (`f3bc23c`): Phase 0, PM, H1 |
| Committed on the branch, awaiting the group 2 merge | P1 (`1ff0991`), P2 (`1379b69`) |
| Working on | **P4a** — protocol rebase onto `envelope.Define` |
| Branch carrying it | `feat/security-group-3` (cut next) |
| On `main` | Phase 0, PM, H1, P1, P2, P3 — groups 1 and 2 merged |
| Awaiting merge | nothing |
| Suite | 109 cases, all passing |
| Last updated | 2026-08-01 |
@@ -64,7 +64,7 @@ group boundary.
the manifest gained a third permission, `supervise`, which names an authorized
supervising task per contract and is deliberately **open-only**, leaving P2's
bind attestation and every refusal it makes untouched (suite 109/109)
- [ ] **merge** group 2 → main, push
- [x] **merge** group 2 → main, push
- [ ] **P4a** — clean protocols rebased onto `Define` (vfs, block, display, scanout, input)
- [ ] **P4b** — misfit protocols rebased (device-manager, power, usb-transfer)
- [ ] **P4c** — harness subscriber lift + badge-scoped per-client integers