Third driver converted. It no longer claims the id from argv[1] — the
manager holds the device and names it in the call that creates the process,
so it is held before the driver's first instruction.
display-reattach is the case that matters here: it kills the driver and
watches the compositor re-attach to the fresh scanout. It passes, so the
restart path survives the fused grant — the manager re-takes the device when
the driver dies and hands it to the replacement.
ps2-bus and discovery are NOT converted, and the reason is recorded as open
question 9 rather than worked around. Both need a device nobody assigned
them. ps2-bus ignores its argv[1] entirely: it finds the controller by
walking the table for PNP0303, then claims a second device, the PNP0F13
mouse node, which it also finds itself — so it holds two devices and was
assigned at most one, while system_spawn carries one. discovery claims the
acpi-tables node it locates itself, because it is what produces the device
tree and there is nothing to assign at that point.
One thing worth checking before designing an answer: devices.csv maps both
PS/2 hardware ids to ps2-bus, so the manager may already be spawning two
instances where the driver expects one. If so the fix is smaller than it
looks.
D6 stays blocked — closing device_claim with these two still depending on it
would stop the machine booting.
Suite 118/118.
Delegation moves out of onHello and into the spawn itself. The manager holds
the hardware and names it in the call that creates the driver; the kernel
checks the device is the caller's to give, then hands it over as part of
making the child.
The reason is the window. A transfer after spawning always leaves an
interval in which the child is running and does not yet hold its device. It
would close on QEMU every time and open occasionally on a machine with
different core counts and timing — the exact failure shape this track exists
to delete, and not one worth introducing while removing the others. Fused
into the spawn there is no interval: the child does not exist until it holds
the device.
Ownership is checked BEFORE the child is created, so a refusal leaves
nothing running rather than a driver without the hardware it was spawned
for. The IOMMU confinement moves with the device, as it does on the transfer
path. systemCall6 is added for the sixth argument; r9 was free, and abi
gains a no_device sentinel matching the protocol's.
No driver had to change to receive a device, which is what makes this
better than requiring every driver to hello: ps2-bus keeps its legacy
status, and discovery — which has no assignment at all, since it is what
produces the device tree — is unaffected.
The attacker fixture now tries the spawn as a back door: name someone else's
device, and both the spawn and any child must be refused. Verifying that
assertion exposed a bug in the fixture itself. The kernel case's pass marker
was "device-authority: ok", which matches the FIRST per-assertion line, so
its wait loop exited before any failure was printed — the case would have
passed with failures in it, and had been able to since D2. The verdict lines
now carry a distinct VERDICT prefix, and with the ownership check removed
the case genuinely fails. A green test that cannot go red is worse than no
test.
Suite 118/118.
maximum_devices = 64 is gone. It was a guess about someone else's computer,
and because it was shared, one driver's enumeration starved every other —
which is how an AMD Ryzen booted with a working display, no USB and no
storage. The table now grows from the kernel heap. It was always built after
heap.init; nothing ever prevented this except it having been written static
first.
What replaces it is an allowance charged to the registrar, so a driver
looping device_register exhausts its own and every other driver carries on.
It is declared as what it is — a runaway detector, NOT a security boundary.
A quota generous enough never to bite a real machine is still generous
enough to be unpleasant, and it is not trying to be the defence; delegation
is. What this catches is a legitimate driver in a loop, early, attributably,
and without collateral. Reaching 4096 is a bug report, not a tuning request.
The initial block is 8, deliberately small. Sizing it for a typical machine
would mean the growth path never ran on the hardware we test on and only
woke up on someone else's larger machine — the exact failure shape this
track exists to stop. At 8 it grows several times every boot; disabling
growth now fails the suite with the HPET not fitting, which is the Ryzen
failure in miniature.
The comptime coupling assert added earlier fired, and was right to. confined
(one slot per device id) and domains (the IOMMU's own translation pool) were
sized by the same constant only because device ids happened to stop at 64
too. Two unrelated quantities: confined now grows with the device table,
while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi
report how many domains they support, and reading it is phase 4. The assert
existed for exactly this and did its job.
Suite 118/118.
pci-bus joins usb-xhci-bus in receiving its device from the manager rather
than claiming the id it found in argv[1]. Its hello moves ahead of the ECAM
mapping, since that is where the bridge now arrives, and its hello was
already mandatory so nothing about its failure behaviour changes.
isDelegated compared whole strings, which silently missed this driver: the
manager records the boot-snapshot match as the bare "pci-bus" and a
devices.csv match as the full "/system/drivers/pci-bus". pci-bus was then
neither claiming nor delegated and died on "ECAM mmio_map failed". It now
matches on the last path component. Reintroducing the whole-string compare
breaks usb-xhci-bus instead of pci-bus — the two spellings swap which driver
loses — so usb-hid is the case that catches it, not pci-scan.
pci-scan asserts the delegation on the initial bring-up AND after the
restart drill, with the device id backreferenced so both must name the same
device. That is what proves the manager re-takes a device when its driver
dies and hands it to the replacement, which is the property the whole
supervision design rests on.
The remaining three claimants are NOT converted, and the plan records why
rather than working around it. ps2-bus and the acpi service never hello at
all, which device-manager.md states deliberately ("legacy drivers ... not
yet required to hello"), so delegating to them means either promoting them
out of legacy or giving the grant a delivery point that is not hello.
virtio-gpu hellos best-effort by design — "standalone bring-up has no
manager" — and delegation would make it mandatory. Both are decisions, not
mechanical steps.
Consequence: D6 is blocked, because device_claim cannot be closed off while
three claimants still depend on it. D7-D9 are unaffected — they concern what
the kernel stores and how its table is sized.
Suite 118/118.
The first driver to stop claiming its own hardware. The device manager holds
the controller and transfers it in the hello reply, so its matching becomes
authoritative instead of advisory — until now the driver claimed the id it
found in argv[1], and any process could have claimed the same integer first.
The manager claims before it spawns, so there is no window in which anything
else could take the device, and transfers in onHello using invocation.sender
— the kernel-stamped task id, which cannot be forged by the caller. hello is
synchronous, so the transfer has completed before the reply lands: no gap
between being told yes and holding the thing.
usb-xhci-bus's hello moves from after controller bring-up to before anything
that needs the device, which is the bring-up reorder the design predicted.
It is the first member of an explicit delegated set, so every unconverted
driver keeps claiming exactly as before and the suite stays green; the set
and device_claim both go at D6. D3 and D4 could not be separated and the
plan records why: the moment the manager claims, any driver still calling
device_claim is refused, and D3 applied to nothing changes no behaviour and
cannot be tested.
This step introduced a regression and the incremental conversion is what
caught it. confineDevice runs inside systemDeviceClaim, so a device arriving
by transfer was never confined for its new owner. Three IOMMU+USB cases
failed on the driver's DMA rings going unbound, and two worse consequences
were latent: a manager death would have torn down a domain a live driver was
using, and a driver death would have leaked one. iommu.reassign now moves
the confinement with the device, keeping the domain and its attachment
intact so it never translates through nothing. Converting all five drivers
at once would have produced the same three failures with five suspects.
A log line of mine claimed "holding controller device N" before anything
verified it — it printed even in the failure case, where the driver held
nothing. Reworded to state only what is known there: where the registers
are.
usb-hid asserts the delegation with the device id backreferenced, so the id
delegated and the id the driver ends up with must match. Emptying the
delegated set fails it with "hello acknowledged" then "mmio_map failed".
usb-hub failed once in a full run and has passed six times since (four
isolated, two full) — recorded in the plan as a suspected instance of the
known intermittent AP fault, not dismissed, since this step did shift boot
timing.
Suite 118/118.
The audit's sharpest finding was structural, not a bug: a fully green suite
had hidden six real defects because it contains no attacker. Every device
case asserts that a driver handed its own hardware can drive it. None asked
what a process handed NOTHING can do.
device-authority-test is that process. It is spawned with no device and
asserts what it therefore cannot do: it cannot give away a device another
task holds, nor a free one, because the kernel's rule is that you may give
away what you hold and the device's state is irrelevant to a process holding
nothing. Asserted across every device the machine actually has, so it cannot
pass by accident of which one happened to be free at boot — six on QEMU,
none of them its.
A positive control runs first. device_enumerate works from this process, so
the refusals below it are decisions rather than a syscall path that is
simply broken here; without it, "everything failed" would read identically
to "the assertions are meaningless". A nonexistent device is refused as
NoSuchDevice rather than NotHeld, because a refusal that cannot name its own
rule is what cost a debugging session on the Ryzen.
What it deliberately does not assert, and says so in its header:
device_claim is still first-come-first-served at this point in the run. That
is the hole D6 closes, and the claim half of the invariant joins this
fixture then. Asserting it now would be writing a test that documents the
bug.
Verified to discriminate: removing the holder check flips "every transfer by
a non-holder is refused" while the positive control keeps passing.
Suite 117 -> 118.
The mechanism behind delegation, which device-manager.md named as the step
after hello: the device manager claims what discovery seeded and hands each
device to the driver it matched, so assignment stops being
first-come-first-served.
It is a MOVE, not a copy. A claim is exclusive (driver-model.md, invariant
1), so the giver stops holding the device the instant the receiver starts.
That is why this is a new syscall rather than the M13 capability path, where
a passed handle is shared refcounted — exclusivity cannot be expressed that
way.
The kernel's whole rule is that you may give away what you hold. It has no
notion of which task is the device manager and deliberately gains none: a
binary name inside the kernel is not something that cannot safely live in
user space. A recipient that does not exist is refused, because a device
moved to nobody would be unreachable for the rest of the boot — nothing
un-holds a device but task death.
Three errnos, each naming its own rule: ENODEV no such device, EPERM you do
not hold it, ESRCH no such recipient.
Nothing uses it yet. The five claimants move across one at a time in D4-D5,
so the suite stays green throughout and a regression names the driver that
caused it.
Ten assertions, verified to discriminate: removing the ownership check flips
four of them, including the giveaway that an illegal transfer then blocks
the legitimate claim behind it.
Suite 116 -> 117.
Both device-setup paths issued a successful Enable Slot and then returned
null if allocateDevice failed, without disabling it. A slot the driver
forgets is one the controller never reissues, so each attempt lost one
permanently for the boot. The hub path did it with no log line at all.
Both now release the slot through a shared disableSlot, extracted from
tearDownDevice, and the hub path warns like the root-port path does.
tearDownDevice also now frees the interface list. That allocation arrived
with the previous commit, so an unplug would have leaked it — found while
reading the teardown path for this fix rather than by a test.
No regression test, and it is recorded as open question 5 rather than
implied. After the slot count became the controller's own figure, reaching
this path needs more devices than the controller has slots: QEMU offers four
against sixty-four. What was verified is that the new path RUNS correctly —
pinning tracking to 2 with four devices attached produced "port 6 setup: no
free device slot", the first two devices enumerated normally, and no Disable
Slot error or timeout appeared, which is how disableSlot reports failure.
Suite 116/116.
max_interfaces was 4. A composite device — a headset, a webcam with audio, a
dock, a multifunction printer — routinely has more, and the fifth did not
merely go missing. parseConfiguration's cap branch had no `else`, so when the
count was reached `current` kept pointing at interface 3 and the fifth
interface's endpoint descriptors were appended to interface 3's array. A
class driver bound to interface 3 could then be handed an endpoint belonging
to something else entirely, and subscribe or bulk-transfer on it. The
alternate-setting arm one line above cleared `current` correctly, which is
what the cap branch should have done.
Interfaces are now counted from the block in a first pass and allocated to
exactly that number, so the ceiling is bNumInterfaces' u8 — the USB
specification's. The missing `else` is added too, though after this the bug
is unreachable by construction: interface_count cannot reach interfaces.len
mid-parse when the list was sized from the same walk.
max_configured_endpoints was max_interfaces * max_endpoints_per_interface =
16, a derived guess that moved whenever either input moved. It is now 31,
which is the xHCI specification's own limit: a Device Context holds a slot
context plus at most 31 endpoint contexts, because the Context Entries field
addressing them is 5 bits.
max_endpoints_per_interface stays at 4 with its reason recorded — the
usb-transfer wire protocol reports exactly max_reported_endpoints (4) per
interface, so widening it alone would change nothing a class driver sees.
Lifting it is a protocol change.
No direct test, and that is written down as open question 5 rather than
glossed. The parser is pure and wants a host unit test, but
usb-xhci-library.zig imports memory, mmio and time so it cannot be a
standalone test root, and QEMU offers nothing that reaches the path — the
largest device available is usb-audio,multi=on at 2 interfaces and 211
bytes. The alternate-setting path that shares the same `current = null`
logic is exercised by that device.
Suite 116/116.
The driver read the first 512 bytes of a configuration block into a fixed
buffer and parsed those. The block's length is the device's own choice
(wTotalLength, a u16), so anything larger was silently cut: interfaces past
the cut did not exist as far as the host was concerned, while the
SET_CONFIGURATION that follows still configured the device for all of them.
A headset is 500-900 bytes, a UVC webcam 1-3 KB, a multifunction printer
600+.
Now allocated at the declared length, so the ceiling is the field's u16 —
the specification's number rather than one of ours. A block shorter than its
own 9-byte header is refused rather than trusted.
The bring-up line reports the declared length and the bytes actually read,
so a truncation can never again be invisible, and usb-large-descriptor
asserts they match with a backreference.
That case has an honest limit, recorded in its comment: QEMU cannot produce
a block over 512 bytes. The boot keyboard, mouse and stick are 34-44, and
the largest device available is usb-audio in multi-channel mode at 211 —
which is exactly why the suite never caught this, and why it cannot now
reproduce the original trigger. What it does catch is the class: any clamp
below the attached device's block fails it, verified by pinning the buffer
to 128 and watching "config block 211 bytes, read 128" turn the case red.
Suite 115 -> 116.
max_devices was 8, with the comment "QEMU presents a handful; a fuller
machine would grow this" — a number chosen against the test rig, waiting for
a real machine, which is the pattern docs/bounds-track-plan.md exists to
stop.
The driver already knew the true figure. It reads HCSPARAMS1.MaxSlots at
bring-up and writes it straight into op_config, so every slot the controller
offers has always been *enabled*; only the array tracking them was 8. QEMU's
xHCI reports 64, so seven eighths of the controller was live and invisible,
and the ninth device — a keyboard, mouse, webcam, headset, hub and two
sticks reach that without trying — disappeared on a hub-attached path that
logs nothing at all.
The array becomes a slice allocated from max_slots at bring-up. A controller
claiming zero slots cannot address anything, so that is now a dead
controller rather than an empty allocation failing mysteriously later. The
Device Context Base Address Array is a page, 511 usable entries, so it
already covered the 255-slot maximum.
The bring-up line reports both numbers, and usb-hid asserts they are equal
with a backreference rather than a magic number, so the test cannot drift
from the hardware. Pinning tracking back to 8 fails it: "64 slots,
tracking 8".
Suite 115/115.
The convention that tunables live in system/parameters.zig with their
reasoning attached predates this and got 2% compliance — 5 of 235. A
convention with no teeth is how a bare `const maximum_devices = 64` reached
an AMD desktop and cost it USB and storage. This is the same rule with a
gate behind it.
tools/check-bounds.py finds every bound-shaped declaration — a `maximum_*`
const with a literal value, or a type with a literal array length — and
requires the five-field block above it: what it counts, who decides its
size, what it protects, what happens at the limit, and how anyone finds out.
The at-limit vocabulary is closed: refuse, degrade, truncate, grow. There is
deliberately no way to spell "silent", no way to spell "drop", and nothing
meaning "allow", so the behaviours that did the damage cannot be written
down. Truncation is legal only carrying a marker the reader can see, which
is why klog_maximum_message qualifies and a USB descriptor cut at 512 bytes
does not.
An array length that names a declared bound is not itself a bound; only
literal lengths are flagged, which pushes ceilings toward having names.
The 273 that predate the rule are allowlisted so this lands without a
tree-wide sweep in front of it, and that list may only shrink: declaring a
bound means deleting its line, and the check fails on a stale entry too.
Nothing may be added.
Wired into `zig build test` and available alone as `zig build bounds`. Not
in the default build — it reads the whole tree, and a red bounds check
should not stop you booting a kernel.
Five are now declared rather than allowlisted. Writing them out is its own
argument: maximum_devices reads "protects: nothing — this is a sizing guess
about someone else's computer", and maximum_tasks now carries the fact that
it has been raised twice, each time by something that outgrew it.
Verified the gate refuses an undeclared bound, a declared one using
forbidden vocabulary, and an allowlist entry that has since been declared.
Suite 115/115.
An AMD Ryzen booted to a working compositor with no USB and no storage,
and the log said only "register refused". A tree-wide audit of every
compile-time ceiling followed: 235 of them, 139 on quantities the machine
or a file decides rather than us, 5 documented anywhere, 171 silent when
reached. docs/fixed-bounds-audit.md has the inventory.
Errno attribution. The errno space was split between the kernel and the
envelope, free to drift; it is now one list in system/abi.zig, restated on
both sides, with a comptime check in library/device/driver where the two
halves are visible. device_register's six refusals and device_claim's three
are distinct codes, so a bus driver can say which rule stopped it, and
BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles
found against registered instead of counting refused functions as found.
Idempotency ordering. The child cap was checked before the identity match,
so a restarted bus was refused its own devices — the supervision restart the
system leans on ratcheted toward a degraded machine. A re-registration
consumes no slot and is now admitted first.
IOMMU fail-closed. confineDevice returned success for a device id past the
confinement table, leaving the device outside every domain while the caller
believed it confined — unreachable only while ids stop at 64, which both the
inventory move and a hardware-reported domain count would change. It refuses
now, and the coupling to the broker's device cap is a comptime assert rather
than a sentence in a comment.
PCI apertures. The bridge's MMIO apertures are derived from the holes in the
firmware memory map, and the derivation copied sub-4 GiB entries into a
fixed [64] array and skipped the rest. A skipped region is not merely lost:
the gap finder concludes it is free, so a real machine's 60-200 entry map
yields an aperture over live RAM, and containment then admits a child BAR
covering kernel memory. Rewritten to walk the map in place, with the hole
finder extracted as a pure function and driven by a synthetic 100-entry map
in a new test case. Both new tests were verified to fail on the old code.
parameters.zig gains the rationale it was missing and loses a stale sentence
pointing at the wrong file; vdso.md documents the errno space, including
EPEER, which had no written meaning anywhere.
docs/os-development/bounds.md is how a ceiling is declared from here.
docs/bounds-track-plan.md is the plan to remove the ones we invented.
Suite 114 -> 115.
A Ryzen 3 3200G triple-faulted on its first timer tick after reaching init.
Three defects in a chain, each hiding the one beneath it.
STAR's SYSRET base was 0x10, so SS came back as base+8 = 0x18 with RPL 0
while CS carried RPL 3. Intel ORs RPL 3 into SS on SYSRET; AMD only does so
for CS. Ring 3 ran fine — RPL is not checked on data access — and died the
moment an interrupt tried to IRETQ back, where SS.RPL must equal CS.RPL.
The base now carries the RPL (0x13), as Linux does.
Two fixes below it, both of which made the first one unreadable:
scheduler() read IA32_GS_BASE and dereferenced it without testing for zero,
so every fault reporter faulted in turn — a panic inside a panic, and the
machine reset before printing anything. Cast after the null test, plus a
re-entrancy guard in the panic handler.
NT is now masked in SFMASK alongside the rest, and isr.s exports
isr_return_iretq at the faulting instruction so a frame dump can say which
IRETQ died and print the CS/SS it was about to load. That dump is what
identified the RPL mismatch.
SMAP makes the rule the copy layer has followed since it was written into a
rule the hardware keeps. A ring-0 read or write of a user page now faults,
so any code that reaches for a user pointer directly fails the first time it
runs rather than the first time someone attacks it — and the suite becomes
the enforcement test, because every case exercises the kernel with the bit
on. Nothing had to be fixed to turn it on, which is the retrospective proof
that the nine stragglers converted earlier were all of them.
The interrupt entry needed one instruction first. Hardware does not clear
the alignment-check flag on its way into a handler, and ring 3 sets that
flag freely, so a process could have taken an interrupt with SMAP suspended
for the duration. The system call path was already covered — its flag mask
clears it — but the interrupt path needed a `clac`, which cannot simply be
assembled in: it is an invalid instruction on a processor without SMAP, and
danos boots on those too. So the entry ships as a three-byte NOP and is
patched at boot, through the physmap, because the kernel maps its own text
read-only.
The ordering that makes that safe is enforced rather than described: the
patch sets a flag, and no core will set the SMAP bit until it is true. A
translation that fails, or bytes that read back wrong through the address
they will actually be fetched from, leave the machine unhardened and saying
so — which is the same posture the IOMMU takes, and better than enforcing
over an entry path that cannot comply. The patch runs before interrupts are
enabled and before any second core exists; a comment says so, because the
three bytes pass through an encoding that must never be executed and a
future change that moves this later has to deal with that first.
Suite 114/114, with a case that reads a user page from ring 0 and requires
the fault, and the multi-core case asserting every core that ran work had
the bit — the same shape SMEP got, for the same reason: CR4 is per-core, and
a hardening is only as wide as its narrowest core.
SYSRETQ with a non-canonical RIP raises a general protection fault in ring
0 — on the kernel stack, an instruction after the swapgs that installed the
user's GS base. It is one of the better-known escalation primitives, and
ring 3 reaches it without any kernel bug at all: the processor saves the
address of the instruction after SYSCALL, so a program whose SYSCALL is the
last two bytes of the last canonical page returns to the first
non-canonical address. The new test does exactly that.
The exit path now sign-extends the return address from bit 47 and compares;
if the value changed, it returns through IRETQ instead, which commits the
privilege change before fetching the new address, so the fault arrives from
ring 3 and the process dies like any other. Four register-only operations
and a branch that a correct program can never take — it could not have
executed at a non-canonical address in the first place. Bit 47 is the right
pivot because danos builds four-level page tables and nothing sets the
five-level bit; a future port must move the pivot, and the comment says so.
SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate,
does not clear the nested-task flag, so the kernel had been running every
system call with whatever ring 3 last chose — harmless while the only exit
was SYSRETQ, and a question worth not having now that one exit is IRETQ.
The kernel is never nested; ring 3 still gets its own flag back.
Suite 113/113. The new case asserts the refusal counter rather than the
dying process: the emulator we test on kills it either way, so only the
counter distinguishes a guard that ran from one that did not.
SMEP turns the classic escalation — divert kernel control flow into a page
the attacker wrote — from a silent takeover into an immediate fault with
the offending address in the log. The bit is per-core state, so it is set
where the syscall MSRs already are: in the per-CPU bring-up both the boot
processor and every application processor run on their way in. A core that
climbed the trampoline without it would be a hole no boot log would show,
which is why the SMP case now reads CR4 on each core it lands on and
requires every one of them to be hardened, not just the one that printed
the banner.
Enabling it that early is only safe because nothing ring 0 executes is
mapped for ring 3, and that had to be established rather than assumed:
kernel text carries only its ELF flags, the physmap is no-execute, the
trampoline page is mapped supervisor and the core running it has not
enabled the bit yet, and the boot processor turns it on while still on the
loader's tables — which map nothing user-accessible at all. The one
indirect call in the kernel takes a kernel address.
The CPUID probing that was scattered across the timer code becomes a small
shared helper, since the feature question is now asked from two places and
each wanted the same maximum-leaf guard. Absence is tolerated and reported,
like the IOMMU: danos still boots on a machine without the feature, and
says which one it is.
The test harness starts asking QEMU for a CPU that has the bit at all —
its default model has neither SMEP nor SMAP, so the code would otherwise
have been unreachable in every run. No case behaved differently under the
richer model.
Suite 112/112, with a new case that maps an executable user page, calls
into it from the kernel, and requires the fault the CPU is supposed to
raise.
Three services had each written the same thing and got it three different
ways: input polled the process list to notice a dead subscriber, and only
when someone else subscribed; the power service never noticed at all; the
device manager noticed drivers but not subscribers. The harness owns the
table now, driven by the events a protocol declares — it registers on the
reserved verb, frames each event once, posts to everyone interested without
waiting on any of them, and reclaims a slot when the kernel says its owner
died. Interest masks moved to the envelope, so a subscriber that wants only
mice asks the same way everywhere.
Two consequences the plan had not foreseen. The device manager now hears a
supervised child's death twice, once as its supervisor and once as a
subscriber, so restart backoff counted every crash twice and gave up after
half as many; it retires the id before counting. And the kernel's published
exit table had eight slots for what is now six subscriptions in a plain
boot, so it holds sixteen.
The other half is a hole the design named early and left standing: a
backend handed out a small integer and then honoured it from anyone. A
process that guessed a file's node id read another client's file; a display
layer had no owner at all, so any client could reconfigure or destroy any
layer; a USB device token was never checked against the client that opened
it. Each is now bound to the task that opened it, and a wrong owner gets
exactly what an unknown id gets — the refusal must not become the oracle
the identical answers elsewhere were designed to remove. Closing a file
changed with it: it used to succeed unconditionally, which would have told
a caller which ids existed.
Suite 111/111, with a new case in which one process holds a file and a
layer, hands both ids to a second process, and finds them untouched after
that process has tried everything with them.
These were the awkward ones. Each began with an operation packed into a
single byte — two of them with a version wedged in beside it — so there was
no wrapping them: the layouts had to be rebuilt. The device manager's own
enumerate and subscribe become the reserved verbs that mean the same thing
everywhere, its replies lose three status structs the envelope already
carries, and a device id becomes the packet's target. Power drops the
version it repeated on every request, because describe is the handshake,
and stops claiming a 64-byte ceiling it never needed for calls. USB moves a
control transfer's data to the packet tail in both directions, which makes
the status length the transferred length and retires a field that had been
saying the same thing twice.
The danger in this one was not the protocols but their readers. Init
recognised a power button by two bytes at the head of a message, the ACPI
service dispatched on the first byte, the xHCI driver read its operation
with a raw integer load, and the HID drivers reinterpreted a report
wholesale — none of which would have failed to compile once the layouts
moved. They would simply have stopped: no shutdown on the power button, no
reports from the keyboard. Every one of them now reads through the
generated types, and the shutdown gate that answers only a subscriber is
the same code it was.
Two sizes were decided by measuring rather than assuming. The child-added
message is both a request and the event broadcast to subscribers, and
alignment rounds it to 48 bytes, which puts its packet exactly on the
64-byte push floor — a test pins that, because a field added carelessly
would now overflow it. The interrupt report gives up eight bytes of inline
room to make space for the header; the two drivers that produce reports
send eight and four.
Suite 110/110.
The folded header stops being a rule in a document and becomes the layout
on the wire. Verbs number from sixteen, leaving describe, enumerate,
subscribe and unsubscribe reserved and answered the same way by every
provider — none of them writes a line to do it. What each protocol used to
carry in a field of its own now travels in the header: a vfs node and a
display layer are the packet's target, and a reply opens with a status the
envelope stamps rather than one each protocol spelled for itself.
Display gains the most. One forty-byte request had served eleven verbs, so
attach_scanout smuggled stride through x, refresh through y and format
through colour, and every coordinate crossed as a bitcast. Per-operation
structs end all three: the fields have their own names and their own signs,
and the tile payload grows to 224 bytes because the prefix shrank. Scanout
loses a message maximum of 64 it had no business declaring — it answers
calls, and the floor for a call is 256 — and virtio-gpu stops hard-coding
that number at its harness.
Two changes are semantic rather than notational. A directory now ends at an
entry with no name, because the fixed part of a reply always travels and a
zero-length reply no longer exists to mean anything. And input joins the
service harness, the last loop in the tree that answered no ping and heard
no terminate; its subscriber table, its pruning and its fan-out are the
same code, and a shutdown now asks it to stop instead of killing it.
A new conformance case reads the registry's own listing and asks every
protocol it finds for its name, its version and its verb count, then offers
a verb nobody defines and requires -ENOSYS — the envelope's promise,
checked against providers rather than against itself. What it cannot reach
in that boot it names on the serial line instead of passing quietly.
Suite 110/110.
The registry consults the open rows it has been parsing since P2, so
reaching a contract now takes a grant as well as a binding. A caller
without one is answered exactly as it would be for a name nobody ever
bound: same status, same empty reply, same absent capability, byte for
byte, and no log line on either path — klog_read is ungated, so a line on
one and not the other would be the oracle the design set out to remove.
Refusal and absence being one answer is what lets a supervisor later
narrow, fake or park a child's namespace without the child learning what
it was denied.
The manifest gains a third permission for a shape the plan did not
foresee: attestation is one hop, but the driver tree is three deep — the
PS/2 keyboard and mouse are spawned by ps2-bus, which the device manager
spawned — so no row could name them and PS/2 input would simply stop.
A supervise grant lets a delegate vouch for what its children *reach*,
never for what they claim; the bind path is untouched, and the laundering
deputy is still refused.
The review found the receive side of a rule this track had already
written down. Every process holds a sendable handle to the registrar —
resolve installs one for anyone who asks — and ipc_reply_wait never asked
who owned the endpoint, so a stranger could dequeue there: take the
provider endpoints riding bind requests, and answer other clients' opens
in the registrar's name. Receiving is the owner's privilege, like binding
a signal or a timer; sending remains anyone's.
Suite 109/109.
A protocol is reached by name now, not by a compile-time integer. Init is
PID 1 and already knows which binary it started, so init serves /protocol
as a vfs backend: bind claims a contract with the provider's endpoint
attached, open answers with that endpoint as the reply's capability, and
readdir lists what is bound with the task and binary behind it. The kernel
reserves the prefix — nothing may mount over it, under it, or unmount it —
and ServiceId, ipc_register and ipc_lookup are gone, their syscall numbers
left vacant.
A bind is authorized by who the caller *is*: the kernel-stamped binary
together with the supervising task's identity, matched against
/system/configuration/protocol.csv. Identity, not spelling — spawn is
ungated, so an attacker can run any bundled binary, and a name-only rule
would have let it launder grants through an init of its own making. A name
a live process holds is refused to everyone else; a dead one's is released.
Three review rounds against a hostile ring-3 process found what 108 green
tests could not, because the suite contains no attacker. Publishing init's
supervision endpoint as the registry put PID 1's mailbox in every process's
hands, where two forged bytes reached the shutdown path: privileged traffic
is now believed only from the task that holds the contract it speaks for.
A capability arriving on a request outlived every path that ignored it,
one handle per call until the table was full — in init, and in the harness
ten services share — so the arriving capability is owned by the turn and
released unless a handler says otherwise. And the kernel let anyone holding
an endpoint handle aim signals, timers, exit notices and interrupts at it:
binding now requires having created it.
Suite 108/108. The new protocol-registry case asserts eleven properties,
each one an attack that must fail.
A new user-memory module owns every kernel touch of a user buffer:
copyFromUser, the new copyToUser, and the resolve behind both. The walk
accumulates the U/S and writable bits down all four levels with the MMU's
own AND rule — folding a 2 MiB leaf in before it resolves and refusing a
1 GiB leaf outright — so a copy honours what ring 3 itself would be
allowed, closing the presence-only trust model the IPC layer carried since
bring-up. It then confirms the frame is physmap-backed, because that is how
the copy reaches it: an mmio_map'd BAR passes the permission walk and would
otherwise fault ring 0 on an alias the physmap never mapped, on the IPC path
as much as the new one.
The nine stragglers that dereferenced user pointers raw now route through
it, so a bad pointer returns -EFAULT where it used to fault the kernel.
The write direction restructures its callees around kernel bounce buffers:
scheduler and devices-broker enumerate from a slot cursor (a task exiting
between chunks can neither duplicate nor lose an entry), klog_read drains
the ring in chunks, and fs_node stages headers and names contiguously.
fs_resolve copies out before installing the endpoint handle, so a faulting
copy cannot strand a capability; its out-capacity bound no longer adds an
unbounded ring-3 length to the base, which wrapped and trapped the kernel's
own overflow check. debug_write reads the caller's message once.
Suite 107/107 (new user-memory case: seven bad pointers refused, each
paired with a sound call that must still succeed).
/etc/init.csv and /etc/devices.csv become /system/configuration/*.csv (the
repo's etc/ moves to system/configuration/, mirroring the runtime tree),
/var/log becomes /system/logs, and /mnt/usb becomes /volumes/usb. The
kernel VFS gains a carve-out so FAT may serve exactly /system/configuration
and /system/logs beneath the initrd-backed /system while /system and /test
themselves stay unshadowable; FAT's single /var mount splits into those two
rewritten mounts. The kvfs readdir check learns /system's third child and
the ramdisk spawn sweep skips the configuration tree.
Suite 106/106.
The module-to-domain table in build-support duplicated what each
domain's build.zig already states with its addModule exports. userBinary
now resolves each named import by searching the packages the binary
declared in its own build.zig.zon (b.available_deps), which also makes
the zon the literal include path: an import can only be satisfied by a
domain the binary claims, and naming a module whose domain is missing
fails the build graph with the domain to declare. build-support is down
to the recipe alone. All build variants green; manifest unchanged.
VT-d and AMD-Vi are x86 hardware, but lived in the architecture-neutral
kernel tree and leaked further: the core's public Kind enum named both
vendors, and the ACPI parser read the VT-d version/capability registers
(raw volatile MMIO inside table discovery). Now the vendor backends
live in architecture/x86_64/ behind architecture.iommu — the core hands
over the discovery facts plus an injected environment (frame allocation
+ the log sink, the same pattern enablePaging uses) and receives the
hardware vtable back, so the backends never import kernel internals and
an ARM port supplies its SMMU with no core change. Discovery keeps
table facts only; the live-unit register check moved into VT-d detect
(version reading zero now stays fail-open). The unused kindOf() is
gone. Log shapes the harness pins (iommu online, DANOS-IOMMU-FAULT)
are unchanged; all five IOMMU QEMU cases pass.
display-client and input-client (files and module names), so a service,
its wire contract, and its client never share a name: `display` the
service, `display-protocol` the contract, `display-client` a program's
view of it. The nine consumers' imports and their packages' declared
lists follow (regenerated from the source scan); build-support's
module_homes table carries the new names.
Every binary's build.zig now names precisely the modules its source
imports (derived by scanning each artifact's sources, transitively
through same-directory files), and its zon carries only the domains
those come from — kernel stays implicit (the root shim + link script
live there). build-support's userBinary resolves each name through one
module-to-domain table (module_homes); Domains/domains()/defaultImports
and the raw recipe entry point are deleted. An undeclared @import is a
compile error (verified: injecting @import("xkeyboard-config") into
logger fails with 'no module named ... available within module
program'), and e.g. xkeyboard-config now appears in exactly two
manifests — the two keyboard drivers. Availability never bloated the
emitted binaries (Zig compiles only what a program imports); this makes
the declared interfaces honest. Production and -Dtest-case manifests
byte-identical; all build variants and standalone package builds
green.
initial_ramdisk's spawn-everything sweep counted the /etc data files
(devices.csv, init.csv) as spawnable programs — BadElf ever since the
boot tree started ferrying them — so it now counts only the /system and
/test trees. The init test waited for two raw user writes before
checking the last write for the heartbeat text; init's boot chatter
(heap ok, the /etc/init.csv lookup) satisfies the count long before the
first beat, so it now waits for the heartbeat itself. Both cases pass
again; these failures predate the build-packages work.
init, fat, display, display-demo, device-manager, input, logger, and
the two discovery fillers (acpi, fdt — each exporting an artifact named
"discovery"; the root -Ddiscovery picks which ships) convert to binary
packages on the pci-bus template. init's serial heartbeat flag rides
the dependency options (the root forwards its -Dserial). fat's and
display's unit tests move into their packages and the root aggregate
delegates to them. Boot-image file list unchanged.
build-support gains the domains-based userBinary: the default import set
(the library/kernel concern modules, driver/service clients, mmio,
acpi-ids, xkeyboard-config) is assembled from the domain packages, and
the root shim + user link script come from the kernel package directory.
system/drivers/pci-bus is the first binary package: a ~15-line
declarative build.zig naming only its extras (device-manager-protocol,
pci-class); the root build consumes the artifact for the boot image and
the driver also builds standalone from its own directory. Boot-image
file list unchanged.
The second hardware backend. The IOMMU core, DMA-region capabilities, and
per-device enforcement are unchanged; this adds AMD-Vi (IVRS) as an
alternative to Intel VT-d (DMAR) under the same Backend vtable.
- parseIvrs records the IOMMU control-register base from the first IVHD;
the platform layer gains iommu_is_amd, and the core picks the backend by
vendor at init. VT-d and AMD-Vi are mutually exclusive on real hardware.
- iommu-amd.zig: a 2 MiB device table (every DTE zeroed = deny-all until a
device is claimed), AMD native-format page tables (4 KiB leaves), a
command buffer (INVALIDATE_DEVTAB_ENTRY / INVALIDATE_IOMMU_PAGES /
COMPLETION_WAIT) and an event log for faults. The DTE forwards
interrupts unmapped, so MSI passthrough works exactly as on VT-d.
- The boot log and the iommu self-test are now vendor-aware.
**UNTESTED on real AMD hardware** — danos is developed on Intel, so this
is validated only against QEMU's amd-iommu, and every log line and doc
says so. QEMU quirk handled: its amd-iommu does not observe the
COMPLETION_WAIT store form, but consumes the command ring synchronously on
the tail-register write, so invalidations are already applied by the time
we poll — the backend warns once and proceeds.
Cases: amd-iommu (detection + scratch-domain walker) and
amd-iommu-usb-storage (full storage stack through AMD device-table
translation with per-grant capabilities), both green. 106/106.
This completes the IOVA/IOMMU-enforcement track: per-device DMA domains on
both vendors, with buffers reachable only through delegated capabilities.
Replaces L2's interim DMA pool (every buffer reachable by every claimed
device) with true per-grant confinement: a device reaches only buffers
whose capability was delegated to its driver.
Kernel:
- DmaRegionObject (handle kind 2): a delegation token naming a dma_alloc'd
region, passable across processes on the IPC cap slot like an endpoint
or shared-memory object. Frames stay owned by the allocating address
space (freed on dma_free/teardown as before); the token carries a `dead`
flag so a stale downstream handle can no longer bind a freed region.
- dma_alloc gains the dma_shareable flag: it returns a capability handle
in r8 and every region is tracked in a registry. A task's own regions
auto-bind into the devices it claims (its rings just work); foreign
buffers are bound explicitly.
- dma_bind / dma_unbind / handle_close syscalls (51-53). dma_bind maps a
held region (or shared-memory) capability into a claimed device's domain;
it is idempotent. handle_close reclaims a table slot (raised 16 -> 32).
- dma_free and task death unmap a region from every domain and invalidate
BEFORE its frames return to the allocator — the stale-IOTLB use-after-
free window, closed structurally.
Protocols (flag-day): block gains attach, usb-transfer gains dma_attach —
each carries a region capability on the cap slot. fat allocates its bounce
buffer shareable and attaches it; usb-storage allocates its transport
buffers shareable, attaches them to the controller, and forwards fat's
capability downstream; usb-xhci-bus binds and closes; virtio-gpu binds its
shared scanout surface. The physical addresses on the wire are unchanged
(identity IOVA), so no register-programming code moved.
Cross-process DMA (fat -> usb-storage -> xHC) now flows only through
delegated capabilities. iommu-usb-storage / iommu-usb-hid / iommu-fault
all green under per-grant enforcement; 104/104 overall (fail-open paths
unchanged).
Replaces L1's shared blanket identity domain with a private translation
domain per claimed PCI function. A device now reaches only:
- the DMA pool: every dma_alloc'd region, mapped into every claimed
device's domain (poolAdd/poolRemove, driven from the dma_alloc and
dma_free syscalls). This keeps the cross-process buffer handoff
working (fat's bounce buffer reaches the xHC) while blocking the
kernel, page tables, process heaps, MMIO, and unallocated RAM.
- its own firmware reserved region (RMRR), seeded at confine time.
The pool is the honest interim: devices can still reach one another's
DMA buffers. The DMA-region capability layer (next) narrows it to
per-grant reachability.
dma_free unmaps from every domain and invalidates BEFORE the frames
return to the allocator, closing the stale-IOTLB use-after-free window.
Driver death tears down its domains (detach + free tables) before the
broker claims and DMA frames are released.
New iommu_fault_drain syscall (+ driver.iommuFaultDrain) forces pending
fault records to the log on demand. The new iommu-fault case proves it:
a claimed e1000e is programmed to DMA-fetch its TX ring from an unmapped
page; VT-d faults the access (bdf 00:03.0 addr 0x1000 reason 0x6) and the
system stays alive. 104/104.
First enforcement step of the IOVA track. A vendor-neutral IOMMU core
(iommu.zig) drives an Intel VT-d backend (iommu-intel.zig) to give DMA a
real translation layer instead of the fail-open free-for-all M16 left.
- Boot posture is now stated explicitly: "iommu online (Intel VT-d)"
with version/agaw/rmrr, or "none present - DMA fail-open (unisolated)".
- DMAR parsing extended to select the INCLUDE_PCI_ALL unit (real Intel
PCs put an iGPU-scoped unit first) and record single-path-endpoint
RMRRs; multi-hop scopes and extra DRHDs are counted and warned, never
silently dropped.
- Translation is enabled at boot into a blanket identity domain (all RAM
+ RMRRs, 2 MiB leaves). PCI functions are enumerated post-boot by the
ring-3 pci-bus driver, so a device is attached to the domain when its
driver claims it (confineDevice, with claim rollback if confinement
fails) and detached on driver death, before broker release and DMA
frame teardown. Unclaimed devices are non-present: their DMA faults.
- Interrupt remapping stays off, so MSI writes to 0xFEE00000 bypass
translation and the interrupt-driven xHC keeps working.
- devices-broker gains pciAddressOf (derives BDF from the config-space
ECAM offset), unclaim, and forEachPciFunction.
Faults are drained and logged rate-limited as DANOS-IOMMU-FAULT.
Cases: iommu extended (translation on, scratch-domain map/resolve/unmap,
zero idle faults); new iommu-usb-storage and iommu-usb-hid run the full
storage + input stacks through translated DMA with MSI intact. 103/103.
library/device/pci is now the complete generic floor a leaf PCI driver
needs, instead of just what virtio-gpu used:
- pci-class: capability IDs, MSI/MSI-X/power-management/PCI-Express
register layouts, extended-capability header decode (host-tested),
per-bit command constants, remaining header offsets.
- pci.Function: header accessors, disableBusMaster + interrupt-disable
helpers, findCapability, programMsi/disableMsi, MsiX vector-table
struct, ensurePowerStateD0, functionLevelReset (BAR save/restore),
extended-capability iterator.
Proven by the new pci-caps QEMU case: a pci-cap-test fixture claims an
extra e1000e (PM+MSI+PCIe+MSI-X, no danos driver) and readback-verifies
every surface, including the first driver-side use of msi_bind.
usb-xhci-bus converts from 8 ms event-ring polling to message-signalled
interrupts: plain MSI where offered (real Intel xHC), MSI-X entry 0
otherwise (qemu-xhci has no MSI capability), byte-identical polling as
fallback. The timer survives as a 250 ms port-reconcile/lost-edge tick —
real-hardware USB2 hub debounce still needs it. MSI setup runs BEFORE
controller bring-up: QEMU's xhci only registers the MSI-X vector as used
when IMAN.IE is written while MSI-X is already enabled; interrupts are
silently dropped otherwise (real hardware does not care about the order).
101/101 QEMU cases green; real-hardware smoke passed (mouse works,
boot 2026-07-23T174805Z, plain-MSI branch, vector 33).
Each bus's discovery line now prints the device's would-be /etc/devices.csv row
(bus, base, class, prog_if, vendor, device, subsystem / hid) in uppercase hex,
followed by the human-readable names — so a row for a new driver reads straight
off the boot log, for pci, usb, and acpi alike.
- pci-bus: logFunction moved to registerAndReport (where vendor/device/subsystem
are read from config space) and reformatted to columns + names; subsystem
prints '*' when the function has none. Read unclaimed via the bridge ECAM, so
no per-function claim is needed.
- usb-xhci-bus, acpi: the same column framing on their existing discovery lines.
- qemu_test.py: the acpi-ps2 and acpi-report regexes updated to the new ACPI
format (both verified passing in QEMU).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KJqSiLLchDUUCoXn5jsiwd
Replace init's hardcoded boot_services array with an authoritative,
human-readable service list read at boot, mirroring /etc/devices.csv. Each row
is a service binary path followed by its argv; startup order is file order,
shutdown the reverse. There is no hardcoded fallback — a missing file starts
nothing (the no-ramdisk isolation behavior).
- library/csv: shared CSV helpers (comment strip, field iteration) with unit
tests; device-registry is refactored onto them so both /etc/*.csv files parse
through one place.
- init reads /etc/init.csv into fixed-max static tables (the same pattern as the
device registry) and passes each row's argv straight to spawnSupervised. This
also makes boot-time modes (e.g. device-manager test-usb-restart) expressible
as data rather than hardcoded.
- Diagnose mode (-Ddiagnose omits the display stack) becomes build-time file
selection between etc/init.csv and etc/init-diagnose.csv, so init carries no
comptime service logic; the diagnose build option is dropped from init.
- build.zig: csv module wired; /etc/init.csv bundled into the initrd; the
device-registry tests move to a dedicated block since they now import csv.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KJqSiLLchDUUCoXn5jsiwd
Replace the three hardcoded switch tables (pciDriverForIdentity, hidDriverFor,
usbDriverForIdentity) with an authoritative, human-readable device registry the
manager reads at boot. Matching is most-specific-wins across
base/subclass/prog_if/vendor/device/subsystem/hid, so a precise vendor:device
rule and a generic class rule coexist; an unmatched device is logged, never
guessed. This resolves docs' "matching stays code until the third bus".
- ABI: child_added and DeviceDescriptor gain vendor/device/subsystem; child_added
gains a bus discriminator (BusKind) so PCI and USB class triples match against
the right namespace.
- pci-bus reads vendor/device (config 0x00) and subsystem (0x2C, type-0) and
reports them.
- library/device/registry: freestanding CSV parser + matchDriver() with
specificity scoring; 5 unit tests wired into `zig build test`.
- etc/devices.csv bundled into the initrd; the kernel serves /etc directly, so
the manager reads it before any filesystem service is up (fat starts later).
- virtio-gpu: drop the now-redundant post-spawn 1AF4:1050 re-confirm, since the
registry binds this driver by exact identity.
- Remove the orphaned system/drivers/display driver (unreferenced by build or
registry).
- docs: new devices-csv.md; device-manager.md "matching stays code" resolved.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KJqSiLLchDUUCoXn5jsiwd
The 11 QEMU-suite fixtures lived mixed into system/services/ with their
binaries bundled at /system/tests/<name>. Now the repo path is the boot
path, like every real service: test/system/services/<name>. fat-test
moves out of the fat server's directory into its own; display-demo
stays a boot service.
- kernel VFS: setInitialRamdisk derives one read-only initrd mount per
top-level tree named by the ramdisk entry paths (/system, /test),
registers ancestors generically with self-parented roots, and refuses
backend shadowing of any initrd tree
- EFI loader: the fallback walk also enumerates \test (optional — a
volume without fixtures still boots); manifest and capsule unchanged
- path literals: vfs-test self-open + create probe, process-test
process_enumerate matches, the args-echo argv[0] expectation; the
kvfs case now covers the /test root end to end
- docs: DFHS /test rows, tree diagrams, loader prose, and the location
convention gain the third home; fixed the input-source link
100/100 QEMU cases pass.
Every user binary and the two device-logic library modules (pci, usb) now
`@import` the concern modules directly instead of aliasing through `runtime`:
runtime.ipc/process/time/service/input/block/display -> @import("<module>")
runtime.device / runtime.device_manager -> @import("driver")
runtime.fs -> @import("file-system")
runtime.Thread -> @import("thread").Thread
runtime.system.{write,writeRecord,klog*} -> logging.*
runtime.system.{sleep,timerOnce,wallClock,clock} -> time.*
runtime.system.{spawn*,kill,exit,yield,processes,...}-> process.*
runtime.system.{mmap,munmap,PROT_*} -> memory.*
runtime.dma.* / runtime.shared_memory.* / runtime.allocator -> memory.*
Each consumer keeps its own alias name (e.g. `const device = @import("driver")`),
so call sites are unchanged and there are no collisions with local `driver`
variables. build.zig now injects the concern modules into every user binary via
`default_imports`; pci/usb module import lists were updated to match.
The `runtime` and `system` shims remain for one more step (root.zig still uses
runtime); they are deleted in C5. Nothing but root.zig imports `runtime` now.
Verified: zig build, zig build test, and 17 QEMU cases (smoke, device-manager,
logger, fat-mount, fat-mutations, usb-storage, usb-hid, display-native,
virtio-gpu, input, thread-spawn, thread-mutex, process-kill, shared-memory,
driver-restart, acpi-ps2, pci-scan).
Follow-up cleanup of the just-moved kernel device code:
- power.zig -> acpi.zig. Its reboot() is built entirely on the FADT reset
register (acpi.power_information) plus the legacy 0xCF9/8042 fallbacks — it is
ACPI reboot, so it becomes acpi.reboot (the "P" in ACPI). platform.reboot
still delegates; soft-off/S5 stays the ring-3 acpi service's job as before.
- device-tree.zig -> fdt.zig. It is a discovery *backend* (the ARM/FDT parser,
a sibling of acpi.zig), not part of the model — renaming it to its actual
subject kills the confusing device-tree / device-model DeviceTree name clash
and makes acpi.zig + fdt.zig read as the two parallel backends.
- Flatten: device-model.zig, acpi.zig, fdt.zig, platform.zig move out of the
system/kernel/devices/ subdir up into system/kernel/, joining devices-broker.zig
(already flat). The kernel dir is a flat pile by convention (only architecture/
is a subdir), so the subdir — and its poor "devices" name — is gone.
Pure restructure; "platform" module name unchanged, all cross-file deps are
relative siblings that moved together. zig build + test green; smoke, discovery,
acpi-parse, acpi-ps2, acpi-report, power-button, orderly-shutdown, reboot pass.
With the shared device data (device-abi, pci-class, usb-abi, usb-ids, acpi-ids,
aml) now in library/device/, what remained in system/devices/ was purely
kernel-internal: the firmware-discovery machinery and the rich pointer-based
device model (platform, device-model, acpi, device-tree, power), reached only
through the "platform" module by three kernel files. It belongs with the kernel.
Move it to system/kernel/devices/. A pure relocation: the "platform" module
name is unchanged and every cross-dir dependency is a module import, so only the
one b.path and some comments move. The top-level split is now clean —
system/kernel/ is the kernel, library/device/ the shared device libraries,
system/{drivers,services} the userspace. The runtime /system/devices concept
(the virtual device tree) is unaffected; system/kernel/devices/ is its
implementation.
Also fixes two comment refs that still pointed device-abi at its pre-Wave-1a
home (system/devices/) — it lives at library/device/model/ now.
zig build + test green; smoke, discovery, acpi-parse, acpi-ps2, acpi-report pass.
mmio is device-driver code, so it joins the other domains under
library/device/mmio/ (module name "mmio" unchanged — a pure relocation, only
the build paths move). And its abbreviated function names are spelled out per
docs/coding-standards.md:
read -> readRegister mb -> memoryBarrier
write -> writeRegister rmb -> readMemoryBarrier
wmb -> writeMemoryBarrier
All call sites updated (virtio-gpu, usb-xhci-library, pci.Function); the two
display-driver placeholders import mmio but use nothing, so they're untouched.
Docs (driver-model graph, README layout, drivers.md, the FHS note) follow the
new path and names.
zig build + test green; virtio-gpu, display-native, display-reattach, usb-hid,
usb-hub, usb-storage, pci-scan pass.
Finish the direct-import convention: a protocol is now aliased to its module
name in snake_case in every file that imports one — consumers and the runtime
client wrappers alike. This kills the last of the alias variety (the generic
`protocol`, plus `power` and `transfer`), so `device_manager_protocol.Hello`
means the same thing everywhere and grepping a module name finds all its uses.
Pure rename (identifier uses only; prose references left untouched via a
guarded pattern). 14 files, 299/299 lines.
zig build + test green; 16 QEMU cases pass (smoke, device-manager,
driver-restart, device-list, pci-scan, usb-hid, usb-storage, display-service,
display-native, display-reattach, input, acpi-ps2, vfs, fat-mount,
power-button, orderly-shutdown).
power-protocol -> library/protocol/power/power-protocol.zig. Its two consumers
(init and the acpi discovery service) import it directly; the
runtime.power_protocol re-export and the runtime module import are dropped (no
runtime client speaks it). The now-empty system/services/power/ is removed.
This completes library/protocol/: every driver<->service wire contract lives
there and is imported by module name, no protocol is re-exported through
runtime, and the alias chaos (dp/sp, mixed protocol/<x>_protocol) is resolved.
virtio-gpu-protocol stays a driver-private relative import (a hardware command
set, not a driver<->service seam — like virtio-pci.zig beside it).
zig build + test green; power-button, orderly-shutdown pass.
display-protocol -> library/protocol/display/, scanout-protocol ->
library/protocol/scanout/. Consumers import both directly, killing the cryptic
dp/sp aliases in virtio-gpu (now display_protocol / scanout_protocol) and the
runtime.display_protocol / runtime.scanout_protocol re-exports. runtime.display
(the client) still imports display-protocol by name; scanout has no runtime
client, so its runtime module import is dropped too.
zig build + test green; display-service, virtio-gpu, display-native,
display-reattach pass.
device-manager-protocol -> library/protocol/device-manager/. Its consumers
(pci-bus, usb-xhci-bus, the acpi discovery service, device-manager, device-list,
crash-test, and the not-yet-built intel-integrated display sub-driver) import the
module directly instead of through runtime.device_manager_protocol, which is
deleted. runtime.device_manager (the hello client) already imported the module
by name and is unchanged.
zig build + test green; device-manager, driver-restart, device-list, pci-scan pass.
input-protocol -> library/protocol/input/input-protocol.zig. Its five
consumers (usb-hid keyboard/mouse, ps2 keyboard/mouse, the input service) now
import the module directly and alias it input_protocol, replacing the
runtime.input_protocol re-export (deleted) and the inconsistent local `protocol`
aliases. runtime.input (the client) still imports the module by name.
zig build + test green; input, acpi-ps2, usb-hid pass.
vfs-protocol -> library/protocol/vfs/vfs-protocol.zig. fat, its only consumer,
now imports the module directly (const vfs_protocol = @import("vfs-protocol"))
instead of through runtime.vfs_protocol, and the runtime re-export is deleted.
The runtime's own client (runtime.fs) still imports the module by name.
zig build + test green; vfs, fat-mount pass.