Author SHA1 Message Date
Daniel Samson ded555961c Merge usb-superspeed-hub: hubs work on real hardware (M22)
USB hub support, root-caused and validated on the user's real machine:
a Keychron keyboard and ROG mouse now enumerate through a Genesys
compound USB3 hub (its USB2 companion + transaction translators) and
type on danos.

Fixes found by reading logs off the stick, each invisible to QEMU:
- SuperSpeed hubs have change bits a USB2 hub lacks; not clearing them
  spun the driver. Clear them (gated on speed) + a per-tick service cap.
- The boot scan runs ~3 ms after the controller reset, too early for a
  USB2 connection to debounce; the SuperSpeed devices train instantly
  and hid the problem. Poll the root ports on the tick (edge-triggered).
- Full-speed devices failed Address Device with a USB Transaction Error:
  USB 2.0 requires a 10 ms reset-recovery interval before SET_ADDRESS.
  Wait it out, and retry Address Device on a transaction error.

Plus richer logs: readable PORTSC decode, USB link-state and speed
names, device string descriptors (maker/product), and a readable name
for every class/subclass/protocol enum in usb-ids. Full suite 92/92.

Open follow-up: 'port 11 enumeration failed' — a third device that did
not come up; to be investigated separately.
2026-07-21 23:48:29 +01:00
Daniel Samson 802d51ba74 usb-ids: human-readable name for every class/subclass/protocol enum
Every enum in usb-ids now has a name function, so any code (logs, tools)
can print the readable value instead of a raw byte: hub.protocolName,
hid.subclassName/protocolName, mass_storage.subclassName/protocolName,
communications.subclassName, wireless_controller.subclassName/
protocolName, miscellaneous.subclassName/protocolName,
application_specific.subclassName (alongside the existing className,
speedName, linkStateName).

interfaceName is now comprehensive across classes — it decodes the
well-known triples a log shows: 'HID boot keyboard', 'Mass Storage
(Bulk-Only)', 'Hub (SuperSpeed)', 'Bluetooth' (E0/01/01), etc. So the
bus log reads e.g.

  interface 0: Mass Storage (Bulk-Only) (8/6/80) registered as device 30
  interface 0: HID boot keyboard (3/1/1) registered as device 33

Host-tested with usb-ids (the name decodings are pinned). Vendor-ID
naming deliberately left out (needs a data table); this is pure standard
class-code decoding.
2026-07-21 23:43:01 +01:00
Daniel Samson 570f6f545c usb: quiet the investigation diagnostics now that hubs work on real HW
The compound-hub investigation is resolved (Keychron keyboard + ROG
mouse enumerate through the Genesys USB2 companion hub and type on real
hardware). Turn the debugging spam back down for main: the full 22-port
PORTSC dump runs only when a scan finds NOTHING (a 'why is this empty'
aid), not every boot; a downstream hub port logs only when a device
actually appears or leaves, not for every empty seed-sweep port. The
readable device names, class names, and port-change events stay.
2026-07-21 23:33:12 +01:00
Daniel Samson da5f404041 usb: post-reset recovery delay + Address Device retry on transaction error
Real-hardware root-cause, from the now-readable log: the tick-poll fix
DID catch the user's full-speed devices (USB2 root ports 3, 9, 11
'device appeared — Full-speed'), but every one failed 'Address Device
completion code 4' = USB Transaction Error. Cause: danos addressed the
device immediately after the port reset, but USB 2.0 (spec 7.1.7.5)
requires a reset-recovery interval (TRSTRCY, 10 ms) before a device
answers SET_ADDRESS. QEMU tolerates the omission; real full-speed
devices do not.

setupDevice now waits 10 ms after a port reset before addressing, and
addressDeviceCommand retries up to 3 times on a transaction error,
re-resetting a root-port device between attempts (xHCI 4.6.5 recovery).
QEMU USB cases still green. Real-hardware confirmation pending — this is
the specific fix for the user's keyboard/mouse failing to address.
2026-07-21 23:25:19 +01:00
Daniel Samson ab732dc455 usb: readable PORTSC decode, link-state names, and device string descriptors
Make the USB diagnostic logs scannable by eye:
- usb-ids gains speedName + linkStateName (USB3 PORTSC link states: U0,
  RxDetect, Polling, ...), host-tested with usb-ids.
- The PORTSC dump decodes the register instead of printing hex flags:
    PORTSC[3] 0x00021203: connected, enabled, link=U0, power=on, SuperSpeed
    PORTSC[5] 0x00020ee1: connected, disabled, link=Polling, power=on, High-speed
  (the raw word stays for reference). A USB2 device reads connected+
  disabled+Polling until reset — so the user's next log shows at a glance
  whether the companion port ever reaches that state.
- The library reads STRING descriptors (readString, UTF-16LE -> ASCII,
  English langid), and the device line now names the maker + product:
    port 5 device: Hub "Genesys Logic USB3.0 Hub" (0x05e3:0x0626), 1 interface(s)
  instead of a bare vendor/product id pair.

Full USB case set green (hub regexes follow the new 'device: <class>
"<maker> <product>"' format). Branch diagnostics for the SuperSpeed
compound-hub investigation.
2026-07-21 23:19:18 +01:00
Daniel Samson 4ea4a040d2 usb-ids: className/interfaceName helpers; readable device classes in the bus log
The logs read 'class 9/0/0' where they could say 'Hub'. usb-ids gains
className (a device/interface class byte -> 'Hub', 'HID', 'Mass Storage',
...) and interfaceName (a HID boot interface -> 'HID boot keyboard' /
'HID boot mouse'), both host-tested with usb-ids. The bus driver's
enumerate + register lines now read, e.g.:

  port 21 device: Hub vendor 0x05e3 product 0x0626, 1 interface(s)
  port 21 interface 0: Hub (9/0/0) registered as device 47
  port 5 interface 0: HID boot keyboard (3/1/1) registered as device 32

so a real-hardware log is scannable by eye. Full USB case set green
(hub regexes follow the new 'device: <name>' format).
2026-07-21 23:15:16 +01:00
Daniel Samson 0b10a637ae usb: poll root ports on the tick — catch a late USB2 connection (edge-triggered)
Real-hardware root-cause: the user's keyboard works in the firmware boot
menu but dies under danos. The port dump shows why — danos scans the
root ports ~3 ms after the controller reset, but a USB2 connection needs
~100 ms to debounce, so the boot scan sees the USB2 companion hub's port
empty (all USB2 ports ccs=0), while the SuperSpeed devices (hub, storage)
train instantly and DO show up. danos then relied on a Port Status
Change EVENT to catch the late USB2 connection, which does not fire
reliably on this hardware.

The tick now polls every root port and reconciles on the empty->connected
EDGE (previous-state tracked per port, seeded from the boot scan so
already-up ports never re-fire), bringing up a device that appears after
the scan without depending on the event. Edge-triggered so a port that
fails to enumerate is not retried every 8 ms. Branch-only until the
keyboard is confirmed on real hardware; the diagnostic dumps stay for
the next boot.
2026-07-21 23:06:49 +01:00
Daniel Samson 10956c6660 usb: [diagnostic] dump the xECP USB2/USB3 port map + all PORTSC
The USB2 companion hub carrying the user's full-speed keyboard never
appears — only the SuperSpeed hub (port 19) and SS storage (port 21)
connect. To find where the USB2 root ports are and whether the companion
is presenting on one, walk the xECP Supported Protocol capabilities
(logging each USB 2.0 / 3.0 root-port range) and dump every port's raw
PORTSC unconditionally (ccs/ped/pls/pp/speed). QEMU confirms the dump
(USB2 ports 5-8, USB3 ports 1-4). Branch-only diagnostic.
2026-07-21 22:56:25 +01:00
Daniel Samson 18408b0666 usb: [diagnostic] dump all root-port PORTSC + log port-change events
Branch-only debugging for the real SuperSpeed compound hub: the spin fix
(27f87cb) stopped the hang — the SS hub's 4 SuperSpeed ports now enumerate
(all empty, correct: the full-speed keyboard isn't a SuperSpeed device) —
but the USB2 COMPANION hub, where the keyboard actually lives, never
appears. Only root ports 19 (SS hub) and 21 (SS storage) connect.

Dump every root port's raw PORTSC at scan (connect/enable/link-state/
speed) and log every root port-change event, so the next boot shows
whether the companion is connected on a USB2 port we misread, connects
late as a port-change event, or is simply absent. Diagnostic logging —
to be removed once the companion path works.
2026-07-21 22:50:11 +01:00
Daniel Samson 27f87cb5ba usb: don't spin on a SuperSpeed hub's change bits; log hub port status
Real-hardware finding: a SuperSpeed hub (the user's Genesys, 4 SS
downstream ports) HUNG the bus driver at boot — it stopped logging with
no fault, and the hub's port power dropped. Cause: a SuperSpeed hub has
change bits a USB2 hub lacks (link-state, BH-reset), and the B4 servicing
— written and tested against QEMU's USB2 hub — never cleared them, so
the hub's status-change endpoint re-reported the same port forever and
the driver spun servicing it, never reaching the USB2 companion hub
where the full-speed keyboard actually lives.

hubPortStatusAck now clears the SuperSpeed-only change features too
(gated on hub speed; a USB2 hub is unaffected), and the tick services a
bounded batch of hub changes (32) before yielding — so even if a hub's
change bits misbehave, the driver cannot spin the machine. A per-port
status log line surfaces exactly what a hub reports, so the next real
boot shows the SS hub's port states and whether the USB2 companion
enumerates. QEMU usb-hub cases still green (the SS clears are gated off
for QEMU's USB2 hub).
2026-07-21 22:44:46 +01:00
Daniel Samson 7ee6033fa4 usb-hid: echo typed characters to the log — a simple keyboard check
There was no way to confirm a keyboard actually registers keys on real
hardware (enumeration binding the driver only proves the device came
up). The keyboard driver now echoes each decoded printable character
(and newline) to its log: type a known phrase and read it back from
usb-hid-keyboard.log, or watch it appear live on screen in a -Ddiagnose
boot (where the kernel console is a log sink). It exercises the whole
path — HID report -> diff -> layout -> character — not just enumeration,
and works identically for a keyboard behind a hub.

A usb-key-echo QEMU case injects a phrase via QMP send-key and asserts
it echoes (stable across repeated runs). Full suite 92/92.
2026-07-21 22:27:47 +01:00
Daniel Samson e52ae242fc usb: hub-downstream disconnect teardown + hub-behind-hub recursion (B4c)
Completes hub support (docs/usb-hub.md). On the bus tick, a downstream
hub-port change now dispatches: a connect enumerates the new device
(B4b), a disconnect tears the old one down — recursively, since a hub
that leaves takes its whole subtree with it (children first), reporting
each interface ChildRemoved and Disable-Slotting the device. Route
strings compose across tiers, so a device two hubs deep enumerates with
a two-tier route.

QEMU's hub — unlike root-port hot-plug — DOES raise downstream
status-change events, so both paths are harness-tested: usb-hub (a
keyboard behind a hub binds usb-hid-keyboard), usb-hub-nested (a
keyboard two hubs deep), usb-hub-unplug (device_del behind the hub tears
it down). Full suite 91/91.

Real-hardware validation of the user's SuperSpeed Genesys hub with
full-speed devices on its USB2 companion is flagged for the user (QEMU's
USB2 hub does not model the compound USB3 hub).
2026-07-21 22:16:42 +01:00
Daniel Samson 9da1a9899c usb: enumerate devices behind a hub — route strings + transaction translators (B4b)
The core of hub support (docs/usb-hub.md): a device on a hub's
downstream port now reaches its class driver exactly like one on a root
port. The hub's interrupt status-change endpoint is armed with an
in-process subscription — completions set the hub's pending-port mask
instead of queuing a class-driver report — and setupHub seeds every
downstream port pending so a STATIC topology (a device present at
power-on) enumerates without waiting on the initial interrupt edge.

On the bus tick, each pending hub port is serviced: GET_STATUS + clear
the change bits, and on a connect reset the port, read the speed, then
enable a slot and Address Device with the composed route string
((parent_route<<4)|port), the inherited root port, and the parent hub's
slot/port as the TRANSACTION TRANSLATOR — so the controller routes a
full/low-speed device's split transactions through the hub's TT. The
device then enumerates and registers its interfaces through the normal
path (a compact topology-unique port key keeps the id tag within its
8-byte cap), recursing setupHub if it is itself a hub.

Verified in QEMU (a keyboard behind a USB2 hub on a second controller):
'hub slot 1 port 1 device vendor 0x0627 ... usb-hid-keyboard: ok
(device 35)'. Full suite 89/89. The user's SuperSpeed Genesys hub +
full-speed keyboard/mouse on its USB2 companion is the real-hardware
target, flagged separately.
2026-07-21 22:06:54 +01:00
Daniel Samson 92b03b8d07 usb: recognize and configure hubs, power downstream ports (B4a)
First slice of hub support (docs/usb-hub.md), handled IN the xhci-bus
because a device behind a hub is reached by the CONTROLLER via a route
string in its slot context — topology only exists inside the driver
that owns the controller. When the scan enumerates a class-9 device it
now calls setupHub: read the hub descriptor for the downstream port
count and TT arrangement, tell the controller the slot is a hub
(Configure Endpoint with the slot add-flag sets the Hub bit, Number of
Ports, and MTT/TT-Think-Time per xHCI 4.6.6), SET_HUB_DEPTH for a
SuperSpeed hub so it can compose route strings, and SET_FEATURE
PORT_POWER every downstream port.

The Device gains topology fields (route string, root port, parent hub
slot/port for the TT) and buildAddressInputContext now fills them, so
the Address Device path is ready for downstream devices — those are
enumerated in B4b. A root-port device gets route 0 and its own port as
the chain root, unchanged behavior.

Verified in QEMU with a hub on a second controller and a device behind
it (a usb-hub QEMU case, since the boot controller's auto-assigned
devices collide on the low ports): 'hub slot 1: 8 downstream ports
powered (USB2 single-TT)'. Full suite 89/89. The user's SuperSpeed
Genesys hub is the real-hardware target, flagged separately (QEMU's
USB2 hub doesn't model the compound USB3 hub).
2026-07-21 21:56:59 +01:00
Daniel Samson 983b4ed05a usb: don't 'correct' SuperSpeed EP0 max packet size — it's an exponent
Regression from the B3 MPS0 work, caught on real hardware: a SuperSpeed
hub (and the SuperSpeed boot stick) failed to enumerate. bMaxPacketSize0
(device descriptor byte 7) is a LITERAL size for USB 2.0 and below
(8/16/32/64) but an EXPONENT for SuperSpeed (9 = 2^9 = 512). The refresh
read the exponent 9 as a size and issued Evaluate Context to set EP0 to
9 bytes, corrupting the control endpoint so every following transfer
failed. SuperSpeed's EP0 is fixed at 512 and needs no correction, so
the refresh is now skipped for speed >= 4; only full/low/high speed,
where the field is a literal that can differ from the speed default,
still run it. QEMU tolerated the wrong MPS0 — real silicon does not
(the class of bug the harness can't reach, flagged real-HW-pending).

This likely also explains intermittent no-storage/no-logs on a normal
boot: the boot stick is SuperSpeed and hit the same corruption. Full
suite green (orderly-shutdown's lone failure was its known QMP-timing
flake — three clean re-runs).
2026-07-21 21:26:56 +01:00
Daniel Samson 0ac07dadc9 usb: hot-plug plumbing, all ports powered, interrupter enabled (M20)
B3 — the runtime lifecycle a hot-pluggable bus needs, plus two init
fixes that runtime device arrival depends on:

- pump() now handles PORT STATUS CHANGE events (silently dropped
  before): it queues the port, and the bus driver brings the port up
  (a device arrived) or tears it down (a device left) on its tick —
  reporting each interface ChildRemoved to the device manager, which
  prunes the node, notifies watchers, and lets the class driver's world
  end honestly, then Disable Slot frees the controller-side state.
- ALL root-hub ports are powered at init, not just those with a
  boot-time device: an unpowered port (PP=0) cannot signal a later
  connect, so a hot-plug would never be seen.
- the interrupter is enabled (IMAN.IE + USBCMD.INTE) while the ring
  stays polled — some controllers only WRITE runtime events to the
  ring when the interrupter is enabled.

Real-hardware validation is flagged for the user: QEMU's qemu-xhci does
not raise a runtime port-change event to a polling driver on device_add,
so the end-to-end hot-plug path can't be exercised in the harness (the
port-change handling itself IS proven — a late boot device's PSCE is
caught and acked). The harness gained qmp_sequence (multi-step QMP
injection with arguments) for when a drivable case exists. Full suite
88/88; the working USB path (enumeration, HID, storage) is unregressed
by the port-power and interrupter changes.
2026-07-21 21:01:41 +01:00
Daniel Samson 4091ea6912 test: boot a per-run copy of the USB image — runs must not inherit leftovers
The guest mutates its boot volume (fat tests create files, the logger
writes /var/log) and QEMU is hard-killed after a match, so booting the
build artifact in place let one run's leftovers fail the next — a stale
TESTDIR tripping the mkdir-duplicate refusal was the lone 87/88 failure
in an otherwise green suite — and dirtied the build cache's own output.
Each case now boots a fresh copy in the work directory.
2026-07-21 20:37:59 +01:00
Daniel Samson d8cf533b73 usb: correct EP0's max packet size from the device; sample speed after reset
Two real-hardware correctness fixes from the M20 list, both invisible
to QEMU's forgiving controller:

refreshMaxPacketSize0 existed only as a comment. The EP0 context kept
the SPEED-DEFAULT max packet size (full-speed: 8) even when the device
declares 16/32/64 — real controllers fault the very first full
descriptor read on the mismatch, which is the standing suspect for
full-speed mice dying on the user's PC. Enumeration now probes the
device descriptor's first 8 bytes (deliverable at any legal MPS0),
and issues Evaluate Context to correct EP0 before any longer transfer,
naming the failure and codes if the controller refuses.

And a USB2 port's PORTSC speed field is only meaningful once the port
reset ENABLES the port: the speed (and the MPS0 default derived from
it) is now sampled after the reset instead of trusting the connect-time
read.
2026-07-21 20:35:36 +01:00
Daniel Samson 1638845a4b usb/fat: transfer events matched by slot+endpoint; storage failures heal
B2 — the 1-in-3 boot-time READ CAPACITY failure, root-caused: the xHCI
library's awaitTransfer claimed ANY unclaimed transfer event as its own
completion. An interrupt-endpoint event whose TRB pointer no longer
matched the armed subscription (an error or stale completion from the
keyboard/mouse polling concurrently with storage bring-up) fell through
and was misread as the bulk transfer's completion — desynchronizing the
mass-storage bulk protocol in controller state that SURVIVED driver
restarts, so every retry failed too. Awaited transfers now match the
event's slot id and endpoint DCI; foreign events are dropped and named.
Twelve consecutive runs of the previously-flaky cases pass; the full
suite is green with none of its old intermittents.

B1 — and when storage does fail transiently, the system now heals
instead of giving up forever: a nonzero exit maps to ExitReason.aborted
(a deliberate FAILURE exit — supervisors restart those with backoff,
unlike a clean .exited), usb-storage exits nonzero when a PRESENT
device fails bring-up, and the fat service no longer blocks its harness
polling for a block device and then dies — it serves immediately
(requests fail politely), retries on a 500 ms timer, and mounts
whenever storage appears, including after a driver restart.
2026-07-21 20:27:55 +01:00
Daniel Samson 7082699f5f gitignore: /var/log — real-hardware log pulls stay out of the tree
(An earlier append landed on a line missing its newline, mangling
'.github/' and silently un-ignoring var/ — which let a stray add sweep
pulled logs into the tree. Repaired, and the logs untracked again.)
2026-07-21 20:09:29 +01:00
Daniel Samson f0611ef8ac logger: logger.log becomes the completeness receipt
Its only content was the redundant announce line — the useful shutdown
marker was written AFTER the final drain, so it reached serial and the
ring but never a file, and a truncated boot could only be inferred from
what was missing. The marker now enters the ring BEFORE the drain, so
the drain carries it into logger.log: a directory whose logger.log ends
with 'shutting down; final flush' is complete through shutdown; one
without it was cut early. The sequence-count epilogue stays serial-only,
after the drain, by design.
2026-07-21 20:05:13 +01:00
Daniel Samson eb6e8edafe apic: time-bound the TSC warp check — 47 s of AP bring-up becomes ~0.3 s
First real per-process logs off the stick (the logging track paying for
itself): kernel.log showed 16-core bring-up costing 47 s — per-core gaps
of 2-21 s — on a machine whose clocksource is the TSC, so every AP runs
the pairwise warp check. QEMU always picks HPET, so the harness never
executed this path at all.

The check was bounded by ITERATIONS: 1<<20 warp ticks, each a locked
read-modify-write on a cacheline two cores fight over — microseconds
under real contention, not the nanosecond the '~1 ms' comment assumed —
and the 1<<32-PAUSE rendezvous 'bound' is ~2 minutes on modern Intel
(PAUSE ~140 cycles). Both are now bounded by TIME measured on the TSC
itself: ~5 ms of pairwise hammering per core (Linux's check_tsc_warp
budget — ample to catch a lagging TSC) and a ~100 ms rendezvous window.
2026-07-21 19:47:53 +01:00
Daniel Samson 59ba95a315 usb: don't reset an enabled SuperSpeed port; name every setup failure
Real-PC diagnose boot (photo + OCR): the boot stick connects at
SuperSpeed on port 21 and 'device setup failed' lands in the SAME
millisecond — an instant failure, not a timeout. setupDevice reset
every port unconditionally; that is required to enable USB2 ports, but
a SuperSpeed port that trained its link is ALREADY enabled (xHCI
advances USB3 ports to Enabled, no reset — spec 4.3), and driving a hot
reset into the live link drops PED mid-reset on real silicon. QEMU
tolerates the spurious reset, which is why the harness never saw it.

An enabled speed>=4 port now skips the reset (a not-yet-enabled SS link
still gets one). Every setup step names its failure — port reset with
the PORTSC value, Enable Slot, device-slot exhaustion, Address Device
with its completion code — so the on-screen transcript of the next
failure identifies the exact xHCI command instead of one blanket line.
block.open's give-up window drops 60 s -> 30 s (the slowest observed
healthy chain completed at ~24 s); a machine whose stick failed setup
should not sit a further minute pretending otherwise.
2026-07-21 19:37:24 +01:00
Daniel Samson f587e7e05e console: page-wrap instead of scrolling — never read the framebuffer
The scroll path copied every pixel row up by one glyph height, READING
video memory — and VRAM reads are uncached-slow on real hardware.
Measured on the 16-core PC: ~90 seconds to bring the cores online,
almost entirely boot-transcript lines each paying a whole-screen scroll
copy. (QEMU never shows this: its 'VRAM' is host RAM.)

When the screen fills, the console now clears and restarts at the top —
writes only, once per screenful. The transcript reads the same as it
streams; only the scrollback illusion is gone, which a boot console
never needed.
2026-07-21 19:22:59 +01:00
Daniel Samson b541921218 build: -Ddiagnose — boot without the display so the transcript stays on screen
The display service claiming the framebuffer suppresses the on-screen
boot transcript — correctly in normal operation, but on a serial-less
machine being debugged, the timeline vanishes just when it matters. A
diagnose image (zig build -Ddiagnose=true) has init skip the display
service and demo: the timestamped transcript stays on screen
indefinitely, and the power button still runs the orderly shutdown (so
the logger's files land when storage works).

First use immediately caught a real bug IN QEMU: an intermittent (~1
in 3) usb-storage READ CAPACITY failure after a ~21 s stall — after
which usb-storage and fat both exit cleanly and nothing retries: one
transient early-boot USB failure leaves the system permanently without
storage (and therefore without logs). That no-retry policy is a prime
suspect for real hardware never mounting /var, and is Track B's first
work item.
2026-07-21 19:15:14 +01:00
Daniel Samson a91365b3d9 kernel: the boot transcript on screen, every line timestamped
Without serial and without working USB storage, a slow real-hardware
boot is undiagnosable — 'stabbing in the dark'. Two changes end that:

The log renderer stamps every line with boot-relative seconds
([  12.045] ...), so every surface — serial, debugcon, and now the
screen — is a readable timeline. And the framebuffer console registers
as an ordinary log sink at boot: kernel AND userspace lines (device
bring-up, fat mounts, logger announcements) show live on screen until
the display service claims the framebuffer, which flips the console's
suppression and silences the sink automatically — the display-owns-the-
screen design is unchanged in normal operation; the console now simply
narrates the part of boot that happens before there IS a display.

On the machine that motivated this, the next boot will show by eye
where the minutes go — including whether the USB chain ever brings
storage up, and whether screen drawing itself crawls (the latent
non-write-combining framebuffer suspect: if these very lines paint
slowly, that's the answer).
2026-07-21 19:09:34 +01:00
Daniel Samson ffa45edc8b boot: the capsule — one-file system image first, manifest and walk as fallbacks
Real-firmware finding: the per-file /system tree walk boots in seconds
under OVMF but stalls for MINUTES on real firmware — the cost is not
bytes (USB 3 moves the ~4 MB instantly) but firmware filesystem
OPERATIONS: ~30 opens, each an uncached directory-chain walk in an
unoptimized firmware FAT driver. This is why every real OS loader
(winload, GRUB) reads many files through its own filesystem code over
Block I/O rather than the firmware's file protocol.

The loader now reads boot/system.img — the bundled binaries packed into
ONE v2 initial_ramdisk (tools/pack-system-image.py, derived from the
same bundled list in the same build graph, so tree and capsule cannot
drift) — with a single open + sequential read, the one firmware file
I/O shape that is fast everywhere. The manifest (open each listed path
by name) and the tree walk remain as fallbacks, so a hand-assembled
stick without the capsule still boots. The running system is identical
in all three cases: the kernel receives the same in-RAM table.

The load phase now brackets itself with unconditional on-screen
breadcrumbs ('EFI: loading the system...' / '...starting the kernel'),
because this phase stalling behind a silent black screen — kernel
status is serial-only by design — already cost a real-hardware
debugging session.

Direction (settled with the user): this capsule becomes the BOOTSTRAP
capsule — kernel + init + the storage-bring-up set — once a
spawn-from-memory syscall lets init and the device manager load
everything else from the stick's real file tree at runtime through
danos's own storage stack: file-granular updates (rebuild one binary,
copy one file), the initramfs shape.
2026-07-21 18:57:33 +01:00
Daniel Samson d446ddd2ed boot: survive hand-written sticks — case-fold the /system walk, skip host litter
The loader ENUMERATES /system now, so it sees whatever the stick's
directory entries literally store — and a hand-copied stick differs
from our generated image: firmware returns bare 8.3 short entries
UPPERCASE (INIT, SYSTEM), and host OSes leave litter next to every file
(macOS '._' AppleDouble forks, .fseventsd). Verified in QEMU/OVMF: an
uppercase-stored volume booted to a dead kernel-only system before this
change and boots fully (display up, logger writing /var/log) after.

The walk now lowers ASCII names (the danos tree is canonically
lowercase; FAT lookups are case-insensitive by definition), skips any
dot-prefixed entry, and treats malformed or unreadable entries as
skip-this-file instead of abort-the-whole-walk. Reader.find compares
case-insensitively as belt and braces.
2026-07-21 18:01:12 +01:00
Daniel Samson 0a4388c3bc Merge branch 'claude/shm-runtime-meaning-4fce7c'
# Conflicts:
#	build.zig
#	system/drivers/virtio-gpu/virtio-gpu.zig
2026-07-21 17:27:39 +01:00
Daniel Samson 0045fd87ba naming: shm → shared_memory — 'shm' is a Unix clipping, not an acronym
Per docs/coding-standards.md (no Unix-abbreviation exception): syscalls
shared_memory_create/map/physical, kernel SharedMemoryObject + handlers,
runtime.shared_memory (library/runtime/shared-memory.zig), the
shared-memory-server/-client test services, the shared-memory QEMU case,
and docs incl. vdso.md's danos_shared_memory_*. 87/87 QEMU tests pass.
2026-07-21 17:26:15 +01:00
Daniel Samson ad0fd52cb8 Merge per-process-logging: tagged ring + logger to /var/log + kernel VFS root
Every process now logs through std.log into a kernel-stamped record
ring, and the logger service persists one file per binary path under
/var/log/<boot-stamp>/ on the flash volume — the real-hardware capture
path (no serial needed). Supporting architecture: the boot medium's
/system tree replaces the packed ramdisk (the EFI loader walks it),
task names are binary paths, FAT creates long file names with a fast
cluster allocator, and the VFS root moved into the kernel
(fs_resolve/fs_node/fs_mount/fs_unmount; resolve + redirect — the
kernel names, userspace filesystems serve). The userspace vfs server
is retired. Full QEMU suite green (88 cases).
2026-07-21 17:13:32 +01:00
Daniel Samson 65a44a5568 docs: per-process logging + the kernel VFS root
logging.md rewritten around the pipeline (std.log -> leveled debug_write
-> tagged ring -> logger service -> /var/log/<boot-stamp>/<binary-path>.log),
why the buffer is a kernel ring rather than a logging server (the
storage stack must log; early boot needs the buffer anyway), the
announce-once and counted-loss disciplines, and the accepted gaps.
vfs-protocol.md: the router is the kernel (fs_resolve names, backends
serve data); endpoint comes from resolve, not a registry lookup; mount/
unmount retired from the wire; the rewrite-prefix mount semantics.
vdso.md: klog_status + the fs naming calls join the future table; the
no-file-I/O line sharpened (kernel resolves names, never blocks on a
filesystem). syscall.md: the placeholder status blurb replaced with the
real call families.
2026-07-21 16:36:27 +01:00
Daniel Samson e186858315 vfs: the root moves into the kernel — resolve + redirect cutover
runtime.fs now routes every path through fs_resolve: kernel-served
/system nodes are read via fs_node (tokens, no open state); everything
under a userspace mount goes straight to the owning backend's endpoint
with the kernel-rewritten mount-relative path — one syscall of naming,
then the unchanged vfs-protocol rendezvous, public API untouched. mkdir/
unlink/rename resolve-then-forward (rename checks both paths land on
the SAME backend); mount is the fs_mount syscall.

The fat server mounts twice — /mnt/usb from the volume root and /var
from its /var subtree — so the logger now writes the FHS path
/var/log/<boot-stamp>/... and swapping the persistent medium later
touches only fat's two mount calls. With clients holding fat's node ids
directly, fat records each handle's owner, checks it, and sweeps a dead
client's handles via the published exit events (the old router's
pattern, now where the state actually lives).

The userspace vfs server and its router die; ServiceId.vfs=1 stays
reserved-retired; protocol.zig moves to system/vfs-protocol.zig (the
wire contract is backend-only now). vfs-test becomes the ring-3 proof
of the kernel VFS (own-binary ELF magic through /system, read-only
refusals, listing); vfs-client-death becomes the fat sweep test over
the full storage chain, with a ring-scanning check (the last-write
buffer is too racy under a chattering tree).
2026-07-21 16:27:06 +01:00
Daniel Samson 7c5645fe48 kernel: VFS root — mount table, /system from the initrd, fs syscalls
system/kernel/vfs.zig is the resolve+redirect router: fs_resolve (#46)
walks the kernel mount table; a path under the kernel-backed /system
mount (the initrd, seeded by setInitialRamdisk with a derived directory
table) yields a permanent stateless node token served by fs_node (#47)
— read/status/readdir with copy-out, initrd reads lock-free — while a
path under a userspace mount yields the backend's endpoint (installed
in the caller's table, DEDUPLICATED so 16 slots can't be exhausted by
repeated resolves) plus the rewritten mount-relative path; the caller
then speaks the unchanged vfs-protocol rendezvous directly. The kernel
never blocks on a userspace filesystem, holds no open-file state, and
refuses create-intent on the immutable /system.

fs_mount (#48) is the syscall form of the old router's op-6 cap-pass
(possession of the backend handle is the capability; an optional
rewrite prefix maps the mount into the backend's namespace — how /var
will reach the flash volume); fs_unmount (#49) removes one. A dead
backend's mount clears lazily on resolve.

Dormant this milestone: the userspace vfs still serves runtime.fs
unchanged; the kvfs QEMU case covers the kernel side (resolution, ELF
magic read-through, /system listing, read-only + unknown refusals)
until the M-G cutover exercises the syscalls end-to-end.
2026-07-21 16:15:05 +01:00
Daniel Samson 127ea2dad9 logger: per-process log files under /var/log/<boot-stamp>/
The logger service drains the tagged ring every 250 ms and demultiplexes
it into one file per process on the flash volume:

    /mnt/usb/var/log/2026-07-21T150434Z/system/services/fat.log
    [     0.214] mounted FAT (fat32, 128992 clusters, partition lba 0)

The boot stamp is the wall-clock anchor from klog_status (FAT-safe, no
colons); the kernel's records go to kernel.log; each line carries the
record's monotonic timestamp and level. Storage is best-effort and
late — the ring buffers a whole boot until the mount appears, then the
first drain writes the backlog, storage-stack records included. Lost
records surface as '-- N records lost --' from sequence gaps. Files
close (= fat's device cache flush) after a 2 s quiet period, bounding
data-at-risk without per-record flush thrash. The logger announces
itself once — a periodic status line would feed the stream it drains.

init spawns the logger last, so the reverse-order shutdown stops it
first and its final drain runs over a live storage chain; log-flush and
init's own DANOS.LOG shutdown flush retire (superseded). runtime.fs
makePath treats components at or above a mount point as router names —
create is best-effort per prefix, the final verdict is exists(path).
2026-07-21 16:05:46 +01:00
Daniel Samson 30d6ea622a fat: long-file-name creation, fast cluster allocation, makePath
createFile/createDirectory now build long-name chains: a mangled STEM~N
8.3 alias (collision-checked per directory), the standard rotate-add
checksum, and 13-UCS-2-per-entry pieces written last-logical-first into
a contiguous free-slot run (found sector-wise, growing the directory as
today). Write order is chain first, 8.3 entry last, so an interrupted
create leaves only skippable orphans. Uppercase-compliant 8.3 names
keep the bare-entry fast path; lowercase names now get a chain so their
exact case survives — matching tools/make-fat-image.py. rename stays
8.3-only (documented; nothing needs more yet).

allocateCluster drops its from-cluster-2 rescan (measured ~1 s/cluster
on a part-full volume, a 37 s shutdown flush in the lost M19 build):
a next-free hint (rewound by frees) plus a FAT-sector LBA cache turn
the scan into one device read per FAT sector.

mkdir now refuses an existing name (no duplicate entries), and
runtime.fs gains makePath (mkdir -p) for the logger's nested per-boot
directories. Engine tests cover the log-directory shape, ~N collisions,
chain unlink/reuse, the 8.3 fast path, and the checksum.
2026-07-21 15:59:06 +01:00
Daniel Samson d0c1b3e45b kernel: serialize test markers through the log lock; render one write per line
tests.zig wrote markers straight to serial, racing user-process records
rendered on other cores — visible as 16-byte UART-FIFO interleave once
the tagged renderer emitted several writes per record. Markers now go
through kernel log print (same serial sink, now under the log lock), the
renderer composes each line into one buffer and hits each sink once, and
the vfs-client-death needle drops the old self-written 'vfs: ' prefix.
2026-07-21 15:53:45 +01:00
Daniel Samson f480c5d790 runtime: std.log for every user binary — kernel-stamped attribution
library/runtime/log.zig wires std.log to the tagged ring: the root shim
installs std_options for every binary (programs may override), logFn
formats one line per record and emits it with its level via
debug_write — the payload no longer carries the process's name; the
kernel stamps identity structurally and the serial renderer prints the
'<binary path>: ' prefix, so the transcript keeps its shape.

Migrate every writeLine/logLine/log helper family (init, vfs, fat,
device-manager, acpi, pci-bus, ps2-bus x3, usb-hid x2, usb-storage,
usb-xhci-bus, virtio-gpu — 104 call sites) to std.log.info, dropping
the hand-written prefixes. A leveled record is a complete line by
contract (raw emissions may still build lines from pieces). Test
fixtures and the display/input services keep raw writes for now — their
ring records are attributed by the kernel regardless.

Harness regexes follow the renamed prefixes (usb-hid-keyboard,
discovery) and path-named restart lines.
2026-07-21 15:39:51 +01:00
Daniel Samson 9f18d8340e kernel: tagged log ring — per-line pid/name/level records, klog_status
Replace the linear keep-earliest RAM buffer with a 512 KiB ring of
framed records (log-ring.zig, host-tested): every debug_write becomes
one record per payload line, stamped by the kernel with the sender's
pid, task name (its binary path), level, per-boot sequence number, and
monotonic timestamp. Attribution is structural — a payload cannot forge
another sender's tag, and newline injection lands inside the forger's
own next record. Oldest records are overwritten when full; sequence
gaps make the loss countable.

debug_write gains a level argument (err/warn/info/debug/raw; old
two-arg callers clamp to raw). klog_read becomes a stream-offset read
that fails once the cursor falls behind the ring's tail; the new
klog_status (#45) returns the cursors plus the boot wall-clock anchor —
what the logger service will name per-boot log directories with.

The log now guards itself with a dedicated spinlock (BKL -> log lock
order, never the reverse); panic paths use a bounded try-acquire and
fall back to sinks-only. Serial rendering keeps the historical
transcript byte-identical for kernel and legacy raw output; leveled
records get a kernel-rendered name prefix. log-flush/init's interim
drains start at the ring tail and write framed records until the logger
service replaces them.
2026-07-21 15:30:28 +01:00
Daniel Samson 0a84e52bf8 tests: match task names by basename; path-tolerant harness regexes
processRunning compared literal names against task names that are now
full binary paths; the acpi-ps2/usb-report harness regexes assumed
unprefixed driver names in device-manager spawn/restart lines.
2026-07-21 15:22:58 +01:00
Daniel Samson 1cf9985da6 boot: fix ramdisk paths lost to FAT name uppercasing; widen driver-name buffers
The EFI loader now ENUMERATES /system rather than opening files by name,
so the on-disk name must be exact, not merely case-insensitively
findable. make-fat-image.py stored 8.3-fitting lowercase names as bare
uppercase short entries ('init' -> INIT), garbling the ramdisk paths;
now any name whose case the 8.3 entry cannot reproduce gets a long-name
chain carrying the exact name.

device-manager's driver table stored names in 24 bytes — silently
truncating '/system/drivers/usb-xhci-bus' and failing the spawn; widen
to 64 (= abi.maximum_process_name). The usb-report harness regex learns
that restart lines name drivers by path.
2026-07-21 15:13:19 +01:00
Daniel Samson 2d0858caf6 boot: the /system tree is the system image — loader-built ramdisk, spawn by path
Retire the build-time ramdisk packer and the packed initial-ramdisk.img.
make-fat-image.py now lays every user binary out at its FHS path on the
boot volume (system/services, system/drivers, system/tests), and the EFI
loader walks \system at boot, packing what it finds into an in-RAM v2
initial_ramdisk whose entry names are full FHS paths. init rides the
table like every other binary: its dedicated handoff fields are gone and
the kernel spawns PID 1 via the same lookup as everyone else
(process.spawnBundled).

system_spawn resolves names by exact path first, then unique basename,
and normalizes argv[0] to the stored path — so task names (and, next,
the tagged log ring's attribution) are honest binary paths everywhere.
initrd v2 rejects the old magic so a stale image fails loudly.

Groundwork for per-process logging (/var/log/<boot-stamp>/<binary-path>.log)
and the kernel-VFS /system mount.
2026-07-21 14:47:26 +01:00
Daniel Samson ec6e888076 display: pace the frame clock by the panel's EDID refresh rate
Both EDID moments the system has are now captured and carried to the
compositor's frame clock:

- EFI: the loader derives refresh from the preferred detailed timing
  (pixel clock / total pixels) while GOP is still alive — the only moment
  it is readable — and hands it through the boot handoff into the
  display0 node's DisplayInfo (new refresh_hz field, 0 = unknown).
- GPU: the virtio-gpu driver derives the same figure from its own EDID
  read and carries it in the attach_scanout announce (request.y).

updateFrameClock() re-derives the interval from the active backend's
info at bring-up and again on every backend change — the boot
framebuffer's clock dies with the GOP floor at upgrade, replaced by the
GPU's rate. Unknown rate defaults to 60 Hz; the result is clamped to
[30, 120] Hz so a mis-parsed EDID can neither starve nor flood the
compositor. Rate only, never phase: without vblank, presents still
free-run (docs/display-v2.md, 'Fenced is not vsync').

Observed in QEMU: OVMF exposes no EDID for the VGA adapter, so the GOP
floor logs 'frame clock 62 Hz (default)' (real firmware does expose it);
the virtio-gpu EDID advertises 75 Hz and the upgrade logs 'frame clock
76 Hz (panel EDID)'. The display kernel test asserts refresh_hz rides
the seeded node; all 7 display QEMU cases pass.
2026-07-21 11:37:29 +01:00
Daniel Samson 4f02f75602 display: a ~60 Hz frame clock — presents are scheduled, not immediate
Client 'present' requests and cursor pokes no longer repaint on the spot:
they accumulate damage and arm a one-shot 16 ms timer, and the tick
composites everything pending as one frame. A fast mouse previously
turned every input event into a full present (100+/s); now any number of
draws and moves inside one interval coalesce into a single repaint.

No backend has a real vblank to pace by (docs/display-v2.md, 'Fenced is
not vsync'), so this is the software stand-in — the same strategy Linux
uses atop virtio-gpu. Bring-up paths that need pixels synchronously
(initialise, self-checks) still present directly.

onNotification now handles the message and timer badge bits
independently: one coalesced badge can carry both, and the old
either/or dispatch would have dropped a tick.

All six display QEMU cases pass.
2026-07-21 11:26:49 +01:00
Daniel Samson 16618d2cdc display: stop calling the fenced present 'vsync' — it isn't
The virtio-gpu present fence completes when the device has consumed the
frame: real completion feedback, and tear-freedom by snapshot semantics.
It is not a vblank — base virtio-gpu 2D has no display-refresh event at
all (Linux fakes one with a timer), so nothing paces presents to the
monitor. The code and docs claimed vsync anyway; now they don't.

- backend.hasVsync -> hasFencedPresent, with an honest doc comment
- marker 'display: vsync present ok' -> 'display: fenced present ok'
  (display-modeset test expectation updated, passes)
- display-v2.md gains a 'Fenced is not vsync' note; the vsync claims in
  both v2 docs are corrected
- true vsync arrives with a native driver's vblank IRQ, or approximated
  by a compositor frame clock
2026-07-21 11:24:23 +01:00
Daniel Samson 15107f54be build: run-x86-64-gpu — the interactive twin of the display-native test
Boots the usual VGA/GOP floor plus a virtio-gpu-pci adapter, so the
device-manager stack spawns the driver and the compositor upgrades to the
native fenced-present backend interactively. QEMU shows one head per
adapter: the live output is the virtio-gpu console in the View menu (the
VGA head freezes at the moment of upgrade). 512M, matching display-native.
2026-07-21 11:22:07 +01:00
Daniel Samson 23f915c593 display: damage-rect list + tile-grid trackers, vectorizable pixel loops, wide WC stores
Tearing mitigation for the GOP floor, attacking the copy window from three sides:

- Damage is no longer one bounding box. Two trackers, A/B-switchable at
  compile time (display.zig damage_mode): DamageList (free-form dirty rects,
  overlap-merged) and TileGrid (fixed 64-px tiles, exact O(1) marking, runs
  coalesced back into rects). Far-apart changes — the cursor here, an
  animating layer there — no longer unite into one huge repaint.
- fillRect/composite/blitTile now work in row spans (@memset/@memcpy), so
  the compiler vectorizes them and ReleaseSafe bounds checks drop to per-row.
- The back->front present streams 8-byte volatile stores (presentSpan);
  Backend.present takes the rect list, so each present copies only what
  changed, faster.

Host tests cover both trackers; the display QEMU cases all pass.
2026-07-21 11:22:00 +01:00
daniel 981a4af7e0 Merge display-mouse-thread: the compositor tracks the mouse (Shape A)
The display becomes danos's first multi-threaded service. A mouse-listener
thread runs beside the compositor loop; the compositor stays the single owner
of the framebuffer, and the cursor it draws is a top-z layer moved from the
listener's accumulated position over a single-slot latest-value CursorChannel
(Thread.Mutex) with a coalesced self-ipc.send poke that wakes the parked
replyWait (docs/display.md, docs/threading.md).

Two kernel-level findings this surfaced, both fixed:
  - IPC handles are per-Task, not per-address-space — a thread reaches another
    thread's endpoint via ipc.lookup() to install its own handle.
  - Concurrent IPC from two threads raced unlocked kernel state (flaky #GP);
    create_ipc_endpoint / ipc_register / ipc_lookup now take the big kernel
    lock like call/reply_wait/send already did (the kernel heap has no lock of
    its own yet — heap.zig: "a lock comes with threads/SMP").

Also decouples display-demo: it no longer draws a cursor (the service owns it)
or blocks its animation loop on a mouse read, so a normal boot shows one
responsive cursor and animation that runs independent of the mouse. The
display-demo test now spawns the input service to keep that independence honest.

New test display-cursor (smp:4) drives synthetic motion end to end. Verified:
zig build, zig build test, and the full 29-case QEMU guardrail suite all green.
2026-07-21 02:41:45 +01:00
73 changed files with 5389 additions and 1826 deletions
+2 -1
View File
@@ -6,4 +6,5 @@ zig-out/
.idea/ .idea/
.claude/ .claude/
.github/ .github/
/var/log/
+252 -49
View File
@@ -2,6 +2,7 @@ const std = @import("std");
const uefi = std.os.uefi; const uefi = std.os.uefi;
const elf = std.elf; const elf = std.elf;
const boot_handoff = @import("boot-handoff"); const boot_handoff = @import("boot-handoff");
const initial_ramdisk = @import("initial-ramdisk");
const build_options = @import("build_options"); const build_options = @import("build_options");
const BootInformation = boot_handoff.BootInformation; const BootInformation = boot_handoff.BootInformation;
const GraphicsOutput = uefi.protocol.GraphicsOutput; const GraphicsOutput = uefi.protocol.GraphicsOutput;
@@ -15,11 +16,11 @@ const MemoryMapSlice = uefi.tables.MemoryMapSlice;
/// The kernel image: /system/kernel. /// The kernel image: /system/kernel.
const kernel_file_name = std.unicode.utf8ToUtf16LeStringLiteral("system\\kernel"); const kernel_file_name = std.unicode.utf8ToUtf16LeStringLiteral("system\\kernel");
/// The init program: /system/services/init. /// The user binaries: everything under /system except the kernel itself. The
const init_file_name = std.unicode.utf8ToUtf16LeStringLiteral("system\\services\\init"); /// loader walks this tree and packs it into the in-RAM initial_ramdisk image —
/// the volume's file structure is the single source of truth (no packed image
/// The initial-ramdisk (the VFS server + drivers), in /boot. /// artifact on disk).
const initial_ramdisk_file_name = std.unicode.utf8ToUtf16LeStringLiteral("boot\\initial-ramdisk.img"); const system_directory_name = std.unicode.utf8ToUtf16LeStringLiteral("system");
/// Physical page size, and the sentinel UEFI uses to seek to end-of-file. /// Physical page size, and the sentinel UEFI uses to seek to end-of-file.
const page_size = 4096; const page_size = 4096;
@@ -64,20 +65,14 @@ fn boot() !noreturn {
const entry = try loadKernel(bs, &boot_information); const entry = try loadKernel(bs, &boot_information);
// Best effort: a volume without /system/services/init still boots (kernel-only). // Best effort: a volume without a /system tree of user binaries still boots
loadInit(bs, &boot_information) catch |err| { // (kernel-only). The tree — init included — becomes the initial_ramdisk.
log("EFI: no /system/services/init ("); loadSystemTree(bs, &boot_information) catch |err| {
log("EFI: no /system binaries (");
logBytes(@errorName(err)); logBytes(@errorName(err));
log(") - booting without user space\r\n"); log(") - booting without user space\r\n");
}; };
// Best effort: the initial_ramdisk (VFS server + drivers) is optional too.
loadInitialRamdisk(bs, &boot_information) catch |err| {
log("EFI: no initial_ramdisk (");
logBytes(@errorName(err));
log(")\r\n");
};
// Build the page tables the kernel starts life on: identity + a physmap of // Build the page tables the kernel starts life on: identity + a physmap of
// low RAM, plus the higher-half kernel image once it links high. Allocated // low RAM, plus the higher-half kernel image once it links high. Allocated
// now, while boot services (and the memory map) are still stable — nothing // now, while boot services (and the memory map) are still stable — nothing
@@ -97,7 +92,7 @@ fn boot() !noreturn {
} }
/// A display resolution in pixels. /// A display resolution in pixels.
const Resolution = struct { width: u32, height: u32 }; const Resolution = struct { width: u32, height: u32, refresh_hz: u32 };
/// Switch the GPU to the monitor's native resolution (when we can determine it) /// Switch the GPU to the monitor's native resolution (when we can determine it)
/// and read the resulting graphics mode into our own framebuffer description. /// and read the resulting graphics mode into our own framebuffer description.
@@ -128,6 +123,10 @@ fn queryFramebuffer(bs: *uefi.tables.BootServices) !boot_handoff.Framebuffer {
// Each pixel is 32 bits, so the byte pitch is 4 * pixels-per-row. // Each pixel is 32 bits, so the byte pitch is 4 * pixels-per-row.
.pitch = info.pixels_per_scan_line * 4, .pitch = info.pixels_per_scan_line * 4,
.format = try pixelFormat(info.pixel_format), .format = try pixelFormat(info.pixel_format),
// The refresh rate rides the EDID preferred timing. If the firmware kept a
// non-native mode it may not describe that mode exactly — but it is the panel's
// own clock, a far better frame-clock seed than a hardcoded 60 Hz.
.refresh_hz = if (native) |n| n.refresh_hz else 0,
}; };
} }
@@ -176,10 +175,12 @@ fn nativeResolution(bs: *uefi.tables.BootServices, handles: []uefi.Handle) ?Reso
return null; return null;
} }
/// Parse the native resolution from a raw EDID block. The first Detailed Timing /// Parse the native resolution and refresh rate from a raw EDID block. The first
/// Descriptor (at byte 54) is the preferred — i.e. native — mode by convention; /// Detailed Timing Descriptor (at byte 54) is the preferred — i.e. native — mode by
/// its active pixel counts are split across low bytes and the high nibbles of /// convention; its active pixel counts are split across low bytes and the high nibbles
/// later bytes. /// of later bytes. The refresh rate is derived, not stored: the descriptor carries the
/// pixel clock (10 kHz units) and the active+blanking extents, and
/// refresh = clock / (horizontal total × vertical total).
fn edidNative(edid: []const u8) ?Resolution { fn edidNative(edid: []const u8) ?Resolution {
if (edid.len < 128) return null; if (edid.len < 128) return null;
// Every EDID begins with this fixed 8-byte header. // Every EDID begins with this fixed 8-byte header.
@@ -193,7 +194,13 @@ fn edidNative(edid: []const u8) ?Resolution {
const w = @as(u32, dtd[2]) | (@as(u32, dtd[4] & 0xf0) << 4); const w = @as(u32, dtd[2]) | (@as(u32, dtd[4] & 0xf0) << 4);
const h = @as(u32, dtd[5]) | (@as(u32, dtd[7] & 0xf0) << 4); const h = @as(u32, dtd[5]) | (@as(u32, dtd[7] & 0xf0) << 4);
if (w == 0 or h == 0) return null; if (w == 0 or h == 0) return null;
return .{ .width = w, .height = h };
const clock_hz = (@as(u64, dtd[0]) | (@as(u64, dtd[1]) << 8)) * 10_000;
const h_blank = @as(u64, dtd[3]) | (@as(u64, dtd[4] & 0x0f) << 8);
const v_blank = @as(u64, dtd[6]) | (@as(u64, dtd[7] & 0x0f) << 8);
const total = (@as(u64, w) + h_blank) * (@as(u64, h) + v_blank);
const refresh: u32 = if (total == 0) 0 else @intCast((clock_hz + total / 2) / total);
return .{ .width = w, .height = h, .refresh_hz = refresh };
} }
/// Open the kernel on the volume we booted from, read it into a pool buffer, /// Open the kernel on the volume we booted from, read it into a pool buffer,
@@ -357,11 +364,36 @@ fn handoff(cr3: u64, entry: usize, boot_information: *const BootInformation) nor
unreachable; unreachable;
} }
/// Read a whole file off the boot volume into a pool buffer that outlives the // --- the /system tree -> initial_ramdisk ------------------------------------
/// loader. The buffer is deliberately NOT freed: it's LoaderData, which the
/// memory-map conversion classifies as reserved, so the kernel identity-maps it /// Cap on bundled binaries. Generous: the tree carries ~30 today.
/// and reads from there. Returns the buffer (pointer + length). const maximum_bundled = 64;
fn loadFile(bs: *uefi.tables.BootServices, name: [*:0]const u16) ![]u8 {
/// How deep the walk goes below /system ("/system/services/x" is depth 1).
const maximum_tree_depth = 3;
/// One binary discovered under /system: its FHS path (UTF-8, '/'-separated,
/// NUL-free) and its contents in a transient pool buffer.
const Bundled = struct {
path: [initial_ramdisk.maximum_name]u8,
path_len: usize,
data: []align(8) u8,
};
/// Gather the boot volume's user binaries into an in-RAM v2 initial_ramdisk
/// image, entries named by full FHS path — the volume's file structure is the
/// single source of truth (no packed ramdisk artifact; init travels in the
/// table like everything else).
///
/// Two strategies, most portable first:
/// 1. /system/manifest (written by the build): each listed path is opened BY
/// NAME — the case-insensitive lookup every firmware FAT driver gets
/// right, and the only file access the pre-tree loader ever used.
/// 2. No manifest: ENUMERATE the /system tree. Portable in principle, but
/// firmware differs in what names enumeration returns (bare 8.3 entries
/// come back uppercase on some drivers), so this is the fallback for
/// hand-assembled sticks, not the primary path.
fn loadSystemTree(bs: *uefi.tables.BootServices, boot_information: *BootInformation) !void {
const loaded = (try bs.handleProtocol(uefi.protocol.LoadedImage, uefi.handle)) orelse const loaded = (try bs.handleProtocol(uefi.protocol.LoadedImage, uefi.handle)) orelse
return error.NoLoadedImage; return error.NoLoadedImage;
const device = loaded.device_handle orelse return error.NoBootDevice; const device = loaded.device_handle orelse return error.NoBootDevice;
@@ -371,40 +403,211 @@ fn loadFile(bs: *uefi.tables.BootServices, name: [*:0]const u16) ![]u8 {
const root = try fs.openVolume(); const root = try fs.openVolume();
defer _ = root.close() catch {}; defer _ = root.close() catch {};
const file = try root.open(name, .read, .{}); // Unconditional breadcrumb (con_out, independent of -Dserial): this phase
defer _ = file.close() catch {}; // is where a slow firmware stalls, and a silent black screen here already
// cost a real-hardware debugging session.
log("EFI: loading the system...\r\n");
// The capsule (boot\system.img) first: one open + one sequential read is
// the only firmware file I/O shape that is fast everywhere. It is already
// the kernel's wire format — hand it over as-is.
if (loadCapsule(bs, root, boot_information)) {
log("EFI: system image loaded, starting the kernel\r\n");
return;
}
var list: [maximum_bundled]Bundled = undefined;
var count: usize = 0;
loadByManifest(bs, root, &list, &count) catch {
count = 0; // a torn manifest read leaves partial entries; start over
};
if (count == 0) {
const system_directory = try root.open(system_directory_name, .read, .{});
defer _ = system_directory.close() catch {};
try walkDirectory(bs, system_directory, "/system", 0, &list, &count);
}
if (count == 0) return error.NoBinaries;
// Assemble the v2 image: header, entry table, then the blobs.
const table_end = @sizeOf(initial_ramdisk.Header) + count * @sizeOf(initial_ramdisk.Entry);
var total: usize = table_end;
for (list[0..count]) |e| total += e.data.len;
const image = try bs.allocatePool(.loader_data, total); // survives the handoff
std.mem.bytesAsValue(initial_ramdisk.Header, image[0..@sizeOf(initial_ramdisk.Header)]).* = .{
.magic = initial_ramdisk.magic,
.count = @intCast(count),
};
var offset: usize = table_end;
for (list[0..count], 0..) |e, i| {
var record = initial_ramdisk.Entry{ .name = @splat(0), .offset = offset, .len = e.data.len };
@memcpy(record.name[0..e.path_len], e.path[0..e.path_len]);
const slot = image[@sizeOf(initial_ramdisk.Header) + i * @sizeOf(initial_ramdisk.Entry) ..][0..@sizeOf(initial_ramdisk.Entry)];
std.mem.bytesAsValue(initial_ramdisk.Entry, slot).* = record;
@memcpy(image[offset..][0..e.data.len], e.data);
offset += e.data.len;
_ = bs.freePool(e.data.ptr) catch {};
}
boot_information.initial_ramdisk_base = @intFromPtr(image.ptr);
boot_information.initial_ramdisk_len = total;
log("EFI: /system tree loaded, starting the kernel\r\n");
}
/// The boot capsule: the bundled binaries as one v2 initial_ramdisk image.
const capsule_file_name = std.unicode.utf8ToUtf16LeStringLiteral("boot\\system.img");
/// Load boot\system.img whole and hand it to the kernel unmodified — it is
/// already the initial_ramdisk wire format. Returns false (capsule absent or
/// unreadable or wrong magic) to let the caller fall back to per-file loading.
fn loadCapsule(bs: *uefi.tables.BootServices, root: *uefi.protocol.File, boot_information: *BootInformation) bool {
const file = root.open(capsule_file_name, .read, .{}) catch return false;
defer _ = file.close() catch {};
const image = readWholeFile(bs, file) catch return false;
if (image.len < @sizeOf(initial_ramdisk.Header) or
std.mem.bytesToValue(initial_ramdisk.Header, image[0..@sizeOf(initial_ramdisk.Header)]).magic != initial_ramdisk.magic)
{
_ = bs.freePool(image.ptr) catch {};
return false;
}
boot_information.initial_ramdisk_base = @intFromPtr(image.ptr);
boot_information.initial_ramdisk_len = image.len;
return true;
}
/// The manifest path, and a scratch limit for its UTF-16 conversion.
const manifest_file_name = std.unicode.utf8ToUtf16LeStringLiteral("system\\manifest");
/// Load every binary the manifest lists, opening each path by name from the
/// volume root. A listed-but-unopenable file is skipped (the kernel reports the
/// absence); a missing manifest errors so the caller falls back to the walk.
fn loadByManifest(bs: *uefi.tables.BootServices, root: *uefi.protocol.File, list: *[maximum_bundled]Bundled, count: *usize) !void {
const manifest_handle = try root.open(manifest_file_name, .read, .{});
var manifest_open = true;
defer if (manifest_open) {
_ = manifest_handle.close() catch {};
};
const manifest = try readWholeFile(bs, manifest_handle);
_ = manifest_handle.close() catch {};
manifest_open = false;
defer _ = bs.freePool(manifest.ptr) catch {};
var lines = std.mem.tokenizeAny(u8, manifest, "\r\n");
while (lines.next()) |line| {
if (line.len < 2 or line[0] != '/') continue;
if (line.len >= initial_ramdisk.maximum_name) continue;
if (count.* == maximum_bundled) return;
// "/system/services/init" -> UTF-16 "system\services\init".
var name16: [initial_ramdisk.maximum_name]u16 = undefined;
var i: usize = 0;
for (line[1..]) |c| {
name16[i] = if (c == '/') '\\' else c;
i += 1;
}
name16[i] = 0;
const file = root.open(@ptrCast(name16[0..i :0]), .read, .{}) catch continue;
defer _ = file.close() catch {};
const data = readWholeFile(bs, file) catch continue;
var entry: *Bundled = &list[count.*];
@memcpy(entry.path[0..line.len], line);
entry.path_len = line.len;
entry.data = data;
count.* += 1;
}
}
/// Recursively collect the regular files below `directory` into `list`. Top-level
/// files (depth 0) are skipped: the only one is /system/kernel, which loadKernel
/// has already consumed and which is not a spawnable user binary.
fn walkDirectory(
bs: *uefi.tables.BootServices,
directory: *uefi.protocol.File,
prefix: []const u8,
depth: usize,
list: *[maximum_bundled]Bundled,
count: *usize,
) !void {
// Each read() on a directory yields one EFI_FILE_INFO; zero bytes means done.
var info_buffer: [1024]u8 align(8) = undefined;
while (true) {
const n = try directory.read(&info_buffer);
if (n == 0) return;
const info: *const uefi.protocol.File.Info.File = @ptrCast(@alignCast(&info_buffer));
const name16 = info.getFileName();
// Convert the (ASCII in practice) UTF-16 name. A hostile-shaped entry
// (too long, non-ASCII) is SKIPPED, never fatal — one odd file on a
// hand-written stick must not cost the whole boot. Names are lowered:
// the danos tree is canonically lowercase and FAT lookups are
// case-insensitive, but firmware ENUMERATION returns whatever the
// directory stores — an 8.3 short entry comes back uppercase ("INIT"),
// which would otherwise poison every path comparison downstream.
var name_buffer: [initial_ramdisk.maximum_name]u8 = undefined;
var name_length: usize = 0;
var name_ok = true;
while (name16[name_length] != 0) : (name_length += 1) {
if (name_length == name_buffer.len) {
name_ok = false;
break;
}
const c = name16[name_length];
if (c > 0x7F) {
name_ok = false;
break;
}
name_buffer[name_length] = std.ascii.toLower(@intCast(c));
}
if (!name_ok) continue;
const name = name_buffer[0..name_length];
// Skip dot entries: "." / ".." and host-OS litter (macOS "._*" AppleDouble
// resource forks, ".fseventsd", ".Spotlight-V100") a copied-onto stick
// accumulates — none of it is a danos binary.
if (name.len == 0 or name[0] == '.') continue;
if (info.attribute.directory) {
if (depth == maximum_tree_depth) continue;
var child_prefix: [initial_ramdisk.maximum_name]u8 = undefined;
const child = try std.fmt.bufPrint(&child_prefix, "{s}/{s}", .{ prefix, name });
const child_directory = try directory.open(name16, .read, .{});
defer _ = child_directory.close() catch {};
try walkDirectory(bs, child_directory, child, depth + 1, list, count);
continue;
}
if (depth == 0) continue; // /system/kernel — already loaded, not bundled
if (count.* == maximum_bundled) return error.TooManyBinaries;
var entry: *Bundled = &list[count.*];
const path = std.fmt.bufPrint(&entry.path, "{s}/{s}", .{ prefix, name }) catch continue; // path too long: skip the file, keep the boot
entry.path_len = path.len;
const file = directory.open(name16, .read, .{}) catch continue;
defer _ = file.close() catch {};
entry.data = readWholeFile(bs, file) catch continue; // unreadable/empty: skip
count.* += 1;
}
}
/// Read an open file completely into a fresh pool buffer that survives the
/// handoff (LoaderData is classified reserved, so the kernel identity-maps it).
fn readWholeFile(bs: *uefi.tables.BootServices, file: *uefi.protocol.File) ![]align(8) u8 {
try file.setPosition(seek_end); try file.setPosition(seek_end);
const size: usize = @intCast(try file.getPosition()); const size: usize = @intCast(try file.getPosition());
try file.setPosition(0); try file.setPosition(0);
if (size == 0) return error.EmptyFile; if (size == 0) return error.EmptyFile;
const image = try bs.allocatePool(.loader_data, size); // survives the handoff const buffer = try bs.allocatePool(.loader_data, size);
var read_total: usize = 0; var read_total: usize = 0;
while (read_total < size) { while (read_total < size) {
const n = try file.read(image[read_total..]); const n = try file.read(buffer[read_total..]);
if (n == 0) return error.UnexpectedEof; if (n == 0) return error.UnexpectedEof;
read_total += n; read_total += n;
} }
return image[0..size]; return buffer[0..size];
}
/// Ferry the init program (/system/services/init) to the kernel. The kernel does the ELF
/// loading itself (into ring-3 mappings) — the loader just carries the bytes.
fn loadInit(bs: *uefi.tables.BootServices, boot_information: *BootInformation) !void {
const image = try loadFile(bs, init_file_name);
boot_information.init_base = @intFromPtr(image.ptr);
boot_information.init_len = image.len;
progress("EFI: /system/services/init loaded\r\n");
}
/// Ferry the initial_ramdisk (the VFS server + drivers) to the kernel, same as init.
fn loadInitialRamdisk(bs: *uefi.tables.BootServices, boot_information: *BootInformation) !void {
const image = try loadFile(bs, initial_ramdisk_file_name);
boot_information.initial_ramdisk_base = @intFromPtr(image.ptr);
boot_information.initial_ramdisk_len = image.len;
progress("EFI: initial_ramdisk loaded\r\n");
} }
/// Validate the ELF, copy every PT_LOAD segment to its physical address, and /// Validate the ELF, copy every PT_LOAD segment to its physical address, and
+185 -119
View File
@@ -221,18 +221,23 @@ fn addKernel(
return exe; return exe;
} }
/// Assemble the bootable FAT32 image (the in-repo Python builder) holding what /// One user binary and its FHS home on the boot volume (and in zig-out).
/// the firmware and loader need off the ESP: the EFI stub, `kernel`, `init`, and const BundledBinary = struct { path: []const u8, binary: std.Build.LazyPath };
/// the initial-ramdisk. Factored so the serial-enabled `run-x86-64` variant can
/// bundle its own serial kernel while sharing the loader, init, and ramdisk — all /// Assemble the bootable FAT32 image (the in-repo Python builder) holding the
/// built once per invocation (the loader's boot breadcrumbs and init's heartbeat /// EFI stub, the kernel, and every user binary at its FHS path — the volume's
/// both follow the top-level -Dserial). Returns the image's LazyPath. /// /system tree IS the system image; the EFI loader walks it at boot and builds
/// the in-RAM initial_ramdisk from it. Factored so the serial-enabled
/// `run-x86-64` variant can bundle its own serial kernel while sharing the
/// loader and user tree (the loader's boot breadcrumbs and init's heartbeat both
/// follow the top-level -Dserial). Returns the image's LazyPath.
fn addBootImage( fn addBootImage(
b: *std.Build, b: *std.Build,
kernel_bin: std.Build.LazyPath, kernel_bin: std.Build.LazyPath,
efi_bin: std.Build.LazyPath, efi_bin: std.Build.LazyPath,
init_bin: std.Build.LazyPath, manifest: std.Build.LazyPath,
initial_ramdisk_img: std.Build.LazyPath, capsule: std.Build.LazyPath,
bundled: []const BundledBinary,
) std.Build.LazyPath { ) std.Build.LazyPath {
const mk_fat = b.addSystemCommand(&.{"python3"}); const mk_fat = b.addSystemCommand(&.{"python3"});
mk_fat.addFileArg(b.path("tools/make-fat-image.py")); mk_fat.addFileArg(b.path("tools/make-fat-image.py"));
@@ -242,10 +247,14 @@ fn addBootImage(
mk_fat.addFileArg(efi_bin); mk_fat.addFileArg(efi_bin);
mk_fat.addArg("system/kernel"); mk_fat.addArg("system/kernel");
mk_fat.addFileArg(kernel_bin); mk_fat.addFileArg(kernel_bin);
mk_fat.addArg("system/services/init"); mk_fat.addArg("system/manifest");
mk_fat.addFileArg(init_bin); mk_fat.addFileArg(manifest);
mk_fat.addArg("boot/initial-ramdisk.img"); mk_fat.addArg("boot/system.img");
mk_fat.addFileArg(initial_ramdisk_img); mk_fat.addFileArg(capsule);
for (bundled) |item| {
mk_fat.addArg(item.path);
mk_fat.addFileArg(item.binary);
}
return fat_image; return fat_image;
} }
@@ -361,7 +370,7 @@ pub fn build(b: *std.Build) void {
// is the first "protocol module" (see docs/driver-model.md); usb/block will // is the first "protocol module" (see docs/driver-model.md); usb/block will
// expose theirs the same way. // expose theirs the same way.
const vfs_protocol_module = b.addModule("vfs-protocol", .{ const vfs_protocol_module = b.addModule("vfs-protocol", .{
.root_source_file = b.path("system/services/vfs/protocol.zig"), .root_source_file = b.path("system/vfs-protocol.zig"),
}); });
// The input wire protocol: the input service's public interface, exposed as its own // The input wire protocol: the input service's public interface, exposed as its own
@@ -443,8 +452,9 @@ pub fn build(b: *std.Build) void {
}, },
}); });
// The initial_ramdisk container format, shared by the kernel (unpacks it) and the // The initial_ramdisk container format, shared by the kernel (unpacks it) and
// build-time packer tools/make-initial-ramdisk.py (produces it). No dependencies. // the EFI loader (packs it in RAM from the boot volume's /system tree). No
// dependencies.
const initial_ramdisk_module = b.addModule("initial-ramdisk", .{ const initial_ramdisk_module = b.addModule("initial-ramdisk", .{
.root_source_file = b.path("system/initial-ramdisk.zig"), .root_source_file = b.path("system/initial-ramdisk.zig"),
}); });
@@ -458,6 +468,10 @@ pub fn build(b: *std.Build) void {
// QEMU test harness (test/qemu_test.py, which asserts on serial markers) turn // QEMU test harness (test/qemu_test.py, which asserts on serial markers) turn
// it on; a flashable `zig build` image leaves it out. See serial.zig. // it on; a flashable `zig build` image leaves it out. See serial.zig.
const serial = b.option(bool, "serial", "Compile the serial-console log sink into the kernel (default: off; run-x86-64 and the test harness enable it)") orelse false; const serial = b.option(bool, "serial", "Compile the serial-console log sink into the kernel (default: off; run-x86-64 and the test harness enable it)") orelse false;
// The diagnose boot: init skips the display service (and demo), so the
// on-screen boot transcript is never suppressed — the full timestamped
// timeline stays on the screen for real-hardware debugging by eye.
const diagnose = b.option(bool, "diagnose", "Boot without the display service so the timestamped boot transcript stays on screen (real-hardware debugging)") orelse false;
// --- Kernel: freestanding x86_64 ELF, jumped to by the bootloader --- // --- Kernel: freestanding x86_64 ELF, jumped to by the bootloader ---
// SSE2 is part of the x86_64 baseline and UEFI leaves it enabled at handoff, // SSE2 is part of the x86_64 baseline and UEFI leaves it enabled at handoff,
@@ -501,16 +515,14 @@ pub fn build(b: *std.Build) void {
// the heartbeat stays present under test. // the heartbeat stays present under test.
const init_options = b.addOptions(); const init_options = b.addOptions();
init_options.addOption(bool, "serial", serial); init_options.addOption(bool, "serial", serial);
init_options.addOption(bool, "diagnose", diagnose);
programModule(init_exe).addImport("build_options", init_options.createModule()); programModule(init_exe).addImport("build_options", init_options.createModule());
const init_install = b.addInstallArtifact(init_exe, .{ .dest_dir = .{ .override = .{ .custom = "system/services" } } });
b.getInstallStep().dependOn(&init_install.step);
// --- initial_ramdisk: a bundle of extra user binaries (VFS server + drivers) --- // --- the rest of the /system tree: services, drivers, test fixtures ---
// Each is built by the same user-binary recipe, then packed into one image by // Each is built by the same user-binary recipe and laid out at its FHS path on
// the host-side make-initial-ramdisk tool. The bootloader ferries the image to the kernel, // the boot volume (see `bundled` below). The EFI loader walks the tree at boot
// which unpacks it and spawns each program (system/initial-ramdisk.zig). // and hands the kernel an in-RAM initial_ramdisk of it (system/initial-ramdisk.zig).
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig"); const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs-test/vfs-test.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig"); const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig");
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig"); const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig"); const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
@@ -540,8 +552,8 @@ pub fn build(b: *std.Build) void {
const display_exe = addThreadedUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display", "system/services/display/display.zig"); const display_exe = addThreadedUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display", "system/services/display/display.zig");
const display_demo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display-demo", "system/services/display-demo/display-demo.zig"); const display_demo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display-demo", "system/services/display-demo/display-demo.zig");
const virtio_gpu_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "virtio-gpu", "system/drivers/virtio-gpu/virtio-gpu.zig"); const virtio_gpu_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "virtio-gpu", "system/drivers/virtio-gpu/virtio-gpu.zig");
const shm_server_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "shm-server", "system/services/shm-server/shm-server.zig"); const shared_memory_server_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "shared-memory-server", "system/services/shared-memory-server/shared-memory-server.zig");
const shm_client_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "shm-client", "system/services/shm-client/shm-client.zig"); const shared_memory_client_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "shared-memory-client", "system/services/shared-memory-client/shared-memory-client.zig");
const fat_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat-test", "system/services/fat/fat-test.zig"); const fat_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat-test", "system/services/fat/fat-test.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig"); const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
// The PCI bus driver decodes each function's class triple to human names in its // The PCI bus driver decodes each function's class triple to human names in its
@@ -579,98 +591,88 @@ pub fn build(b: *std.Build) void {
const input_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig"); const input_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig");
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig"); const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig");
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig"); const process_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig");
const log_flush_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "log-flush", "system/services/log-flush/log-flush.zig"); const logger_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "logger", "system/services/logger/logger.zig");
// The first multi-threaded binary: exercises runtime.Thread over the thread ABI // The first multi-threaded binary: exercises runtime.Thread over the thread ABI
// (docs/threading.md). Built threaded so its shared-memory poll is real. // (docs/threading.md). Built threaded so its shared-memory poll is real.
const thread_test_exe = addThreadedUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "thread-test", "system/services/thread-test/thread-test.zig"); const thread_test_exe = addThreadedUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "thread-test", "system/services/thread-test/thread-test.zig");
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool // Every user binary and its FHS home on the boot volume. There is no packed
// (the container format is trivial, and Python sidesteps std API churn). Args: // ramdisk artifact any more: make-fat-image.py lays each binary out at this
// make-initial-ramdisk.py <out> [<name> <file>]... — one name/file pair per binary. // path on the image, and the EFI loader walks /system at boot and builds the
const mk_run = b.addSystemCommand(&.{"python3"}); // in-RAM initial_ramdisk table from the tree — the volume's file structure is
mk_run.addFileArg(b.path("tools/make-initial-ramdisk.py")); // the single source of truth. Entry names (and hence argv[0] and task names)
const initial_ramdisk_img = mk_run.addOutputFileArg("initial-ramdisk.img"); // are these paths with a leading slash.
mk_run.addArg("vfs"); const bundled = [_]BundledBinary{
mk_run.addFileArg(vfs_exe.getEmittedBin()); .{ .path = "system/services/init", .binary = init_exe.getEmittedBin() },
mk_run.addArg("vfs-test"); .{ .path = "system/services/fat", .binary = fat_exe.getEmittedBin() },
mk_run.addFileArg(vfstest_exe.getEmittedBin()); .{ .path = "system/services/display", .binary = display_exe.getEmittedBin() },
mk_run.addArg("ps2-bus"); .{ .path = "system/services/display-demo", .binary = display_demo_exe.getEmittedBin() },
mk_run.addFileArg(ps2_bus_exe.getEmittedBin()); .{ .path = "system/services/device-manager", .binary = device_manager_exe.getEmittedBin() },
mk_run.addArg("ps2-keyboard"); .{ .path = "system/services/input", .binary = input_exe.getEmittedBin() },
mk_run.addFileArg(ps2_keyboard_exe.getEmittedBin()); .{ .path = "system/services/discovery", .binary = discovery_exe.getEmittedBin() },
mk_run.addArg("ps2-mouse"); .{ .path = "system/services/logger", .binary = logger_exe.getEmittedBin() },
mk_run.addFileArg(ps2_mouse_exe.getEmittedBin()); .{ .path = "system/drivers/ps2-bus", .binary = ps2_bus_exe.getEmittedBin() },
mk_run.addArg("usb-xhci-bus"); .{ .path = "system/drivers/ps2-keyboard", .binary = ps2_keyboard_exe.getEmittedBin() },
mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin()); .{ .path = "system/drivers/ps2-mouse", .binary = ps2_mouse_exe.getEmittedBin() },
mk_run.addArg("usb-hid-keyboard"); .{ .path = "system/drivers/usb-xhci-bus", .binary = usb_xhci_bus_exe.getEmittedBin() },
mk_run.addFileArg(usb_hid_keyboard_exe.getEmittedBin()); .{ .path = "system/drivers/usb-hid-keyboard", .binary = usb_hid_keyboard_exe.getEmittedBin() },
mk_run.addArg("usb-hid-mouse"); .{ .path = "system/drivers/usb-hid-mouse", .binary = usb_hid_mouse_exe.getEmittedBin() },
mk_run.addFileArg(usb_hid_mouse_exe.getEmittedBin()); .{ .path = "system/drivers/usb-storage", .binary = usb_storage_exe.getEmittedBin() },
mk_run.addArg("usb-storage"); .{ .path = "system/drivers/virtio-gpu", .binary = virtio_gpu_exe.getEmittedBin() },
mk_run.addFileArg(usb_storage_exe.getEmittedBin()); .{ .path = "system/drivers/pci-bus", .binary = pci_bus_exe.getEmittedBin() },
mk_run.addArg("fat"); .{ .path = "system/tests/vfs-test", .binary = vfstest_exe.getEmittedBin() },
mk_run.addFileArg(fat_exe.getEmittedBin()); .{ .path = "system/tests/fat-test", .binary = fat_test_exe.getEmittedBin() },
mk_run.addArg("fat-test"); .{ .path = "system/tests/shared-memory-server", .binary = shared_memory_server_exe.getEmittedBin() },
mk_run.addFileArg(fat_test_exe.getEmittedBin()); .{ .path = "system/tests/shared-memory-client", .binary = shared_memory_client_exe.getEmittedBin() },
mk_run.addArg("display"); .{ .path = "system/tests/crash-test", .binary = crash_test_exe.getEmittedBin() },
mk_run.addFileArg(display_exe.getEmittedBin()); .{ .path = "system/tests/device-list", .binary = device_list_exe.getEmittedBin() },
mk_run.addArg("display-demo"); .{ .path = "system/tests/input-source", .binary = input_source_exe.getEmittedBin() },
mk_run.addFileArg(display_demo_exe.getEmittedBin()); .{ .path = "system/tests/input-test", .binary = input_test_exe.getEmittedBin() },
mk_run.addArg("virtio-gpu"); .{ .path = "system/tests/args-echo", .binary = args_echo_exe.getEmittedBin() },
mk_run.addFileArg(virtio_gpu_exe.getEmittedBin()); .{ .path = "system/tests/process-test", .binary = process_test_exe.getEmittedBin() },
mk_run.addArg("shm-server"); .{ .path = "system/tests/thread-test", .binary = thread_test_exe.getEmittedBin() },
mk_run.addFileArg(shm_server_exe.getEmittedBin()); };
mk_run.addArg("shm-client");
mk_run.addFileArg(shm_client_exe.getEmittedBin());
mk_run.addArg("pci-bus");
mk_run.addFileArg(pci_bus_exe.getEmittedBin());
mk_run.addArg("crash-test");
mk_run.addFileArg(crash_test_exe.getEmittedBin());
mk_run.addArg("thread-test");
mk_run.addFileArg(thread_test_exe.getEmittedBin());
mk_run.addArg("device-list");
mk_run.addFileArg(device_list_exe.getEmittedBin());
mk_run.addArg("discovery");
mk_run.addFileArg(discovery_exe.getEmittedBin());
mk_run.addArg("device-manager");
mk_run.addFileArg(device_manager_exe.getEmittedBin());
mk_run.addArg("input");
mk_run.addFileArg(input_exe.getEmittedBin());
mk_run.addArg("input-source");
mk_run.addFileArg(input_source_exe.getEmittedBin());
mk_run.addArg("input-test");
mk_run.addFileArg(input_test_exe.getEmittedBin());
mk_run.addArg("args-echo");
mk_run.addFileArg(args_echo_exe.getEmittedBin());
mk_run.addArg("process-test");
mk_run.addFileArg(process_test_exe.getEmittedBin());
mk_run.addArg("log-flush");
mk_run.addFileArg(log_flush_exe.getEmittedBin());
// Also install the packed binaries to their FHS homes, so zig-out is a true image // The boot manifest: the FHS path of every bundled binary, one per line. The
// of the filesystem — even though at boot they arrive inside the initial-ramdisk. // EFI loader reads THIS by name and opens each listed path by name — FAT
for ([_]struct { *std.Build.Step.Compile, []const u8 }{ // name lookup is case-insensitive and firmware-portable, unlike directory
.{ vfs_exe, "system/services" }, // ENUMERATION, whose returned names vary by firmware (bare 8.3 entries come
.{ device_manager_exe, "system/services" }, // back uppercase on some FAT drivers). The tree walk remains only as the
.{ input_exe, "system/services" }, // loader's fallback for hand-assembled sticks without a manifest.
.{ ps2_bus_exe, "system/drivers" }, var manifest_text: std.ArrayListUnmanaged(u8) = .empty;
.{ ps2_keyboard_exe, "system/drivers" }, for (bundled) |item| {
.{ ps2_mouse_exe, "system/drivers" }, manifest_text.append(b.allocator, '/') catch @panic("OOM");
.{ usb_xhci_bus_exe, "system/drivers" }, manifest_text.appendSlice(b.allocator, item.path) catch @panic("OOM");
.{ usb_hid_keyboard_exe, "system/drivers" }, manifest_text.append(b.allocator, '\n') catch @panic("OOM");
.{ usb_hid_mouse_exe, "system/drivers" },
.{ usb_storage_exe, "system/drivers" },
.{ fat_exe, "system/services" },
.{ display_exe, "system/services" },
.{ log_flush_exe, "system/services" },
}) |entry| {
const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } });
b.getInstallStep().dependOn(&step.step);
} }
const manifest_files = b.addWriteFiles();
const manifest_file = manifest_files.add("manifest", manifest_text.items);
const manifest_install = b.addInstallFileWithDir(manifest_file, .prefix, "system/manifest");
b.getInstallStep().dependOn(&manifest_install.step);
// The initial-ramdisk itself installs to /boot (with the loaders). // The boot capsule: the same bundled list packed into ONE file (v2
const initial_ramdisk_install = b.addInstallFile(initial_ramdisk_img, "boot/initial-ramdisk.img"); // initial_ramdisk format), because a single open + sequential read is the
b.getInstallStep().dependOn(&initial_ramdisk_install.step); // only firmware file I/O shape that is fast everywhere — a per-file tree
// walk measured MINUTES on real firmware. The loader tries this first,
// then the manifest, then the walk; the running system cannot tell the
// difference (it always receives the same in-RAM table). Derived from the
// tree in the same build graph, so the two cannot drift.
const mk_capsule = b.addSystemCommand(&.{"python3"});
mk_capsule.addFileArg(b.path("tools/pack-system-image.py"));
const capsule_img = mk_capsule.addOutputFileArg("system.img");
for (bundled) |item| {
mk_capsule.addArg(item.path);
mk_capsule.addFileArg(item.binary);
}
const capsule_install = b.addInstallFile(capsule_img, "boot/system.img");
b.getInstallStep().dependOn(&capsule_install.step);
// Install every bundled binary to its FHS home, so zig-out is a true image of
// the filesystem — the same tree make-fat-image.py lays out on the boot volume.
for (bundled) |item| {
const install = b.addInstallFileWithDir(item.binary, .prefix, item.path);
b.getInstallStep().dependOn(&install.step);
}
// Boot methods live in boot/, one per way of getting the kernel running. // Boot methods live in boot/, one per way of getting the kernel running.
// Each is its own binary/entry (a loader is built for its own target); today // Each is its own binary/entry (a loader is built for its own target); today
@@ -692,8 +694,10 @@ pub fn build(b: *std.Build) void {
}), }),
.optimize = optimize, .optimize = optimize,
.imports = &.{ .imports = &.{
// The bootloader speaks only the handoff contract — never the user ABI. // The bootloader speaks the handoff contract and the ramdisk
// container it packs the /system tree into — never the user ABI.
.{ .name = "boot-handoff", .module = boot_handoff_module }, .{ .name = "boot-handoff", .module = boot_handoff_module },
.{ .name = "initial-ramdisk", .module = initial_ramdisk_module },
.{ .name = "build_options", .module = loader_options_module }, .{ .name = "build_options", .module = loader_options_module },
}, },
}), }),
@@ -706,11 +710,11 @@ pub fn build(b: *std.Build) void {
// --- danos-usb.img: the bootable FAT32 USB image --- // --- danos-usb.img: the bootable FAT32 USB image ---
// Format a real FAT32 image (the in-repo Python builder, no external tools) // Format a real FAT32 image (the in-repo Python builder, no external tools)
// holding exactly what the firmware and bootloader need off the ESP: the EFI // holding the EFI stub, the kernel, and the whole /system tree of user
// stub, the kernel, init, and the initial-ramdisk. QEMU presents this image as // binaries at their FHS paths. QEMU presents this image as a USB mass-storage
// a USB mass-storage device the guest boots from (see run-x86-64 and the test // device the guest boots from (see run-x86-64 and the test harness), and the
// harness), and the danos fat driver mounts the same image at /mnt/usb. // danos fat driver mounts the same image at /mnt/usb.
const fat_image = addBootImage(b, exe.getEmittedBin(), efiexe.getEmittedBin(), init_exe.getEmittedBin(), initial_ramdisk_img); const fat_image = addBootImage(b, exe.getEmittedBin(), efiexe.getEmittedBin(), manifest_file, capsule_img, &bundled);
const fat_image_install = b.addInstallFile(fat_image, "danos-usb.img"); const fat_image_install = b.addInstallFile(fat_image, "danos-usb.img");
b.getInstallStep().dependOn(&fat_image_install.step); b.getInstallStep().dependOn(&fat_image_install.step);
@@ -719,7 +723,7 @@ pub fn build(b: *std.Build) void {
// log captured to serial0 — without baking serial into the image users flash. // log captured to serial0 — without baking serial into the image users flash.
// Built lazily (only when `run-x86-64` is requested), and never installed. // Built lazily (only when `run-x86-64` is requested), and never installed.
const exe_serial = addKernel(b, kernel_target, optimize, kernel_modules, test_case, true); const exe_serial = addKernel(b, kernel_target, optimize, kernel_modules, test_case, true);
const fat_image_serial = addBootImage(b, exe_serial.getEmittedBin(), efiexe.getEmittedBin(), init_exe.getEmittedBin(), initial_ramdisk_img); const fat_image_serial = addBootImage(b, exe_serial.getEmittedBin(), efiexe.getEmittedBin(), manifest_file, capsule_img, &bundled);
// `zig build check-fat-image` — validate the produced image is a real FAT32 // `zig build check-fat-image` — validate the produced image is a real FAT32
// with the EFI stub present (the builder's own --verify, no external tools). // with the EFI stub present (the builder's own --verify, no external tools).
@@ -850,6 +854,53 @@ pub fn build(b: *std.Build) void {
const run_efi_step = b.step("run-x86-64", "Boot the x86-64 kernel in QEMU (UEFI/OVMF); serial0 is logged to zig-out/qemu-test/run-x86-64-serial0-<timestamp>.log"); const run_efi_step = b.step("run-x86-64", "Boot the x86-64 kernel in QEMU (UEFI/OVMF); serial0 is logged to zig-out/qemu-test/run-x86-64-serial0-<timestamp>.log");
run_efi_step.dependOn(&run_efi.step); run_efi_step.dependOn(&run_efi.step);
// --- run-x86-64-gpu: the same boot plus a virtio-gpu adapter ---
// The VGA device still supplies the boot (GOP) framebuffer the compositor starts
// on; the virtio-gpu function is discovered by the device-manager stack, its
// driver announces a shared scanout, and the compositor upgrades off the GOP
// floor to fenced, tear-free native presents (docs/display-v2.md).
// This is the interactive twin of the `display-native` test case, and 512M
// matches it (the whole driver stack + the compositor's surfaces at once).
// QEMU shows one head per adapter: pick the virtio-gpu head in the View menu
// to watch the native output.
const run_gpu = b.addSystemCommand(&.{
"qemu-system-x86_64",
"-device",
"qemu-xhci,id=xhci",
"-device",
"usb-mouse,bus=xhci.0",
"-device",
"usb-kbd,bus=xhci.0",
"-machine",
"q35",
"-m",
"512M",
"-drive",
b.fmt("if=pflash,format=raw,readonly=on,file={s}", .{ovmf_code}),
});
run_gpu.addArg("-drive");
run_gpu.addPrefixedFileArg("if=pflash,format=raw,file=", vars_out);
run_gpu.addArg("-drive");
run_gpu.addPrefixedFileArg("if=none,id=bootusb,format=raw,file=", fat_image_serial);
run_gpu.addArgs(&.{
"-device",
"usb-storage,bus=xhci.0,drive=bootusb,removable=on,bootindex=0",
"-net",
"none",
"-vga",
"none",
"-device",
"VGA,edid=on,xres=1280,yres=720",
"-device",
"virtio-gpu-pci",
});
const gpu_serial_log = b.fmt("{s}/run-x86-64-gpu-serial0-{s}.log", .{ log_dir, timestamp(b) });
run_gpu.addArgs(&.{ "-serial", b.fmt("file:{s}", .{gpu_serial_log}) });
run_gpu.step.dependOn(&make_log_dir.step);
const run_gpu_step = b.step("run-x86-64-gpu", "Boot in QEMU with a virtio-gpu adapter: the compositor upgrades to fenced (tear-free) native presents; watch the virtio-gpu head in QEMU's View menu");
run_gpu_step.dependOn(&run_gpu.step);
// const run_cmd = b.addRunArtifact(exe); // const run_cmd = b.addRunArtifact(exe);
// const run_step = b.step("run", "Run the app"); // const run_step = b.step("run", "Run the app");
// run_step.dependOn(&run_cmd.step); // run_step.dependOn(&run_cmd.step);
@@ -867,6 +918,7 @@ pub fn build(b: *std.Build) void {
for ([_][]const u8{ for ([_][]const u8{
"system/boot-handoff.zig", "system/boot-handoff.zig",
"system/abi.zig", "system/abi.zig",
"system/initial-ramdisk.zig", // v2 path-named entries: find/basename/magic
"system/devices/device-abi.zig", "system/devices/device-abi.zig",
"system/devices/pci-class.zig", // class/subclass/prog-IF name decoding "system/devices/pci-class.zig", // class/subclass/prog-IF name decoding
"system/devices/acpi-ids.zig", // _HID name decoding "system/devices/acpi-ids.zig", // _HID name decoding
@@ -879,8 +931,7 @@ pub fn build(b: *std.Build) void {
"system/drivers/usb-hid/hid-report.zig", // HID boot-report keyboard/mouse decode "system/drivers/usb-hid/hid-report.zig", // HID boot-report keyboard/mouse decode
"system/drivers/usb-storage/bulk-only-transport.zig", // CBW/CSW wrapper sizes "system/drivers/usb-storage/bulk-only-transport.zig", // CBW/CSW wrapper sizes
"system/drivers/usb-storage/scsi.zig", // SCSI CDB encodings (big-endian) "system/drivers/usb-storage/scsi.zig", // SCSI CDB encodings (big-endian)
"system/services/vfs/path.zig", // mount-prefix path matching "system/vfs-protocol.zig", // NodeKind / DirectoryEntry sizes + op values
"system/services/vfs/protocol.zig", // NodeKind / DirectoryEntry sizes + op values
"system/services/fat/on-disk.zig", // FAT on-disk struct sizes + type detection "system/services/fat/on-disk.zig", // FAT on-disk struct sizes + type detection
"system/services/fat/engine.zig", // FAT read/write over a RAM-backed image "system/services/fat/engine.zig", // FAT read/write over a RAM-backed image
"system/services/display/compositor.zig", // Rect math + fill/composite/blit-tile "system/services/display/compositor.zig", // Rect math + fill/composite/blit-tile
@@ -913,6 +964,21 @@ pub fn build(b: *std.Build) void {
}); });
test_step.dependOn(&b.addRunArtifact(xkb_tests).step); test_step.dependOn(&b.addRunArtifact(xkb_tests).step);
// The tagged kernel log ring: append/wrap/reclaim/sequence-gap behavior over
// a RAM buffer. Needs the `abi` module (record header layout), so it doesn't
// fit the plain loop above.
const log_ring_tests = b.addTest(.{
.root_module = b.createModule(.{
.root_source_file = b.path("system/kernel/log-ring.zig"),
.target = target,
.optimize = optimize,
.imports = &.{
.{ .name = "abi", .module = abi_module },
},
}),
});
test_step.dependOn(&b.addRunArtifact(log_ring_tests).step);
// runtime.time's Instant/Duration arithmetic. time.zig pulls in system.zig (the // runtime.time's Instant/Duration arithmetic. time.zig pulls in system.zig (the
// syscall wrappers), which needs the `abi` module, so it doesn't fit the plain // syscall wrappers), which needs the `abi` module, so it doesn't fit the plain
// loop above. // loop above.
+2 -2
View File
@@ -134,7 +134,7 @@ Prove the pipeline end-to-end from a separate process.
every IPC message at `MESSAGE_MAXIMUM` = 256 — so `replyWait` rejected the oversized every IPC message at `MESSAGE_MAXIMUM` = 256 — so `replyWait` rejected the oversized
receive buffer with `-E2BIG` and the serve loop had been *spinning* since D2 (unseen, receive buffer with `-E2BIG` and the serve loop had been *spinning* since D2 (unseen,
as D2/D3 matched init-time heartbeats). Set it to 256; `blit_tile` is now explicitly as D2/D3 matched init-time heartbeats). Set it to 256; `blit_tile` is now explicitly
a small-tile path (≤ 54 px inline), larger bitmaps being the deferred shm surface. a small-tile path (≤ 54 px inline), larger bitmaps being the deferred shared-memory surface.
**Gate (met):** `python3 test/qemu_test.py display-demo` spawns the service + `display-demo`; **Gate (met):** `python3 test/qemu_test.py display-demo` spawns the service + `display-demo`;
the demo drives a run of frames of motion through the layer client API and logs the demo drives a run of frames of motion through the layer client API and logs
@@ -172,7 +172,7 @@ documented (docs/display.md): no runtime mode-setting (native backend) and no tr
## Deferred (explicitly not in this plan) ## Deferred (explicitly not in this plan)
- **Shared-memory surfaces** — generalize M13 capability passing to memory objects - **Shared-memory surfaces** — generalize M13 capability passing to memory objects
(`shm_create`/`shm_map`), so bitmap clients hand the compositor a rendered surface (`shared_memory_create`/`shared_memory_map`), so bitmap clients hand the compositor a rendered surface
instead of drawing commands. The compositor's layer model already anticipates it. instead of drawing commands. The compositor's layer model already anticipates it.
- **Native backend (Bochs DISPI, then virtio-gpu)** — behind the same internal backend - **Native backend (Bochs DISPI, then virtio-gpu)** — behind the same internal backend
interface as the dumb framebuffer: EDID mode list + runtime resolution/bpp change + interface as the dumb framebuffer: EDID mode list + runtime resolution/bpp change +
+25 -23
View File
@@ -6,11 +6,11 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run,
## Locked decisions (do not relitigate) ## Locked decisions (do not relitigate)
- **First native backend = virtio-gpu** (VM standard: mode-set + present/flush + vsync). - **First native backend = virtio-gpu** (VM standard: mode-set + fenced present/flush).
- **Dynamic hot-attach**: boot on GOP, upgrade to native when the driver **announces** - **Dynamic hot-attach**: boot on GOP, upgrade to native when the driver **announces**
(push, not polling); re-attach across driver restarts; GOP is the floor for "no driver (push, not polling); re-attach across driver restarts; GOP is the floor for "no driver
ever," not a live fall-back after a reprogram. ever," not a live fall-back after a reprogram.
- **v2 builds the `shm` capability** (endpoints → memory objects), shared with the future - **v2 builds the shared-memory capability** (endpoints → memory objects), shared with the future
client-surface path. client-surface path.
- The compositor's layers/back-buffer/damage are **unchanged**; only scanout is pluggable. - The compositor's layers/back-buffer/damage are **unchanged**; only scanout is pluggable.
@@ -26,7 +26,7 @@ kebab-case file names, no `Co-Authored-By` trailers. New user binaries go throug
**Every gate is serial-checkable — no screenshots** (this plan is built to run unattended). **Every gate is serial-checkable — no screenshots** (this plan is built to run unattended).
Where "does it actually display" would otherwise need a human eyeball, the code **reads its Where "does it actually display" would otherwise need a human eyeball, the code **reads its
own pixels back**: the scanout resource is CPU-visible RAM (shm-backed) and the back buffer own pixels back**: the scanout resource is CPU-visible RAM (shared-memory-backed) and the back buffer
is cacheable, so a driver/compositor can write a known value, read it back, and log a is cacheable, so a driver/compositor can write a known value, read it back, and log a
pass/fail — and a virtio `resource_flush` is confirmed by the device **acking it on the pass/fail — and a virtio `resource_flush` is confirmed by the device **acking it on the
used ring**. Those two together (pixel-readback + flush-ack) are the automated stand-in for used ring**. Those two together (pixel-readback + flush-ack) are the automated stand-in for
@@ -46,7 +46,7 @@ Extract scanout from the compositor so today's path becomes one backend among fu
- [x] `system/services/display/backend.zig`: a `Backend` tagged union with `info()`, - [x] `system/services/display/backend.zig`: a `Backend` tagged union with `info()`,
`surface()` (the cacheable compose target), `present(damage)`, and capability flags `surface()` (the cacheable compose target), `present(damage)`, and capability flags
(`canModeSet`/`hasVsync`, both false for GOP). (`canModeSet`/`hasFencedPresent`, both false for GOP).
- [x] The v1 GOP path is now `backend.Gop` (claims the `display` node, WC-maps the LFB, - [x] The v1 GOP path is now `backend.Gop` (claims the `display` node, WC-maps the LFB,
keeps the cacheable back buffer, `present` = the damage-rect WC copy). display.zig keeps the cacheable back buffer, `present` = the damage-rect WC copy). display.zig
composes into `backend.surface()` and calls `backend.present(damage)` — no LFB or composes into `backend.surface()` and calls `backend.present(damage)` — no LFB or
@@ -57,25 +57,25 @@ Extract scanout from the compositor so today's path becomes one backend among fu
**Gate (met):** `display-service` + `display-demo` pass **unchanged** (pure refactor; GOP **Gate (met):** `display-service` + `display-demo` pass **unchanged** (pure refactor; GOP
is the only backend), and `zig build test` stays green. is the only backend), and `zig build test` stays green.
## V2 — The `shm` cross-process memory capability (kernel) ✅ ## V2 — The shared-memory cross-process capability (kernel) ✅
- [x] [abi.zig](../system/abi.zig): `shm_create` (34) / `shm_map` (35) syscalls + a - [x] [abi.zig](../system/abi.zig): `shared_memory_create` (34) / `shared_memory_map` (35) syscalls + a
`shm_test` service id. Handlers in process.zig: `shm_create(len)` allocates contiguous, `shared_memory_test` service id. Handlers in process.zig: `shared_memory_create(len)` allocates contiguous,
zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability
handle, maps them into the caller's shm arena → returns virtual_address + handle; `shm_map(cap)` handle, maps them into the caller's shared-memory arena → returns virtual_address + handle; `shared_memory_map(cap)`
maps the same physical pages into the receiver. Reclaimed on death (see below). maps the same physical pages into the receiver. Reclaimed on death (see below).
- [x] The capability core (ipc-synchronous.zig) is now **kind-tagged**: `scheduler.Task`'s - [x] The capability core (ipc-synchronous.zig) is now **kind-tagged**: `scheduler.Task`'s
handle table holds `HandleObject{kind, ptr}`; `closeHandles` and `shareCapability` handle table holds `HandleObject{kind, ptr}`; `closeHandles` and `shareCapability`
dispatch by kind, so an `ShmObject` rides an `ipc_call` `send_cap` exactly like an dispatch by kind, so a `SharedMemoryObject` rides an `ipc_call` `send_cap` exactly like an
endpoint and frees only when its last capability drops. `mapUserSharedInto` (paging) endpoint and frees only when its last capability drops. `mapUserSharedInto` (paging)
maps WB-cacheable + `device_grant`, so a sharer's teardown never frees the shared maps WB-cacheable + `device_grant`, so a sharer's teardown never frees the shared
frames — the object owns them. frames — the object owns them.
- [x] `library/runtime/shm.zig` (+ barrel export): `create(len) -> Region{ptr, handle, len}`, - [x] `library/runtime/shared-memory.zig` (+ barrel export): `create(len) -> Region{ptr, handle, len}`,
`map(handle) -> ptr`. `map(handle) -> ptr`.
**Gate (met):** `python3 test/qemu_test.py shm` — `shm-client` creates a region, writes a **Gate (met):** `python3 test/qemu_test.py shared-memory` — `shared-memory-client` creates a region, writes a
pattern, and passes its capability to `shm-server` as an `ipc_call` send_cap; the server pattern, and passes its capability to `shared-memory-server` as an `ipc_call` send_cap; the server
`shm_map`s it and reads the **same bytes** back → `shm: shared 4096 bytes ok`. Guardrail: `shared_memory_map`s it and reads the **same bytes** back → `shared-memory: shared 4096 bytes ok`. Guardrail:
`ipc`/`ipc-call`/`ipc-cap`, `supervision`, `dma`, `usermem`, `display-service`, and host `ipc`/`ipc-call`/`ipc-cap`, `supervision`, `dma`, `usermem`, `display-service`, and host
tests all still pass — the handle-table change broke no existing IPC. tests all still pass — the handle-table change broke no existing IPC.
@@ -88,7 +88,7 @@ tests all still pass — the handle-table change broke no existing IPC.
VERSION_1, and stand up the control virtqueue in coherent DMA. `virtio-gpu-protocol.zig` VERSION_1, and stand up the control virtqueue in coherent DMA. `virtio-gpu-protocol.zig`
+ `virtio-pci.zig` for the control/transport structs (host-tested sizes). + `virtio-pci.zig` for the control/transport structs (host-tested sizes).
- [x] Create a 2D scanout resource backed by a coherent DMA region (V4 swaps this for the - [x] Create a 2D scanout resource backed by a coherent DMA region (V4 swaps this for the
shm-shared surface), `attach_backing`, `set_scanout` to scanout 0, `transfer_to_host_2d` shared-memory surface), `attach_backing`, `set_scanout` to scanout 0, `transfer_to_host_2d`
+ `resource_flush` of a test pattern, and wait on the used ring. + `resource_flush` of a test pattern, and wait on the used ring.
- [x] Register a `scanout` service (`ServiceId.scanout` = 11). - [x] Register a `scanout` service (`ServiceId.scanout` = 11).
@@ -102,7 +102,7 @@ end without a screenshot (the used-ring ack is the device confirming it consumed
## V4 — The native backend + hot-attach ✅ ## V4 — The native backend + hot-attach ✅
- [x] `backend.VirtioGpu` in the compositor: `surface()` = the shared `shm` scanout surface - [x] `backend.VirtioGpu` in the compositor: `surface()` = the shared-memory scanout surface
(the compositor composes straight into the device's resource backing; x86 DMA is (the compositor composes straight into the device's resource backing; x86 DMA is
coherent, so the cacheable shared pages need no flush), `present(damage)` = a `present` coherent, so the cacheable shared pages need no flush), `present(damage)` = a `present`
request over the driver's `.scanout` endpoint (→ transfer-to-host + resource flush). request over the driver's `.scanout` endpoint (→ transfer-to-host + resource flush).
@@ -112,7 +112,7 @@ end without a screenshot (the used-ring ack is the device confirming it consumed
driver registered it), switches backend, and re-composites the current frame full-screen. driver registered it), switches backend, and re-composites the current frame full-screen.
The present is deferred to a one-shot timer so it runs *after* the reply unblocks the The present is deferred to a one-shot timer so it runs *after* the reply unblocks the
driver and it serves `.scanout` — presenting inline would deadlock. driver and it serves `.scanout` — presenting inline would deadlock.
- [x] Boot still starts on `backend.Gop`; the upgrade happens on announce. `shm_physical` (a - [x] Boot still starts on `backend.Gop`; the upgrade happens on announce. `shared_memory_physical` (a
new syscall) gives the driver the guest-physical of the shared surface for `attach_backing`. new syscall) gives the driver the guest-physical of the shared surface for `attach_backing`.
**Gate (met):** the `display-native` case (QEMU `-device virtio-gpu-pci`, `mem` bumped since it **Gate (met):** the `display-native` case (QEMU `-device virtio-gpu-pci`, `mem` bumped since it
@@ -123,7 +123,7 @@ confirm the composited frame landed (`display: native present verified`), while
ok` still fires — checked order-independently. Without `-device virtio-gpu-pci` nothing is ok` still fires — checked order-independently. Without `-device virtio-gpu-pci` nothing is
announced and it stays on GOP: the v1 `display-service`/`display-demo` gates pass unchanged. announced and it stays on GOP: the v1 `display-service`/`display-demo` gates pass unchanged.
## V5 — Mode-setting, EDID, and vsync ✅ ## V5 — Mode-setting, EDID, and fenced presents ✅
- [x] The driver negotiates `VIRTIO_GPU_F_EDID` (when offered) and reads the monitor's EDID, - [x] The driver negotiates `VIRTIO_GPU_F_EDID` (when offered) and reads the monitor's EDID,
logging its preferred mode; it offers a small mode list over `.scanout` `get_modes`. The logging its preferred mode; it offers a small mode list over `.scanout` `get_modes`. The
@@ -131,14 +131,16 @@ announced and it stays on GOP: the v1 `display-service`/`display-demo` gates pas
scanout rectangle (no resource/surface churn) — a runtime resolution change. `runtime.display` scanout rectangle (no resource/surface churn) — a runtime resolution change. `runtime.display`
gains `modes()` / `setMode()` (display-protocol `get_modes`/`set_mode`, forwarded to the backend). gains `modes()` / `setMode()` (display-protocol `get_modes`/`set_mode`, forwarded to the backend).
- [x] Every `resource_flush` is issued fenced (`VIRTIO_GPU_FLAG_FENCE`); the device signals the - [x] Every `resource_flush` is issued fenced (`VIRTIO_GPU_FLAG_FENCE`); the device signals the
fence when the frame is on screen, which the used-ring ack the synchronous present waits on fence when it has consumed the frame, which the used-ring ack the synchronous present waits
already gates — a tear-free present. on already gates — a tear-free present. (Completion feedback, **not vblank**: base
- [x] `backend.VirtioGpu` reports `canModeSet` / `hasVsync` = true. virtio-gpu 2D has no display-refresh event, so nothing paces presents to the monitor —
see the "Fenced is not vsync" note in [display-v2.md](display-v2.md).)
- [x] `backend.VirtioGpu` reports `canModeSet` / `hasFencedPresent` = true.
**Gate (met):** the `display-modeset` case (reusing the display-native boot) upgrades to **Gate (met):** the `display-modeset` case (reusing the display-native boot) upgrades to
virtio-gpu, queries the driver's modes, `setMode`s to a different resolution, and confirms the virtio-gpu, queries the driver's modes, `setMode`s to a different resolution, and confirms the
change by reading the backend's geometry back (`display: mode set to {w}x{h}, verified`); the change by reading the backend's geometry back (`display: mode set to {w}x{h}, verified`); the
fenced present path is exercised and confirmed (`display: vsync present ok`) — both from serial, fenced present path is exercised and confirmed (`display: fenced present ok`) — both from serial,
passing 3/3. The driver also logs the EDID preferred mode (`virtio-gpu: EDID preferred mode …`). passing 3/3. The driver also logs the EDID preferred mode (`virtio-gpu: EDID preferred mode …`).
## V6 — Resilience (restart + re-attach) + tests + docs ✅ ## V6 — Resilience (restart + re-attach) + tests + docs ✅
@@ -157,14 +159,14 @@ kills the virtio-gpu driver once after it hellos; the restart policy respawns it
re-announces, and the compositor logs `display: scanout re-attached` after the initial re-announces, and the compositor logs `display: scanout re-attached` after the initial
`display: scanout upgraded to virtio-gpu`, with no CPU exception / panic (the compositor `display: scanout upgraded to virtio-gpu`, with no CPU exception / panic (the compositor
survives) — passing 3/3. All v1 + v2 cases (host tests, `ipc`/`ipc-call`/`ipc-cap`, survives) — passing 3/3. All v1 + v2 cases (host tests, `ipc`/`ipc-call`/`ipc-cap`,
`supervision`, `shm`, `display-service`, `display-demo`, `virtio-gpu`, `display-native`, `supervision`, `shared-memory`, `display-service`, `display-demo`, `virtio-gpu`, `display-native`,
`display-modeset`) pass; default `zig build` is clean. `display-modeset`) pass; default `zig build` is clean.
--- ---
## Deferred (explicitly not in this plan) ## Deferred (explicitly not in this plan)
- **Client-rendered surfaces** — now unblocked by the `shm` capability (V2): an app renders - **Client-rendered surfaces** — now unblocked by the shared-memory capability (V2): an app renders
its own bitmap and hands the compositor a reference. A natural follow-on. its own bitmap and hands the compositor a reference. A natural follow-on.
- **Bochs DISPI backend** — a simpler second native backend (mode-set only, dumb scanout); - **Bochs DISPI backend** — a simpler second native backend (mode-set only, dumb scanout);
slots behind the same interface if wanted. slots behind the same interface if wanted.
+23 -15
View File
@@ -1,8 +1,8 @@
# The display service v2: a pluggable scanout backend # The display service v2: a pluggable scanout backend
**Status: complete (V1–V6).** The compositor boots on the GOP framebuffer and, when a **Status: complete (V1–V6).** The compositor boots on the GOP framebuffer and, when a
virtio-gpu driver announces itself, hot-attaches a native backend over the shared `shm` virtio-gpu driver announces itself, hot-attaches a native backend over the shared-memory
scanout surface — with runtime mode-setting, EDID, and fenced (vsync) presents, and it scanout surface — with runtime mode-setting, EDID, and fenced presents, and it
re-attaches across driver restarts. All serial-gated (see [display-v2-plan.md](display-v2-plan.md)). re-attaches across driver restarts. All serial-gated (see [display-v2-plan.md](display-v2-plan.md)).
v1 ([display.md](display.md)) is a compositor that owns the **GOP framebuffer** — it v1 ([display.md](display.md)) is a compositor that owns the **GOP framebuffer** — it
@@ -25,10 +25,10 @@ The compositor itself (layers, back buffer, damage) does not change. Only the la
scanout backend (selected at runtime — GOP by default, native when it appears) scanout backend (selected at runtime — GOP by default, native when it appears)
│ │
├─ GopBackend the v1 path: WC copy back→front to the firmware LFB. ├─ GopBackend the v1 path: WC copy back→front to the firmware LFB.
│ Always available. No mode-set, no vsync. THE FLOOR. │ Always available. No mode-set, no present fence. THE FLOOR.
│ │
└─ VirtioGpuBackend talks to a virtio-gpu driver process over a `scanout` └─ VirtioGpuBackend talks to a virtio-gpu driver process over a `scanout`
service: present via a shared resource + flush (real vsync), service: present via a shared resource + fenced flush,
EDID mode list, runtime mode-set. EDID mode list, runtime mode-set.
``` ```
@@ -37,8 +37,8 @@ A **backend** is a small interface the compositor calls:
- `surface()` → the pixels to compose into and their geometry `{ptr, pitch, format, w, h}` - `surface()` → the pixels to compose into and their geometry `{ptr, pitch, format, w, h}`
(the LFB for GOP; a shared scanout resource for virtio-gpu), (the LFB for GOP; a shared scanout resource for virtio-gpu),
- `present(damage: Rect)` → make the damaged region visible (a no-op-ish WC copy for GOP; - `present(damage: Rect)` → make the damaged region visible (a no-op-ish WC copy for GOP;
a virtio flush, optionally vsync-fenced, for the native path), a fenced virtio flush for the native path),
- capability queries — `canModeSet`, `hasVsync` — and, when supported, `modes()` / - capability queries — `canModeSet`, `hasFencedPresent` — and, when supported, `modes()` /
`setMode(m)`. `setMode(m)`.
The compositor composes into `surface()` and calls `present(damage)` exactly as it does The compositor composes into `surface()` and calls `present(damage)` exactly as it does
@@ -77,9 +77,9 @@ deferred (docs/display.md, "What v1 does not do"). v2 builds it: the natural gen
of M13 capability-passing from *endpoints* to *memory objects* — of M13 capability-passing from *endpoints* to *memory objects* —
``` ```
shm_create(len) -> {handle, virtual_address} // a shareable, page-aligned RAM region shared_memory_create(len) -> {handle, virtual_address} // a shareable, page-aligned RAM region
… pass `handle` as the send_cap on an ipc_call … … pass `handle` as the send_cap on an ipc_call …
shm_map(cap) -> virtual_address // the receiver maps the same physical pages shared_memory_map(cap) -> virtual_address // the receiver maps the same physical pages
``` ```
The payoff is leverage: the **same** primitive unlocks **both** native GPU drivers *and* The payoff is leverage: the **same** primitive unlocks **both** native GPU drivers *and*
@@ -92,35 +92,43 @@ A new ring-3 driver process (the topology v1 anticipated — "split the driver f
compositor when a second backend arrives"). It claims the virtio-gpu PCI function, and: compositor when a second backend arrives"). It claims the virtio-gpu PCI function, and:
- sets up the **virtqueues** (control + cursor) and the device's config space, - sets up the **virtqueues** (control + cursor) and the device's config space,
- creates a **2D scanout resource** backed by an `shm` region, `attach_backing`s it, - creates a **2D scanout resource** backed by a shared-memory region, `attach_backing`s it,
`set_scanout`s it to a CRTC, and `resource_flush`es damaged rectangles, `set_scanout`s it to a CRTC, and `resource_flush`es damaged rectangles,
- reads **EDID** (the `GET_EDID` control command) for the mode list, and `set_scanout` - reads **EDID** (the `GET_EDID` control command) for the mode list, and `set_scanout`
at a chosen mode for **runtime mode-setting**, at a chosen mode for **runtime mode-setting**,
- registers a `scanout` service and announces to the display service. - registers a `scanout` service and announces to the display service.
Its `resource_flush` is the real **present** — and gives a genuine **vsync/tear-free** Its `resource_flush` is the real **present** — and gives a **fenced, tear-free** path a
path a dumb GOP framebuffer can't. dumb GOP framebuffer can't.
**Fenced is not vsync.** The fence completes when the device has *consumed* the frame:
real completion feedback, and tear-freedom by snapshot semantics (the host displays
discrete transferred frames, never a half-written surface). It is **not** a vblank —
base virtio-gpu 2D has no display-refresh event at all (Linux's driver for this device
fakes one with a software timer), so nothing paces presents to the monitor's refresh.
Refresh-paced presents need either a native driver's vblank interrupt (delivered over
the existing IRQ-as-IPC path) or the compositor's own frame clock.
## What v2 unlocks — and its honest scope ## What v2 unlocks — and its honest scope
Behind the abstraction, a native backend gives runtime **mode-setting** (resolution / Behind the abstraction, a native backend gives runtime **mode-setting** (resolution /
refresh / bpp), **EDID** enumeration, and **vsync**. But only on devices we have a driver refresh / bpp), **EDID** enumeration, and **fenced presents**. But only on devices we have a driver
for — realistically **VMs** (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs for — realistically **VMs** (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs
need per-vendor KMS-class drivers that aren't getting written, so they **stay on GOP** — need per-vendor KMS-class drivers that aren't getting written, so they **stay on GOP** —
which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold: which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold:
the **pluggable architecture** (a driver slots in when one exists) and a **rich, vsync'd the **pluggable architecture** (a driver slots in when one exists) and a **rich, fenced
path in VMs**, where danos development happens. The framebuffer floor never goes away. path in VMs**, where danos development happens. The framebuffer floor never goes away.
## Locked decisions ## Locked decisions
- **First native backend: virtio-gpu** — the VM standard; gives mode-set + a real - **First native backend: virtio-gpu** — the VM standard; gives mode-set + a real
present/flush (and vsync), and exercises the whole pluggable design. Tested with QEMU present/flush (fenced), and exercises the whole pluggable design. Tested with QEMU
`-device virtio-gpu`. `-device virtio-gpu`.
- **Dynamic hot-attach** — boot on GOP, upgrade to native on the driver's announce, - **Dynamic hot-attach** — boot on GOP, upgrade to native on the driver's announce,
re-attach across driver restarts; GOP is the floor for "no driver ever," not a live re-attach across driver restarts; GOP is the floor for "no driver ever," not a live
fall-back after a reprogram. fall-back after a reprogram.
- **Detection = push** (the driver announces to `.display`), not compositor polling. - **Detection = push** (the driver announces to `.display`), not compositor polling.
- **v2 builds the `shm` capability** (endpoints → memory objects), shared with the future - **v2 builds the shared-memory capability** (endpoints → memory objects), shared with the future
client-surface path. client-surface path.
## See also ## See also
+12 -4
View File
@@ -95,7 +95,7 @@ rest of the system hasn't had to face:
┌────────────────────────────────────┬──────────────────────────────────────┐ ┌────────────────────────────────────┬──────────────────────────────────────┐
drawing clients (v1) surface clients (deferred) drawing clients (v1) surface clients (deferred)
runtime.display commands: runtime.display surfaces: runtime.display commands: runtime.display surfaces:
create_layer / configure_layer shm_create → pass as a capability → create_layer / configure_layer shared_memory_create → pass as a capability →
fill_rect / blit_tile / damage the compositor maps & composites the fill_rect / blit_tile / damage the compositor maps & composites the
present client-rendered bitmap directly present client-rendered bitmap directly
``` ```
@@ -196,13 +196,21 @@ shell, a terminal, a cursor, and a wallpaper:
| `fill_rect` | fill a rectangle of a layer with a colour | | `fill_rect` | fill a rectangle of a layer with a colour |
| `blit_tile` | copy a small client-supplied pixel tile into a layer (inline) | | `blit_tile` | copy a small client-supplied pixel tile into a layer (inline) |
| `damage` | mark a region of a layer dirty | | `damage` | mark a region of a layer dirty |
| `present` | composite dirty layers and flush to the screen | | `present` | request a repaint: composited at the next frame-clock tick |
Text is intentionally *not* an operation — a client renders glyphs by blitting tiles Text is intentionally *not* an operation — a client renders glyphs by blitting tiles
(the [PSF font](../system/kernel/font.psf) path the console already uses can move into a (the [PSF font](../system/kernel/font.psf) path the console already uses can move into a
client). Keeping the protocol to rectangles and tiles keeps the compositor small and the client). Keeping the protocol to rectangles and tiles keeps the compositor small and the
policy in the client. policy in the client.
`present` is a *request*, not an immediate flush: the compositor runs a ~60 Hz **frame
clock** (a one-shot kernel timer re-armed on demand), and each tick composites all the
damage accumulated since the last one. Any number of client presents and cursor moves
inside one interval coalesce into a single repaint — the software stand-in for vblank
pacing on backends that have none (all of them today; see
[display-v2.md](display-v2.md), "Fenced is not vsync"). Bring-up paths that must put
pixels on screen synchronously (initialisation, the self-checks) bypass the clock.
## `runtime.display` ## `runtime.display`
Clients speak the protocol through a new [`library/runtime/display.zig`](../library/runtime/runtime.zig), Clients speak the protocol through a new [`library/runtime/display.zig`](../library/runtime/runtime.zig),
@@ -252,8 +260,8 @@ both are clean additions behind the interfaces v1 establishes.
to render into its *own* buffer and hand the compositor a *reference*, not a stream of to render into its *own* buffer and hand the compositor a *reference*, not a stream of
commands. That needs the missing cross-process shared-memory primitive — best built as commands. That needs the missing cross-process shared-memory primitive — best built as
the natural generalization of the existing M13 [capability passing](driver-model.md) the natural generalization of the existing M13 [capability passing](driver-model.md)
from *endpoints* to *memory objects* (`shm_create(len) → {cap, virtual_address}`, pass `cap` on from *endpoints* to *memory objects* (`shared_memory_create(len) → {cap, virtual_address}`, pass `cap` on
an `ipc_call`, receiver `shm_map(cap) → virtual_address`). v1 avoids it because server-owned an `ipc_call`, receiver `shared_memory_map(cap) → virtual_address`). v1 avoids it because server-owned
surfaces already prove the whole pipeline. surfaces already prove the whole pipeline.
- **Runtime mode-setting (a native backend).** Detecting the EDID mode list and changing - **Runtime mode-setting (a native backend).** Detecting the EDID mode list and changing
+78 -86
View File
@@ -1,104 +1,96 @@
# Logging: the diagnostic log vs. the display # Logging
danos separates two things that are easy to conflate: the **diagnostic log** — the Output is a *diagnostic convenience, never a correctness dependency*: the kernel
machine-readable stream of *what the kernel is doing* — and the **display**, the and every service must run correctly with zero output channels. On top of that
framebuffer surface the OS draws on. They are different concerns with different rule, danos has **per-process logging** — every process's output is attributed
lifetimes, so they're different code paths. by the kernel and lands in its own file on the flash volume, which is what makes
a headless real machine (no serial port) debuggable. The display (the
framebuffer surface) is a separate concern and deliberately not a log sink;
`main.zig` mirrors a few user-facing status lines and panics to it explicitly.
The guiding rule: **output is a diagnostic convenience, never a correctness ## The pipeline
dependency.** The kernel must boot and run correctly with *zero* output channels —
no serial, no screen. Logging that can take the kernel down isn't robust; it's a
liability. This is the same [resilience](resilience.md) posture the rest of the
kernel follows.
## The log is multi-sink
`system/kernel/log.zig` is the diagnostic log. It fans a message out to a set of
registered **sinks**, each best-effort and self-guarding:
```zig
log.addSink(arch.serialWrite); // the serial UART
if (arch.debugconPresent()) log.addSink(arch.debugconWrite); // 0xE9 debug console
// later: log.addSink(fs.logWrite); // a file on a ramdisk / USB / SSD
log.write("…"); log.print("x={d}\n", .{x});
```
Properties that matter:
- **No allocation.** The sink table is a fixed array, so the log works before the
heap is up and inside a panic.
- **Best-effort.** A sink whose device is absent is a no-op (e.g. writing to a
missing UART just goes nowhere — the TX-wait is bounded so it can't hang). A
message reaches whatever channels exist; if none do, the kernel runs on, silent.
- **Order-independent.** Every registered sink gets every message. Adding the file
logger later is one `addSink` call and **zero** changes to call sites.
## The framebuffer is *not* a log sink
The framebuffer is a general graphics surface, **not inherently a text terminal**.
Today `system/kernel/console.zig` paints a text grid on it as a *bootstrap* console, but
that's a stop-gap: once the driver machinery exists the framebuffer becomes a proper
**graphics device driver**, and the text crutch goes away. So the log must not assume
it — routing the verbose log through a pixel console would bake in "the OS is text".
Instead the two paths are explicit:
``` ```
verbose diagnostics ──► log ──► serial, debugcon, (file later) process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /var/log/<boot-stamp>/<binary-path>.log
user status / panics ──► status() ──► log (above) + framebuffer (if present) kernel log.print ─┘ │
└▶ serial / 0xE9 sinks (QEMU, -Dserial)
``` ```
A handful of user-facing lines (`kernel initialised`, a panic) go through 1. **Emit.** A program calls `std.log.info("mounted {s}", .{path})` — the
`main.zig`'s `status()` / `statusPrint()`, which write to the log **and** paint the runtime's `logFn` (installed for every binary by the root shim,
framebuffer when one is present. Everything else uses `log.*` and never touches the `library/runtime/log.zig`) formats one line and issues one `debug_write`
screen. `console.write` is a no-op when the firmware gave us no framebuffer. carrying the level. The payload does NOT contain the process's name.
`runtime.system.write` remains as the raw/bring-up path (panics, test
fixtures); raw bytes ride the same ring, attributed all the same.
## Optional framebuffer (headless machines) 2. **Stamp.** The kernel wraps every payload LINE in a record stamped with the
sender's pid, task name (its binary path, e.g. `/system/services/fat`),
level, a per-boot sequence number, and a monotonic timestamp
(`system/kernel/log.zig` + `log-ring.zig`). Attribution is structural — a
payload cannot forge another sender's tag, and an embedded newline just ends
the record, so the forged "prefix" lands inside the forger's own next line.
A framebuffer is not guaranteed — a headless server exposes no UEFI Graphics Output 3. **Retain.** The 512 KiB ring overwrites oldest-first; sequence gaps make any
Protocol. That used to be *fatal* (the loader failed the boot). Now the loader hands loss countable. `klog_read` (#32) copies stream bytes from a free-running
over a "no framebuffer" descriptor (`base == 0`) rather than failing, and offset; `klog_status` (#45) returns the cursors plus the wall-clock time of
`Framebuffer.present()` (in `system/boot-handoff.zig`) gates every on-screen path. A headless, boot. The framing (`abi.KlogRecordHeader`) is 32 bytes + name + payload,
serial-less machine boots and runs correctly — it just goes quiet. 8-byte aligned.
## Last-resort channels (no text output at all) 4. **Render.** Registered sinks (serial under `-Dserial`, the 0xE9 debug
console) get a live transcript: kernel/raw output verbatim, leveled records
as `<binary path>: message` — one composed write per line, under the log's
own spinlock (never the big kernel lock; panic paths try-acquire with a
bound and fall back to sinks-only). Sinks are best-effort and self-guarding;
a serial-less machine just goes quiet.
Two signals bypass the sink list, because they must survive even a total 5. **Persist.** The **logger service** (`system/services/logger`) drains the
output-channel failure: ring every 250 ms and demultiplexes records into one file per source under
`/var/log/<boot-stamp>/`, e.g.
- **`log.checkpoint(code)`** — a one-byte **POST code** to I/O port `0x80` (a POST ```
card or BMC shows it). `main.zig` emits one at each boot milestone (`cp_paging`, /var/log/2026-07-21T150434Z/kernel.log
`cp_heap`, …) and on a fault/panic, so "where did it hang?" is answerable with no /var/log/2026-07-21T150434Z/system/services/fat.log
text output whatsoever. Writing `0x80` is universally safe. /var/log/2026-07-21T150434Z/system/drivers/usb-storage.log
- **`log.recordPanic(msg)`** — stamps the panic message into a fixed record ```
(`log.panic_record`, with a `magic` written last). A post-mortem — an attached
debugger, a RAM dump, or a future file/pstore reader — recovers *what killed it*
even though nothing was on screen.
The panic and CPU-exception handlers fan out to every sink, emit a POST code, and The boot stamp is the RTC anchor from `klog_status` (FAT-safe: no colons; a
drop the breadcrumb — they never assume a console. dead RTC yields the 1970 directory rather than no logs). Each line carries
the record's monotonic timestamp and level. Storage is best-effort and late:
the ring buffers a whole boot many times over, and the first successful
`makePath` of the per-boot directory (also the readiness probe) triggers a
full backlog write. Files close — which is the fat server's SCSI cache
flush — after a ~2 s quiet period, bounding data-at-risk without per-record
flush thrash. At shutdown init stops the logger FIRST (it is last in the
boot order), so its final drain runs over a live storage chain.
## The 0xE9 debug console ## Why a ring in the kernel, not a logging server
Port `0xE9` is the Bochs/QEMU debug console. It's detected safely: the port returns The storage stack must be able to log. If the fat server wrote its own log file
`0xE9` when read if present, and `0xFF` on real hardware, so `debugconPresent()` through the VFS it would rendezvous-deadlock on itself; if processes sent
only enables the sink when it's really there. Under QEMU it's captured with records to a logging server over IPC, early boot would need a buffer that is —
`-debugcon file:…`, giving CI a log channel independent of `-serial`. a ring, one hop later. The kernel ring is that buffer, placed where every
process (and the kernel itself) can reach it with one syscall, before any
service exists. The logger service is a *reader*, not a hop.
## The robustness spectrum Two disciplines keep it honest:
The result handles every combination — framebuffer only, serial only, both, or - the logger announces itself **once** — a periodic status line would feed the
**neither**. With no channels at all the kernel still boots and runs; port-`0x80` very stream it drains;
checkpoints track progress and the panic breadcrumb captures failures. *Runs blind - lost records surface as an explicit `-- N records lost --` line, computed
but correct* is the goal, not *always has output*. from sequence gaps, never silently.
## Related ## Last-resort channels
- [framebuffer.md](framebuffer.md) — the display surface itself (pitch, format), the Unchanged, and independent of the sink list so they survive a total output
thing that becomes a graphics device driver. failure: `checkpoint` (a one-byte POST code on port 0x80) and `recordPanic`
- [efi.md](efi.md) — where the loader captures (or, headless, doesn't capture) the (a fixed breadcrumb record, `log.panic_record`, findable in a RAM dump; magic
framebuffer before `ExitBootServices`. written last so a reader only trusts a complete record).
- [device-interrupts.md](device-interrupts.md) — the serial UART bring-up the log's
primary sink rides on. ## Accepted gaps
- [resilience.md](resilience.md) — why "never let a missing peripheral take the
kernel down" is a core design stance. - A write-spamming process can evict other processes' unread records from the
ring (a per-process quota is future work); the loss is at least visible via
sequence gaps in every affected file.
- `/var/log` files have no privacy until the VFS grows permissions.
- Records emitted after the logger's final shutdown drain reach serial and the
ring but not the files.
+10 -7
View File
@@ -1,15 +1,18 @@
# System Calls # System Calls
System calls (syscalls) are the bridge between your programs and the operating system's restricted core (kernel). System calls (syscalls) are the bridge between your programs and the operating system's restricted core (kernel).
> **Status:** danos has real user processes (M3). User programs enter the kernel > **Status:** danos has real user processes. User programs enter the kernel
> via the `syscall` instruction (STAR/LSTAR/SFMASK set per core; the entry stub in > via the `syscall` instruction (STAR/LSTAR/SFMASK set per core; the entry stub in
> `isr.s` does the `swapgs` + kernel-stack switch and reuses the interrupt > `isr.s` does the `swapgs` + kernel-stack switch and reuses the interrupt
> dispatcher). The `int 0x80` gate is kept alongside as a minimal test path. The > dispatcher); the `int 0x80` gate is kept alongside as a minimal test path.
> current call set is still a placeholder — `0 = exit(code)`, `1 = ping`, > The live table is `system/abi.zig` (private, renumberable — see
> `2 = write(ptr, len)`, `3 = sleep(ms)` (see `system/kernel/process.zig`); the > [vdso.md](vdso.md) for the public boundary): process lifecycle + threads,
> handler dispatches on whether the caller is a scheduled process (its own address > memory (mmap/dma/shared-memory), synchronous + async IPC with capability passing,
> space) or a borrowed test thread. The microkernel set below (IPC_Call / > device access, time, the tagged-log diagnostics (`debug_write` with a level,
> IPC_ReplyWait / Yield) replaces it once a second user server exists. > `klog_read`/`klog_status`), and filesystem NAMING (`fs_resolve`/`fs_node`/
> `fs_mount`/`fs_unmount` — the kernel VFS root routes paths and serves the
> read-only /system initrd mount; file DATA stays with userspace filesystem
> servers over the vfs-protocol, docs/vfs-protocol.md).
## The Mechanism of a Syscall ## The Mechanism of a Syscall
+2 -2
View File
@@ -14,7 +14,7 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run,
- **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an - **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an
idle core still halts ([halting.md](halting.md)). idle core still halts ([halting.md](halting.md)).
- **New syscalls are private**: extend [abi.zig](../system/abi.zig) `SystemCall` after - **New syscalls are private**: extend [abi.zig](../system/abi.zig) `SystemCall` after
`shm_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`, `shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`,
`futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number. `futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number.
- **Restart granularity stays the process** — a faulting thread kills its process; the - **Restart granularity stays the process** — a faulting thread kills its process; the
supervisor restarts the process, which respawns its threads. supervisor restarts the process, which respawns its threads.
@@ -457,7 +457,7 @@ clean.
## Deferred (explicitly not in this plan) ## Deferred (explicitly not in this plan)
- **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a - **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a
physical-address key so two processes share a futex through an [shm](display-v2.md) physical-address key so two processes share a futex through a [shared-memory](display-v2.md)
region. Not needed for intra-process threads. region. Not needed for intra-process threads.
- **Per-thread priorities / affinity distinct from the process** — threads inherit the - **Per-thread priorities / affinity distinct from the process** — threads inherit the
process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep. process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep.
+2 -2
View File
@@ -133,7 +133,7 @@ Deviations from `std.Thread`, called out honestly:
## Kernel primitives (new private syscalls) ## Kernel primitives (new private syscalls)
Four new entries extend [abi.zig](../system/abi.zig) `SystemCall` after Four new entries extend [abi.zig](../system/abi.zig) `SystemCall` after
`shm_physical = 36`, each with a `library/runtime` wrapper: `shared_memory_physical = 36`, each with a `library/runtime` wrapper:
| Syscall | Signature | Purpose | | Syscall | Signature | Purpose |
|---|---|---| |---|---|---|
@@ -210,7 +210,7 @@ Keying: threads share an address space, so a **virtual address within that addre
identifies a futex uniquely; the kernel keys its wait queue by `(address_space_root, virtual_address)`. identifies a futex uniquely; the kernel keys its wait queue by `(address_space_root, virtual_address)`.
Keying by the **physical** address instead (translate `virtual_address -> physical_address` on entry) is a Keying by the **physical** address instead (translate `virtual_address -> physical_address` on entry) is a
deliberate forward door: it lets two *processes* share a futex through an deliberate forward door: it lets two *processes* share a futex through an
[shm](display-v2.md) region later, without changing the API. We start with the [shared-memory](display-v2.md) region later, without changing the API. We start with the
private-per-address-space key and note the physical-key upgrade. private-per-address-space key and note the physical-key upgrade.
No spinning: a contended lock parks the task in the kernel and the core is free to run No spinning: a contended lock parks the task in the kernel and the core is free to run
+111
View File
@@ -0,0 +1,111 @@
# USB hubs (M22)
A hub is USB **bus infrastructure**, not an application peripheral, so hub
topology is handled **inside the `usb-xhci-bus` driver** — the process that owns
the controller's device slots and contexts. A device behind a hub is not reached
by any hub-specific software path: it is reached by the **controller**,
programmed with a *route string* in its slot context. Route strings and slot
contexts are xHCI hardware concepts that only exist inside the controller driver,
so that is where hub handling belongs. Class drivers (HID, storage) stay separate
and unaware — the hub is transparent to them; a keyboard behind a hub reaches the
same `usb-hid-keyboard` driver as one on a root port.
This is a deliberate scoping choice, not a microkernel compromise: the USB *bus*
driver handles USB *bus* topology. The alternative — a separate `usb-hub`
class-driver process plus a cross-process enumeration protocol — would only
shuttle the bus's own topology state (slot ids, route strings, TT linkage) out to
another process and back, since the hub driver cannot build a slot context
itself.
## The compound-hub reality
A USB 3.0 hub is physically **two hubs** sharing each connector: a SuperSpeed hub
and a USB 2.0 companion hub, enumerated as **separate devices on separate root
ports**. A full- or low-speed device plugged into a USB 3.0 hub attaches to the
**USB 2.0 companion**, not the SuperSpeed hub. So supporting full-speed devices
(keyboards, mice) behind a hub means driving the USB 2.0 companion and handling
**transaction translators** — there is no SuperSpeed-only shortcut that reaches a
full-speed keyboard.
## Slot-context fields for a downstream device
`buildAddressInputContext` fills the Slot Context. Today it hardcodes route 0 and
the root-hub port. A downstream device additionally needs:
- **Route String** (Slot Context dword 0, bits 19:0) — 5 tiers × 4 bits, each
tier the downstream hub-port number. Composed as
`route = (parent_route << 4) | hub_port`, capped at the xHCI 5-tier max.
- **Root Hub Port Number** (dword 1, bits 23:16) — the *root* port the whole hub
chain hangs off, inherited from the parent hub (not the hub's own port number).
- **Speed** (dword 0, bits 23:20) — read from the hub's downstream port status
after reset, not assumed.
- **Parent Hub Slot ID** (dword 2, bits 7:0) + **Parent Port Number** (dword 2,
bits 13:8) — the **transaction translator**: set when a full/low-speed device
sits behind a high-speed hub, so the controller routes split transactions
through that hub's TT. For a multi-TT hub, **MTT** (Slot Context dword 0 bit
25) is set and the TT port is the device's own hub port.
## Detection: the status-change interrupt endpoint
A hub has one interrupt IN endpoint that returns a **port-status-change bitmap**
(bit N set = port N changed). The bus arms an interrupt transfer on it (reusing
the controller's existing interrupt-endpoint machinery, but serviced
**in-process** — no class-driver subscription IPC), and on each report:
1. For each changed port, `GET_STATUS` (hub class request) reads the port's
connect/enable/reset state and speed, and `CLEAR_FEATURE(C_PORT_*)`
acknowledges the change.
2. On a **connect**: `SET_FEATURE(PORT_RESET)`, wait for reset-complete via a
later status-change report, read the enabled speed, then `setupDevice` with
the composed route string / root port / TT fields, `enumerate`, and register
the interfaces — exactly the existing path, recursing if the new device is
itself a hub.
3. On a **disconnect**: tear down the downstream device (report each interface
`ChildRemoved`, Disable Slot) — the B3 teardown path, keyed by the device's
route rather than a root port.
## Hub setup (once, when the hub enumerates)
When the bus scan (or a hot-plug bring-up) finds a device of class 9:
1. Read the **hub descriptor** (class GET_DESCRIPTOR, type 0x2A for a USB 3.0
hub / 0x29 for USB 2.0) → downstream port count, characteristics.
2. For a USB 3.0 hub, `SET_FEATURE(BH_PORT_RESET)` semantics and the depth
(`SET_HUB_DEPTH`) so the hub knows its tier for route-string forwarding.
3. `SET_FEATURE(PORT_POWER)` each downstream port.
4. Configure the hub's slot as a hub: **Hub** bit (Slot Context dword 0 bit 26),
**Number of Ports** (dword 1, bits 31:24), **TT Think Time** and **MTT** for a
USB 2.0 multi-TT hub — via an Evaluate/Configure Endpoint on the hub's slot.
5. Arm the status-change interrupt endpoint.
## Testing
QEMU's `usb-hub` is a USB 2.0 single-TT hub. A **static boot topology**
(`-device usb-hub,bus=xhci.0,id=h -device usb-kbd,bus=h.0`) presents the
downstream device connected from the start, so the bus reads it on the first
status-change report — exercising the full path (hub setup, TT slot context,
downstream enumerate, class-driver bind) without needing a hot-plug event. A new
`usb-hub` QEMU case asserts the hub enumerates, the downstream keyboard
enumerates behind it, and `usb-hid-keyboard` binds.
Real-hardware validation (the user's SuperSpeed Genesys hub + full-speed
keyboard/mouse on its USB 2.0 companion) is flagged separately — the compound
USB 3.0 hub path is not modelled by QEMU's USB 2.0 hub.
## Milestones (all complete)
- **B4a** ✓ — hub recognition + setup: detect class 9 in the scan, read the hub
descriptor, configure the slot as a hub, power downstream ports, log the
topology.
- **B4b** ✓ — downstream enumeration: the in-process status-change subscription,
port reset, Address Device with route string + root port + TT fields,
enumerate + register. A full-speed keyboard behind a USB2 hub binds
`usb-hid-keyboard` in QEMU.
- **B4c** ✓ — disconnect teardown (recursive: a hub takes its subtree with it)
and hub-behind-hub recursion (route strings compose across tiers). QEMU's hub
*does* raise downstream status changes, so both connect and disconnect are
harness-tested (`usb-hub`, `usb-hub-nested`, `usb-hub-unplug`).
Real-hardware validation of the user's SuperSpeed Genesys hub with full-speed
devices on its USB 2.0 companion remains pending — QEMU's USB 2.0 hub does not
model the compound USB 3.0 hub.
+8 -5
View File
@@ -119,7 +119,7 @@ One table entry per kernel call, C ABI (System V AMD64), names prefixed
returns are `u64`, errors return as negative values exactly as today. returns are `u64`, errors return as negative values exactly as today.
The calls that return two values in `rax:rdx` today — `dma_alloc` The calls that return two values in `rax:rdx` today — `dma_alloc`
(virtual_address + physical_address), `msi_bind` (address + data), `shm_create` (virtual_address + handle) — (virtual_address + physical_address), `msi_bind` (address + data), `shared_memory_create` (virtual_address + handle) —
become functions returning a two-`u64` struct. The System V ABI returns a become functions returning a two-`u64` struct. The System V ABI returns a
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the 16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
C-ABI spelling of the existing convention, at zero cost. C-ABI spelling of the existing convention, at zero cost.
@@ -129,11 +129,12 @@ Grouped as `abi.zig` groups them:
| Group | Functions | | Group | Functions |
|-------|-----------| |-------|-----------|
| process | `danos_exit`, `danos_yield`, `danos_sleep`, `danos_spawn`, `danos_process_enumerate`, `danos_process_kill`, `danos_process_exit_reason`, `danos_process_subscribe`, `danos_process_signal`, `danos_signal_bind` | | process | `danos_exit`, `danos_yield`, `danos_sleep`, `danos_spawn`, `danos_process_enumerate`, `danos_process_kill`, `danos_process_exit_reason`, `danos_process_subscribe`, `danos_process_signal`, `danos_signal_bind` |
| memory | `danos_mmap`, `danos_munmap`, `danos_dma_alloc`, `danos_dma_free`, `danos_shm_create`, `danos_shm_map`, `danos_shm_physical` | | memory | `danos_mmap`, `danos_munmap`, `danos_dma_alloc`, `danos_dma_free`, `danos_shared_memory_create`, `danos_shared_memory_map`, `danos_shared_memory_physical` |
| ipc | `danos_endpoint_create`, `danos_ipc_register`, `danos_ipc_lookup`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` | | ipc | `danos_endpoint_create`, `danos_ipc_register`, `danos_ipc_lookup`, `danos_ipc_call`, `danos_ipc_reply_wait`, `danos_ipc_send` |
| devices | `danos_device_enumerate`, `danos_device_claim`, `danos_device_register`, `danos_mmio_map`, `danos_irq_bind`, `danos_irq_ack`, `danos_msi_bind`, `danos_io_read`, `danos_io_write` | | devices | `danos_device_enumerate`, `danos_device_claim`, `danos_device_register`, `danos_mmio_map`, `danos_irq_bind`, `danos_irq_ack`, `danos_msi_bind`, `danos_io_read`, `danos_io_write` |
| time | `danos_clock`, `danos_wall_clock`, `danos_timer_bind` | | time | `danos_clock`, `danos_wall_clock`, `danos_timer_bind` |
| diagnostics | `danos_debug_write`, `danos_klog_read` | | diagnostics | `danos_debug_write` (leveled, kernel-stamped records), `danos_klog_read`, `danos_klog_status` |
| filesystem naming | `danos_fs_resolve`, `danos_fs_node`, `danos_fs_mount`, `danos_fs_unmount` (naming only — file DATA still crosses the vfs-protocol IPC, see below) |
The constants that ride alongside the calls — mmap protection bits, DMA The constants that ride alongside the calls — mmap protection bits, DMA
flags, notification badge bits, `ExitReason`, `Signal`, well-known service flags, notification badge bits, `ExitReason`, `Signal`, well-known service
@@ -197,8 +198,10 @@ Phased so every step ships alone (the M-milestone discipline):
across processes stays what it is today: a service behind IPC, or source across processes stays what it is today: a service behind IPC, or source
compiled into each binary. compiled into each binary.
- **No file/device I/O in the vDSO.** The microkernel line doesn't move: the - **No file/device I/O in the vDSO.** The microkernel line doesn't move: the
vDSO wraps the same deliberately tiny table (docs/syscall.md); files are vDSO wraps the same deliberately tiny table (docs/syscall.md). The kernel
still the VFS server's business over IPC. resolves file NAMES (`fs_resolve` — the mount table moved in-kernel), but
file data is still the filesystem server's business over the vfs-protocol
IPC; the kernel never blocks on a userspace filesystem.
- **No fast-path user-mode implementations yet.** Linux's vDSO exists mostly - **No fast-path user-mode implementations yet.** Linux's vDSO exists mostly
to answer `gettimeofday` without a kernel entry. `danos_clock` could one to answer `gettimeofday` without a kernel entry. `danos_clock` could one
day read the calibrated TSC in user mode the same way — the blob is where day read the calibrated TSC in user mode the same way — the blob is where
+26 -18
View File
@@ -1,20 +1,24 @@
# The VFS wire protocol # The VFS wire protocol
> **Status:** built and spoken today between `runtime.fs` (the client) and the > **Status:** built and spoken today between `runtime.fs` (the client) and the
> VFS server (`system/services/vfs`), with mounted backends (the FAT server) > filesystem BACKENDS (the FAT server). The mount router lives in the
> speaking the same protocol behind the router. The Zig source of truth is > **kernel** (`system/kernel/vfs.zig`): `fs_resolve` routes a path and either
> `system/services/vfs/protocol.zig` (the `vfs-protocol` module), whose unit > serves it directly (the read-only /system initrd mount, via `fs_node`) or
> tests pin the sizes and values below. This page is the **language-neutral > redirects the caller to the owning backend's endpoint plus the rewritten
> wire specification** of that contract — what a Rust or C client implements > mount-relative path — after which the client speaks THIS protocol to the
> ([vdso.md](vdso.md) explains why the IPC protocols, not the syscall > backend, unchanged. The Zig source of truth is `system/vfs-protocol.zig`
> numbers, are danos's public ABI). > (the `vfs-protocol` module), whose unit tests pin the sizes and values
> below. This page is the **language-neutral wire specification** of that
> contract — what a Rust or C client implements ([vdso.md](vdso.md) explains
> why the IPC protocols, not the syscall numbers, are danos's public ABI).
## Transport ## Transport
A VFS exchange is one synchronous IPC rendezvous (`ipc_call`, A VFS exchange is one synchronous IPC rendezvous (`ipc_call`,
docs/ipc.md): the client sends one message and blocks; the server replies docs/ipc.md): the client sends one message and blocks; the server replies
with one message. The endpoint is found by well-known service id with one message. The endpoint comes from the kernel's `fs_resolve` — which
(`ipc_lookup`, service id **1** = vfs). also hands back the path rewritten relative to the mount — not from a
registry lookup. (Service id 1, the old userspace router, is retired.)
- A message is at most **256 bytes** (`message_maximum`). - A message is at most **256 bytes** (`message_maximum`).
- A request is a fixed 32-byte **Request** header followed by an inline - A request is a fixed 32-byte **Request** header followed by an inline
@@ -26,8 +30,11 @@ with one message. The endpoint is found by well-known service id
- All integers are **little-endian**; layouts are C layout for x86-64 - All integers are **little-endian**; layouts are C layout for x86-64
(`extern struct`), offsets given below so nothing need be inferred. (`extern struct`), offsets given below so nothing need be inferred.
The kernel never parses any of this — it only moves the bytes The kernel resolves NAMES (the mount table) but never parses these
(docs/syscall.md); files are entirely a user-space affair. messages — it moves the bytes; file state is entirely the backend's affair.
With clients holding backend node ids directly, a backend records each open
handle's owner and sweeps a dead client's handles via the published process
exit events.
## Request header — 32 bytes ## Request header — 32 bytes
@@ -90,13 +97,14 @@ Notes per operation:
position. Each call returns exactly one entry; the client increments the position. Each call returns exactly one entry; the client increments the
cursor by 1. A reply with `len` 0 is end-of-directory. The directory must cursor by 1. A reply with `len` 0 is end-of-directory. The directory must
have been opened with the `directory` flag. have been opened with the `directory` flag.
- **mount** — the one operation that passes a **capability**: the caller - **mount / unmount** — RETIRED from the wire: mounting is the `fs_mount`
(a filesystem server, e.g. FAT) sends its own request endpoint as the syscall now (a filesystem server passes its endpoint handle; possession is
`ipc_call` capability argument, and the router forwards everything under the capability, exactly the trust of the old cap-passing op). The op
the mount point to it — speaking this same protocol, with paths rewritten numbers stay reserved. Mount-prefix semantics are unchanged: prefixes
relative to the mount. Prefixes match at path boundaries only match at path boundaries only (`/mnt/usb` never captures `/mnt/usbextra`),
(`/mnt/usb` never captures `/mnt/usbextra`); the longest matching prefix the longest matching prefix wins, and an optional backend-side rewrite
wins. prefix maps a mount into the backend's namespace (fat serves `/mnt/usb`
from its volume root and `/var` from its `/var` subtree).
- **rename** — same-directory rename only (the router requires old and new to - **rename** — same-directory rename only (the router requires old and new to
resolve under one mount). resolve under one mount).
+11 -1
View File
@@ -54,6 +54,13 @@ pub const Device = struct {
} }
}; };
/// One lookup attempt, no waiting — for a server that retries on its own
/// timer (the fat service) instead of blocking its harness in here.
pub fn tryOpen() ?Device {
if (ipc.lookup(.block)) |handle| return .{ .endpoint = handle };
return null;
}
/// Look up the block device, retrying generously while the USB storage chain /// Look up the block device, retrying generously while the USB storage chain
/// (controller reset, enumeration, mass-storage bring-up) comes up. /// (controller reset, enumeration, mass-storage bring-up) comes up.
pub fn open() ?Device { pub fn open() ?Device {
@@ -61,7 +68,10 @@ pub fn open() ?Device {
// enumeration, mass-storage bring-up) must complete first, which can take // enumeration, mass-storage bring-up) must complete first, which can take
// tens of seconds under emulation. // tens of seconds under emulation.
var attempts: usize = 0; var attempts: usize = 0;
while (attempts < 1200) : (attempts += 1) { // 30 s covers the slowest observed healthy chain (a flaky QEMU enumeration
// completed at ~24 s); a machine whose stick genuinely failed setup should
// not sit a further minute pretending otherwise.
while (attempts < 600) : (attempts += 1) {
if (ipc.lookup(.block)) |handle| return .{ .endpoint = handle }; if (ipc.lookup(.block)) |handle| return .{ .endpoint = handle };
system.sleep(50); system.sleep(50);
} }
+129 -49
View File
@@ -1,5 +1,5 @@
//! runtime.fs — the danos-native file API. A program opens, reads, writes, and //! runtime.fs — the danos-native file API. A program opens, reads, writes, and
//! lists files served by the user-space VFS (system/services/vfs), each call //! lists files through the kernel VFS root (resolve + redirect), each call
//! marshalling a vfs-protocol request over IPC. This is the danos-native layer //! marshalling a vfs-protocol request over IPC. This is the danos-native layer
//! danos programs use directly; it is also where the file operations that later //! danos programs use directly; it is also where the file operations that later
//! become `std.os.danos` are staged (see docs/zig-self-hosting.md). It replaces //! become `std.os.danos` are staged (see docs/zig-self-hosting.md). It replaces
@@ -12,6 +12,7 @@
const std = @import("std"); const std = @import("std");
const ipc = @import("ipc.zig"); const ipc = @import("ipc.zig");
const system = @import("system.zig");
const protocol = @import("vfs-protocol"); const protocol = @import("vfs-protocol");
/// The kind of a filesystem node — re-exported so a caller need not import the /// The kind of a filesystem node — re-exported so a caller need not import the
@@ -60,23 +61,37 @@ pub const OpenOptions = struct {
} }
}; };
// The VFS server endpoint, looked up once by well-known id and cached. // The route to a path: the kernel resolves (fs_resolve) and either serves the
var vfs_handle: ipc.Handle = 0; // node itself (the initrd at /system — a permanent token) or redirects us to
var vfs_resolved = false; // the owning filesystem backend's endpoint, to which we speak the vfs-protocol
fn vfs() ?ipc.Handle { // rendezvous directly with the rewritten mount-relative path.
if (!vfs_resolved) { const Route = union(enum) {
vfs_handle = ipc.lookup(.vfs) orelse return null; kernel: u64,
vfs_resolved = true; backend: struct { handle: ipc.Handle, path: [224]u8, path_len: usize },
fn backendPath(self: *const Route) []const u8 {
return self.backend.path[0..self.backend.path_len];
}
};
fn resolve(path: []const u8, flags: usize) ?Route {
var out: [224]u8 = undefined;
const route = system.fsResolve(path, flags, &out) orelse return null;
switch (route) {
.kernel => |token| return .{ .kernel = token },
.backend => |b| {
var r: Route = .{ .backend = .{ .handle = b.handle, .path = undefined, .path_len = b.path_len } };
@memcpy(r.backend.path[0..b.path_len], out[0..b.path_len]);
return r;
},
} }
return vfs_handle;
} }
const Result = struct { reply: protocol.Reply, payload: []u8 }; const Result = struct { reply: protocol.Reply, payload: []u8 };
// One request/reply round trip: [Request header][send payload] -> VFS -> // One request/reply round trip: [Request header][send payload] -> backend ->
// [Reply header][receive payload]. The receive payload lands in `out`. // [Reply header][receive payload]. The receive payload lands in `out`.
fn transact(request: protocol.Request, send: []const u8, out: []u8) ?Result { fn transact(h: ipc.Handle, request: protocol.Request, send: []const u8, out: []u8) ?Result {
const h = vfs() orelse return null;
var message: [protocol.message_maximum]u8 = undefined; var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request)); @memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const slen = @min(send.len, protocol.maximum_payload); const slen = @min(send.len, protocol.maximum_payload);
@@ -95,13 +110,21 @@ fn transact(request: protocol.Request, send: []const u8, out: []u8) ?Result {
pub const File = struct { pub const File = struct {
node: u64, node: u64,
offset: u64 = 0, offset: u64 = 0,
/// The owning backend's endpoint, or null for a kernel-served node (the
/// read-only /system tree), whose `node` is a permanent fs_node token.
backend: ?ipc.Handle = null,
/// Read up to `buffer.len` bytes at the current offset; returns the count, or /// Read up to `buffer.len` bytes at the current offset; returns the count, or
/// null on error. /// null on error.
pub fn read(self: *File, buffer: []u8) ?usize { pub fn read(self: *File, buffer: []u8) ?usize {
const h = self.backend orelse {
const n = system.fsNodeRead(self.node, self.offset, buffer) orelse return null;
self.offset += n;
return n;
};
const want: u32 = @intCast(@min(buffer.len, protocol.maximum_payload)); const want: u32 = @intCast(@min(buffer.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .read, .node = self.node, .offset = self.offset, .len = want, .flags = 0 }; const request = protocol.Request{ .operation = .read, .node = self.node, .offset = self.offset, .len = want, .flags = 0 };
const r = transact(request, &.{}, buffer) orelse return null; const r = transact(h, request, &.{}, buffer) orelse return null;
if (r.reply.status != 0) return null; if (r.reply.status != 0) return null;
self.offset += r.reply.len; self.offset += r.reply.len;
return r.reply.len; return r.reply.len;
@@ -109,11 +132,13 @@ pub const File = struct {
/// Write `data` at the current offset; returns the count written. A single /// Write `data` at the current offset; returns the count written. A single
/// call is capped at the VFS payload size, so the return may be short — use /// call is capped at the VFS payload size, so the return may be short — use
/// `writeAll` to write the whole slice. Null on error. /// `writeAll` to write the whole slice. Null on error (kernel-served nodes
/// are read-only).
pub fn write(self: *File, data: []const u8) ?usize { pub fn write(self: *File, data: []const u8) ?usize {
const h = self.backend orelse return null;
const want: u32 = @intCast(@min(data.len, protocol.maximum_payload)); const want: u32 = @intCast(@min(data.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .write, .node = self.node, .offset = self.offset, .len = want, .flags = 0 }; const request = protocol.Request{ .operation = .write, .node = self.node, .offset = self.offset, .len = want, .flags = 0 };
const r = transact(request, data[0..want], &.{}) orelse return null; const r = transact(h, request, data[0..want], &.{}) orelse return null;
if (r.reply.status != 0) return null; if (r.reply.status != 0) return null;
self.offset += r.reply.len; self.offset += r.reply.len;
return r.reply.len; return r.reply.len;
@@ -138,27 +163,40 @@ pub const File = struct {
/// This file's metadata. /// This file's metadata.
pub fn attributes(self: *File) ?Attributes { pub fn attributes(self: *File) ?Attributes {
const h = self.backend orelse {
const a = system.fsNodeStatus(self.node) orelse return null;
return .{ .size = a.size, .kind = if (a.kind == system.file_kind_directory) .directory else .regular, .mtime = a.mtime };
};
const request = protocol.Request{ .operation = .status, .node = self.node, .offset = 0, .len = 0, .flags = 0 }; const request = protocol.Request{ .operation = .status, .node = self.node, .offset = 0, .len = 0, .flags = 0 };
var buffer: [@sizeOf(protocol.FileStatus)]u8 = undefined; var buffer: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return null; const r = transact(h, request, &.{}, &buffer) orelse return null;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return null; if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return null;
const status = std.mem.bytesToValue(protocol.FileStatus, buffer[0..@sizeOf(protocol.FileStatus)]); const status = std.mem.bytesToValue(protocol.FileStatus, buffer[0..@sizeOf(protocol.FileStatus)]);
return .{ .size = status.size, .kind = kindFromWire(status.kind), .mtime = status.mtime }; return .{ .size = status.size, .kind = kindFromWire(status.kind), .mtime = status.mtime };
} }
/// Release the VFS's open handle for this file. /// Release the backend's open handle for this file. Kernel-served node
/// tokens are permanent — nothing to release.
pub fn close(self: *File) void { pub fn close(self: *File) void {
const h = self.backend orelse return;
const request = protocol.Request{ .operation = .close, .node = self.node, .offset = 0, .len = 0, .flags = 0 }; const request = protocol.Request{ .operation = .close, .node = self.node, .offset = 0, .len = 0, .flags = 0 };
_ = transact(request, &.{}, &.{}); _ = transact(h, request, &.{}, &.{});
} }
}; };
/// Open (or create, with `.create`) `path`. Returns the open file, or null. /// Open (or create, with `.create`) `path`. Returns the open file, or null.
pub fn open(path: []const u8, options: OpenOptions) ?File { pub fn open(path: []const u8, options: OpenOptions) ?File {
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = options.wireFlags() }; const route = resolve(path, options.wireFlags()) orelse return null;
const r = transact(request, path, &.{}) orelse return null; switch (route) {
if (r.reply.status != 0) return null; .kernel => |token| return .{ .node = token, .backend = null },
return .{ .node = r.reply.node }; .backend => |b| {
const relative = route.backendPath();
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = options.wireFlags() };
const r = transact(b.handle, request, relative, &.{}) orelse return null;
if (r.reply.status != 0) return null;
return .{ .node = r.reply.node, .backend = b.handle };
},
}
} }
/// A path's metadata without keeping it open (open -> status -> close). /// A path's metadata without keeping it open (open -> status -> close).
@@ -190,13 +228,27 @@ pub const Entry = struct {
pub const Directory = struct { pub const Directory = struct {
node: u64, node: u64,
cursor: u64 = 0, cursor: u64 = 0,
backend: ?ipc.Handle = null,
/// Fill `entry` with the next directory entry; false at end of directory or /// Fill `entry` with the next directory entry; false at end of directory or
/// on error. /// on error.
pub fn next(self: *Directory, entry: *Entry) bool { pub fn next(self: *Directory, entry: *Entry) bool {
const h = self.backend orelse {
var buffer: [@sizeOf(system.DirectoryEntryHeader) + 64]u8 = undefined;
const n = system.fsNodeReaddir(self.node, self.cursor, &buffer) orelse return false;
if (n < @sizeOf(system.DirectoryEntryHeader)) return false; // end
const header = std.mem.bytesToValue(system.DirectoryEntryHeader, buffer[0..@sizeOf(system.DirectoryEntryHeader)]);
entry.kind = if (header.kind == system.file_kind_directory) .directory else .regular;
entry.size = header.size;
const nlen = @min(@as(usize, header.name_len), entry.name_buffer.len);
@memcpy(entry.name_buffer[0..nlen], buffer[@sizeOf(system.DirectoryEntryHeader)..][0..nlen]);
entry.name_len = nlen;
self.cursor += 1;
return true;
};
const request = protocol.Request{ .operation = .readdir, .node = self.node, .offset = self.cursor, .len = 0, .flags = 0 }; const request = protocol.Request{ .operation = .readdir, .node = self.node, .offset = self.cursor, .len = 0, .flags = 0 };
var buffer: [protocol.message_maximum]u8 = undefined; var buffer: [protocol.message_maximum]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return false; const r = transact(h, request, &.{}, &buffer) orelse return false;
if (r.reply.status != 0 or r.reply.len == 0) return false; // error or EOF if (r.reply.status != 0 or r.reply.len == 0) return false; // error or EOF
if (r.payload.len < protocol.directory_entry_size) return false; if (r.payload.len < protocol.directory_entry_size) return false;
const header = std.mem.bytesToValue(protocol.DirectoryEntry, r.payload[0..protocol.directory_entry_size]); const header = std.mem.bytesToValue(protocol.DirectoryEntry, r.payload[0..protocol.directory_entry_size]);
@@ -210,9 +262,9 @@ pub const Directory = struct {
return true; return true;
} }
/// Release the VFS's open handle for this directory. /// Release the backend's open handle for this directory.
pub fn close(self: *Directory) void { pub fn close(self: *Directory) void {
var f = File{ .node = self.node }; var f = File{ .node = self.node, .backend = self.backend };
f.close(); f.close();
} }
}; };
@@ -220,13 +272,18 @@ pub const Directory = struct {
/// Open `path` as a directory for listing. Returns null if it isn't one / on error. /// Open `path` as a directory for listing. Returns null if it isn't one / on error.
pub fn openDirectory(path: []const u8) ?Directory { pub fn openDirectory(path: []const u8) ?Directory {
const file = open(path, .{ .directory = true }) orelse return null; const file = open(path, .{ .directory = true }) orelse return null;
return .{ .node = file.node }; return .{ .node = file.node, .backend = file.backend };
} }
// A path-based request that returns only a status (mkdir, unlink). // A path-based request that returns only a status (mkdir, unlink). Kernel-served
// paths (the read-only /system) refuse mutation by construction: the resolve
// must land on a backend.
fn pathOperation(operation: protocol.Operation, path: []const u8) bool { fn pathOperation(operation: protocol.Operation, path: []const u8) bool {
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = 0 }; const route = resolve(path, 0) orelse return false;
const r = transact(request, path, &.{}) orelse return false; if (route != .backend) return false;
const relative = route.backendPath();
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = 0 };
const r = transact(route.backend.handle, request, relative, &.{}) orelse return false;
return r.reply.status == 0; return r.reply.status == 0;
} }
@@ -236,39 +293,62 @@ pub fn makeDirectory(path: []const u8) bool {
return pathOperation(.mkdir, path); return pathOperation(.mkdir, path);
} }
/// Create every missing directory along `path` (mkdir -p). Probes each prefix
/// with `exists` first — a FAT mkdir of an existing name is refused, and the
/// probe keeps the common "already there" case cheap. Returns true when the
/// whole path exists afterwards.
pub fn makePath(path: []const u8) bool {
var end: usize = 0;
while (end < path.len) {
end += 1;
while (end < path.len and path[end] != '/') end += 1;
const prefix = path[0..end];
if (prefix.len == 0 or (prefix.len == 1 and prefix[0] == '/')) continue;
// Best-effort per prefix: components at or above a mount point ("/mnt")
// are router names, not filesystem nodes — they neither exist as nodes
// nor accept mkdir, and that is fine. Only the final verdict counts.
if (!exists(prefix)) _ = makeDirectory(prefix);
}
return exists(path);
}
/// Remove the file at `path`. Returns true on success. Directories are refused /// Remove the file at `path`. Returns true on success. Directories are refused
/// (a separate directory-removal would have to check emptiness). /// (a separate directory-removal would have to check emptiness).
pub fn remove(path: []const u8) bool { pub fn remove(path: []const u8) bool {
return pathOperation(.unlink, path); return pathOperation(.unlink, path);
} }
/// Rename `old_path` to `new_path`. Both must be in the same directory (same- /// Rename `old_path` to `new_path`. Both must resolve to the SAME filesystem
/// directory, 8.3-name rename only for now). Returns true on success. /// backend (same-directory, 8.3-name rename only for now). Returns true on
/// success.
pub fn rename(old_path: []const u8, new_path: []const u8) bool { pub fn rename(old_path: []const u8, new_path: []const u8) bool {
const total = old_path.len + 1 + new_path.len; const old_route = resolve(old_path, 0) orelse return false;
const new_route = resolve(new_path, 0) orelse return false;
if (old_route != .backend or new_route != .backend) return false;
if (old_route.backend.handle != new_route.backend.handle) return false; // cross-filesystem
const old_relative = old_route.backendPath();
const new_relative = new_route.backendPath();
const total = old_relative.len + 1 + new_relative.len;
if (total > protocol.maximum_payload) return false; if (total > protocol.maximum_payload) return false;
var payload: [protocol.maximum_payload]u8 = undefined; var payload: [protocol.maximum_payload]u8 = undefined;
@memcpy(payload[0..old_path.len], old_path); @memcpy(payload[0..old_relative.len], old_relative);
payload[old_path.len] = 0; payload[old_relative.len] = 0;
@memcpy(payload[old_path.len + 1 ..][0..new_path.len], new_path); @memcpy(payload[old_relative.len + 1 ..][0..new_relative.len], new_relative);
const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 }; const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 };
const r = transact(request, payload[0..total], &.{}) orelse return false; const r = transact(old_route.backend.handle, request, payload[0..total], &.{}) orelse return false;
return r.reply.status == 0; return r.reply.status == 0;
} }
/// Mount a filesystem backend (its server endpoint) at absolute path `target`; /// Mount a filesystem backend (its server endpoint) at absolute path `target`;
/// the VFS then routes everything under `target` to that backend. This is the one /// the kernel VFS then routes everything under `target` to that backend.
/// call that hands the VFS a capability (the backend endpoint). Returns true on /// Possession of the endpoint handle is the capability. Returns true on success.
/// success.
pub fn mount(target: []const u8, backend: ipc.Handle) bool { pub fn mount(target: []const u8, backend: ipc.Handle) bool {
const h = vfs() orelse return false; return system.fsMount(target, backend, "");
const request = protocol.Request{ .operation = .mount, .node = 0, .offset = 0, .len = @intCast(target.len), .flags = 0 }; }
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request)); /// As `mount`, with a backend-side rewrite prefix: a path under `target` reaches
const tlen = @min(target.len, protocol.maximum_payload); /// the backend as `rewrite` + the mount-relative tail. How one volume serves two
@memcpy(message[protocol.request_size..][0..tlen], target[0..tlen]); /// mounts ("/mnt/usb" from its root, "/var" from its /var subtree).
var rbuf: [protocol.message_maximum]u8 = undefined; pub fn mountRewritten(target: []const u8, backend: ipc.Handle, rewrite: []const u8) bool {
const result = ipc.callCap(h, message[0 .. protocol.request_size + tlen], &rbuf, backend) catch return false; return system.fsMount(target, backend, rewrite);
if (result.len < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]).status == 0;
} }
+51
View File
@@ -0,0 +1,51 @@
//! The per-process logger: std.log wired to the tagged kernel log ring.
//!
//! A program just calls `std.log.info("mounted {s}", .{path})` (or a scoped
//! logger); this backend formats the line into a fixed buffer and emits ONE
//! `debug_write` record carrying the level. The kernel stamps the record with
//! the sender's pid and task name (its binary path) — the process does NOT put
//! its own name in the payload; attribution is the kernel's, structural and
//! unforgeable. Serial shows the kernel-rendered `<path>: message` line, and
//! the logger service demultiplexes the ring into one file per process.
//!
//! Installed for every user binary by the root shim (library/runtime/root.zig)
//! via `std_options`; a program can override by declaring its own
//! `pub const std_options`.
const std = @import("std");
const system = @import("system.zig");
fn levelOf(comptime level: std.log.Level) system.KlogLevel {
return switch (level) {
.err => .err,
.warn => .warn,
.info => .info,
.debug => .debug,
};
}
pub fn logFn(
comptime level: std.log.Level,
comptime scope: @EnumLiteral(),
comptime format: []const u8,
args: anytype,
) void {
// One record = one line = at most klog_maximum_message bytes of payload.
// On overflow keep what fits and end with "~" so the record is still a
// whole line (the kernel would split an embedded rest anyway).
var buffer: [256]u8 = undefined;
const prefix = if (scope == .default) "" else "(" ++ @tagName(scope) ++ ") ";
const line = std.fmt.bufPrint(&buffer, prefix ++ format, args) catch truncated: {
buffer[buffer.len - 1] = '~';
break :truncated buffer[0..];
};
_ = system.writeRecord(levelOf(level), line);
}
/// The std.Options the root shim installs unless the program overrides it.
/// Debug level: filtering is the log *reader's* job here — the ring is cheap,
/// serial is a dev convenience, and the logger service keeps everything.
pub const default_options: std.Options = .{
.log_level = .debug,
.logFn = logFn,
};
+6
View File
@@ -14,6 +14,12 @@ pub const main = program.main;
/// The panic handler for every safety check in the image (runtime.start.panic). /// The panic handler for every safety check in the image (runtime.start.panic).
pub const panic = runtime.panic; pub const panic = runtime.panic;
/// std.log for every user binary goes to the tagged kernel log ring (the kernel
/// stamps the sender; see runtime.log). A program overrides by declaring its
/// own `pub const std_options`.
pub const std_options: @import("std").Options =
if (@hasDecl(program, "std_options")) program.std_options else runtime.log.default_options;
comptime { comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image _ = &runtime.start._start; // pull the runtime entry shim into the image
} }
+3 -2
View File
@@ -11,6 +11,7 @@
//! nothing to declare per source file. //! nothing to declare per source file.
pub const system = @import("system.zig"); pub const system = @import("system.zig");
pub const log = @import("log.zig");
/// Monotonic time, delays, and deadlines over the kernel clock/sleep/timer syscalls /// Monotonic time, delays, and deadlines over the kernel clock/sleep/timer syscalls
/// — an `Instant`/`Duration` front door, no time service (docs/timers.md). /// — an `Instant`/`Duration` front door, no time service (docs/timers.md).
pub const time = @import("time.zig"); pub const time = @import("time.zig");
@@ -38,9 +39,9 @@ pub const device = @import("device.zig");
pub const dma = @import("dma.zig"); pub const dma = @import("dma.zig");
/// Shared cacheable memory: create a region + capability, pass the capability to another /// Shared cacheable memory: create a region + capability, pass the capability to another
/// process (an `ipc_call` send_cap), map the same pages there. See library/runtime/shm.zig /// process (an `ipc_call` send_cap), map the same pages there. See library/runtime/shared-memory.zig
/// and docs/display-v2.md. /// and docs/display-v2.md.
pub const shm = @import("shm.zig"); pub const shared_memory = @import("shared-memory.zig");
/// USB class-driver client: open a device on the xHCI bus and drive it /// USB class-driver client: open a device on the xHCI bus and drive it
/// (control / interrupt / bulk transfers). See library/runtime/usb.zig. /// (control / interrupt / bulk transfers). See library/runtime/usb.zig.
@@ -1,7 +1,7 @@
//! User-space shared memory: `shm_create` / `shm_map`. A process creates a shareable, //! User-space shared memory: `shared_memory_create` / `shared_memory_map`. A process creates a shareable,
//! zeroed, cacheable RAM region and gets back a pointer plus a **capability handle**; it //! zeroed, cacheable RAM region and gets back a pointer plus a **capability handle**; it
//! passes that handle to another process as an `ipc_call` send_cap, and the receiver //! passes that handle to another process as an `ipc_call` send_cap, and the receiver
//! `shm_map`s it to map the same physical pages. The kernel primitive under the display //! `shared_memory_map`s it to map the same physical pages. The kernel primitive under the display
//! compositor↔native-driver and app↔compositor surface paths (docs/display-v2.md). The //! compositor↔native-driver and app↔compositor surface paths (docs/display-v2.md). The
//! generalization of capability passing from endpoints to memory objects. //! generalization of capability passing from endpoints to memory objects.
@@ -30,7 +30,7 @@ pub fn create(len: usize) ?Region {
asm volatile ("syscall" asm volatile ("syscall"
: [rax] "={rax}" (rax), : [rax] "={rax}" (rax),
[rdx] "={rdx}" (rdx), [rdx] "={rdx}" (rdx),
: [n] "{rax}" (@intFromEnum(abi.SystemCall.shm_create)), : [n] "{rax}" (@intFromEnum(abi.SystemCall.shared_memory_create)),
[a0] "{rdi}" (len), [a0] "{rdi}" (len),
: .{ .rcx = true, .r11 = true, .memory = true }); : .{ .rcx = true, .r11 = true, .memory = true });
if (failed(rax)) return null; if (failed(rax)) return null;
@@ -41,7 +41,7 @@ pub fn create(len: usize) ?Region {
/// `ipc_call` send_cap) into its address space — the same physical pages the creator sees. /// `ipc_call` send_cap) into its address space — the same physical pages the creator sees.
/// Returns the pointer, or null on failure. /// Returns the pointer, or null on failure.
pub fn map(handle: ipc.Handle) ?[*]u8 { pub fn map(handle: ipc.Handle) ?[*]u8 {
const r = sc.systemCall1(.shm_map, handle); const r = sc.systemCall1(.shared_memory_map, handle);
if (failed(r)) return null; if (failed(r)) return null;
return @ptrFromInt(r); return @ptrFromInt(r);
} }
@@ -51,7 +51,7 @@ pub fn map(handle: ipc.Handle) ?[*]u8 {
/// the region length is all a device needs — e.g. a virtio-gpu driver programming an /// the region length is all a device needs — e.g. a virtio-gpu driver programming an
/// `attach_backing`. Returns null on failure. /// `attach_backing`. Returns null on failure.
pub fn physical(handle: ipc.Handle) ?usize { pub fn physical(handle: ipc.Handle) ?usize {
const r = sc.systemCall1(.shm_physical, handle); const r = sc.systemCall1(.shared_memory_physical, handle);
if (failed(r)) return null; if (failed(r)) return null;
return r; return r;
} }
+112 -11
View File
@@ -21,10 +21,34 @@ pub fn yield() void {
_ = sc.systemCall0(.yield); _ = sc.systemCall0(.yield);
} }
/// Write raw bytes to the kernel log (a bring-up diagnostic; real output goes /// The tagged-log level of a record — re-exported so runtime.log and the logger
/// through the console/VFS later). Returns the byte count, or a wrapped -1. /// service don't import `abi` themselves.
pub const KlogLevel = abi.KlogLevel;
pub const KlogStatus = abi.KlogStatus;
pub const KlogRecordHeader = abi.KlogRecordHeader;
pub const klog_record_header_size = abi.klog_record_header_size;
pub const klog_record_alignment = abi.klog_record_alignment;
pub const klog_record_magic = abi.klog_record_magic;
pub const klog_flag_truncated = abi.klog_flag_truncated;
pub const klog_maximum_message = abi.klog_maximum_message;
pub const maximum_process_name = abi.maximum_process_name;
pub const FileAttributes = abi.FileAttributes;
pub const DirectoryEntryHeader = abi.DirectoryEntryHeader;
pub const file_kind_regular = abi.file_kind_regular;
pub const file_kind_directory = abi.file_kind_directory;
/// Write raw bytes to the kernel log (bring-up/panic diagnostics; ordinary
/// output goes through std.log -> writeRecord). The kernel stamps the record
/// with this process's id and name. Returns the byte count, or a wrapped -1.
pub fn write(message: []const u8) usize { pub fn write(message: []const u8) usize {
return sc.systemCall2(.debug_write, @intFromPtr(message.ptr), message.len); return writeRecord(.raw, message);
}
/// Emit one leveled record into the tagged kernel log ring. The kernel stamps
/// pid/name/sequence/timestamp; the payload should be a single line (embedded
/// newlines split into further records).
pub fn writeRecord(level: KlogLevel, message: []const u8) usize {
return sc.systemCall3(.debug_write, @intFromPtr(message.ptr), message.len, @intFromEnum(level));
} }
/// Block the caller for `ms` milliseconds. /// Block the caller for `ms` milliseconds.
@@ -59,14 +83,91 @@ pub fn wallClock() u64 {
return @intCast(sc.systemCall0(.wall_clock)); return @intCast(sc.systemCall0(.wall_clock));
} }
/// Copy bytes out of the kernel's in-memory diagnostic log — the accumulated /// Copy bytes out of the tagged kernel log ring — framed records of everything
/// stream of everything `write` (and the kernel itself) has emitted — starting at /// every process (and the kernel) has emitted — starting at stream offset
/// `offset`, into `out`. Returns the number of bytes copied (0 at end of buffer). /// `offset`, into `out`. Returns the byte count (0 = caught up), or null when
/// A program reads the whole log by looping from offset 0, advancing by the return /// `offset` fell behind the ring's tail (those records were overwritten) or
/// value, until it gets 0. This is how the boot log is persisted to disk on a /// lies past its head; re-sync via `klogStatus`. A reader parses
/// headless/real machine where serial output is otherwise lost. /// [KlogRecordHeader][name][message] frames (8-byte aligned) from the bytes.
pub fn klogRead(offset: usize, out: []u8) usize { pub fn klogRead(offset: u64, out: []u8) ?usize {
return sc.systemCall3(.klog_read, offset, @intFromPtr(out.ptr), out.len); const r = sc.systemCall3(.klog_read, offset, @intFromPtr(out.ptr), out.len);
if (@as(isize, @bitCast(r)) < 0) return null;
return r;
}
/// The log ring's live cursors (oldest retained offset, end of stream, next
/// sequence number) plus the wall-clock time of boot — how a log reader starts,
/// detects loss, and names a per-boot log directory.
pub fn klogStatus() ?KlogStatus {
var status: KlogStatus = undefined;
if (@as(isize, @bitCast(sc.systemCall1(.klog_status, @intFromPtr(&status)))) != 0) return null;
return status;
}
/// Where fs_resolve routed a path: served by the kernel (a permanent node
/// token for fs_node) or by a userspace filesystem backend (an endpoint handle
/// plus the rewritten mount-relative path, returned in the caller's buffer).
pub const FsRoute = union(enum) {
kernel: u64,
backend: struct { handle: usize, path_len: usize },
};
/// Route `path` through the kernel VFS. For a backend route the rewritten
/// mount-relative path lands in `out` (behind a kernel-written length prefix,
/// already stripped here: out[0..path_len] is the path).
pub fn fsResolve(path: []const u8, flags: usize, out: []u8) ?FsRoute {
var rax: usize = undefined;
var rdx: usize = flags; // in: flags (arg #3); out: node token / backend handle
asm volatile ("syscall"
: [rax] "={rax}" (rax),
[rdx] "+{rdx}" (rdx),
: [n] "{rax}" (@intFromEnum(abi.SystemCall.fs_resolve)),
[a0] "{rdi}" (@intFromPtr(path.ptr)),
[a1] "{rsi}" (path.len),
[a3] "{r10}" (@intFromPtr(out.ptr)),
[a4] "{r8}" (out.len),
: .{ .rcx = true, .r11 = true, .memory = true });
if (@as(isize, @bitCast(rax)) < 0) return null;
if (rax == abi.fs_route_kernel) return .{ .kernel = rdx };
if (rax != abi.fs_route_backend) return null;
const path_len = @as(usize, out[0]) | (@as(usize, out[1]) << 8);
if (path_len + 2 > out.len) return null;
std.mem.copyForwards(u8, out[0..path_len], out[2..][0..path_len]);
return .{ .backend = .{ .handle = rdx, .path_len = path_len } };
}
/// Read `out.len` bytes of a kernel-served node at `offset` (fs_node read).
pub fn fsNodeRead(node_token: u64, offset: u64, out: []u8) ?usize {
const r = sc.systemCall5(.fs_node, abi.fs_node_read, node_token, offset, @intFromPtr(out.ptr), out.len);
if (@as(isize, @bitCast(r)) < 0) return null;
return r;
}
/// A kernel-served node's metadata (fs_node status).
pub fn fsNodeStatus(node_token: u64) ?abi.FileAttributes {
var attributes: abi.FileAttributes = undefined;
const r = sc.systemCall5(.fs_node, abi.fs_node_status, node_token, 0, @intFromPtr(&attributes), @sizeOf(abi.FileAttributes));
if (@as(isize, @bitCast(r)) < 0) return null;
return attributes;
}
/// The `cursor`th child of a kernel-served directory (fs_node readdir): fills
/// `out` with [DirectoryEntryHeader][name]; returns total bytes (0 = end).
pub fn fsNodeReaddir(node_token: u64, cursor: u64, out: []u8) ?usize {
const r = sc.systemCall5(.fs_node, abi.fs_node_readdir, node_token, cursor, @intFromPtr(out.ptr), out.len);
if (@as(isize, @bitCast(r)) < 0) return null;
return r;
}
/// Mount a userspace filesystem's endpoint at `prefix`, with an optional
/// backend-side `rewrite` prefix ("" = none). Possession of the endpoint
/// handle is the capability.
pub fn fsMount(prefix: []const u8, backend: usize, rewrite: []const u8) bool {
return sc.systemCall5(.fs_mount, @intFromPtr(prefix.ptr), prefix.len, backend, @intFromPtr(rewrite.ptr), rewrite.len) == 0;
}
pub fn fsUnmount(prefix: []const u8) bool {
return sc.systemCall2(.fs_unmount, @intFromPtr(prefix.ptr), prefix.len) == 0;
} }
/// End the process. Never returns. /// End the process. Never returns.
+98 -7
View File
@@ -58,11 +58,11 @@ pub const SystemCall = enum(u64) {
signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on
process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself) process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself)
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
klog_read = 32, // klog_read(offset, ptr, len) -> bytes copied: copy the kernel RAM log buffer out to a user buffer (for persisting the boot log to disk) klog_read = 32, // klog_read(offset, ptr, len) -> bytes copied: copy tagged log-ring stream bytes from `offset` out to a user buffer; fails once `offset` falls behind the ring's tail (re-sync via klog_status)
wall_clock = 33, // wall_clock() -> Unix epoch seconds (UTC): the RTC wall-clock time, for filesystem timestamps (mtime). Monotonic time is `clock`. wall_clock = 33, // wall_clock() -> Unix epoch seconds (UTC): the RTC wall-clock time, for filesystem timestamps (mtime). Monotonic time is `clock`.
shm_create = 34, // shm_create(len) -> virtual_address (rax), handle (rdx): a shareable, zeroed, cacheable RAM region mapped into this AS; the handle is a capability passed to another process as an ipc_call send_cap (docs/display-v2.md) shared_memory_create = 34, // shared_memory_create(len) -> virtual_address (rax), handle (rdx): a shareable, zeroed, cacheable RAM region mapped into this AS; the handle is a capability passed to another process as an ipc_call send_cap (docs/display-v2.md)
shm_map = 35, // shm_map(cap) -> virtual_address: map the shared region named by a received capability into this address space (the same physical pages the creator sees) shared_memory_map = 35, // shared_memory_map(cap) -> virtual_address: map the shared region named by a received capability into this address space (the same physical pages the creator sees)
shm_physical = 36, // shm_physical(cap) -> physical_address: the guest-physical base of a shared region held by capability, so a driver can program it into a device (e.g. virtio-gpu attach_backing); the pages are contiguous (docs/display-v2.md) shared_memory_physical = 36, // shared_memory_physical(cap) -> physical_address: the guest-physical base of a shared region held by capability, so a driver can program it into a device (e.g. virtio-gpu attach_backing); the pages are contiguous (docs/display-v2.md)
thread_spawn = 37, // thread_spawn(entry, stack_top, arg, exit_endpoint) -> tid: start a task sharing the caller's address space at `entry` on `stack_top`, `arg` in rdi; exit_endpoint (a handle, or no_cap) is notified when it ends — how join waits (docs/threading.md) thread_spawn = 37, // thread_spawn(entry, stack_top, arg, exit_endpoint) -> tid: start a task sharing the caller's address space at `entry` on `stack_top`, `arg` in rdi; exit_endpoint (a handle, or no_cap) is notified when it ends — how join waits (docs/threading.md)
thread_exit = 38, // thread_exit(): end the calling thread, dropping one reference to its address space (destroyed on the last) thread_exit = 38, // thread_exit(): end the calling thread, dropping one reference to its address space (destroyed on the last)
current_core = 39, // current_core() -> index: the dense 0-based index of the core the caller is running on (for parallelism/affinity introspection) current_core = 39, // current_core() -> index: the dense 0-based index of the core the caller is running on (for parallelism/affinity introspection)
@@ -71,6 +71,11 @@ pub const SystemCall = enum(u64) {
thread_self = 42, // thread_self() -> tid: the calling thread's kernel task id (runtime.Thread.getCurrentId) thread_self = 42, // thread_self() -> tid: the calling thread's kernel task id (runtime.Thread.getCurrentId)
thread_join = 43, // thread_join(tid) -> 0: block until the thread with id `tid` has exited (runtime.Thread.join; no per-thread IPC endpoint) (docs/threading.md) thread_join = 43, // thread_join(tid) -> 0: block until the thread with id `tid` has exited (runtime.Thread.join; no per-thread IPC endpoint) (docs/threading.md)
set_thread_pointer = 44, // set_thread_pointer(addr) -> 0: set the caller's thread pointer (user-space TLS base; x86_64 IA32_FS_BASE, aarch64 TPIDR_EL0); restored per task across context switches (docs/threading-plan.md M10) set_thread_pointer = 44, // set_thread_pointer(addr) -> 0: set the caller's thread pointer (user-space TLS base; x86_64 IA32_FS_BASE, aarch64 TPIDR_EL0); restored per task across context switches (docs/threading-plan.md M10)
klog_status = 45, // klog_status(ptr) -> 0: copy a KlogStatus (ring cursors + the boot wall-clock anchor) out to a user buffer
fs_resolve = 46, // fs_resolve(path_ptr, path_len, flags, out_ptr, out_cap) -> route tag (rax: fs_route_*) + node token or backend handle (rdx); a backend resolve writes the rewritten mount-relative path into out (length in r8 via third result)
fs_node = 47, // fs_node(op, node_token, offset, buf_ptr, buf_len) -> bytes/0/-errno: read/status/readdir on a kernel-served node (op values mirror the vfs-protocol Operation numbers)
fs_mount = 48, // fs_mount(prefix_ptr, prefix_len, backend_handle, rewrite_ptr, rewrite_len) -> 0/-errno: mount a userspace filesystem's endpoint at an absolute prefix (possession of the handle is the capability)
fs_unmount = 49, // fs_unmount(prefix_ptr, prefix_len) -> 0/-errno: remove a backend mount
_, _,
}; };
@@ -87,7 +92,7 @@ pub const futex_timed_out: u64 = 2; // the timeout elapsed before a wake
/// never delivered to the faulting process (recovery is restart, not a handler). /// never delivered to the faulting process (recovery is restart, not a handler).
pub const ExitReason = enum(u8) { pub const ExitReason = enum(u8) {
exited = 0, // returned from main / called exit exited = 0, // returned from main / called exit
aborted = 1, // deliberate self-termination (reserved: no abort path yet) aborted = 1, // deliberate FAILURE exit (exit with a nonzero code): supervisors restart these, unlike a clean .exited
segmentation_fault = 2, // page fault segmentation_fault = 2, // page fault
illegal_instruction = 3, // invalid opcode illegal_instruction = 3, // invalid opcode
arithmetic_fault = 4, // divide error, x87 or SIMD fault arithmetic_fault = 4, // divide error, x87 or SIMD fault
@@ -187,11 +192,97 @@ pub const ProcessDescriptor = extern struct {
name: [maximum_process_name]u8, // argv[0] at spawn; empty for kernel tasks name: [maximum_process_name]u8, // argv[0] at spawn; empty for kernel tasks
}; };
// --- the tagged kernel log ring (klog) ---------------------------------------
// Every `debug_write` becomes one RECORD per payload line, stamped by the kernel
// with the sender's pid, task name (its binary path), level, a per-boot sequence
// number, and a monotonic timestamp. `klog_read` copies raw stream bytes — a
// reader parses [KlogRecordHeader][name][message] frames, each padded to
// `klog_record_alignment`. Sequence gaps tell a reader exactly how many records
// the ring overwrote while it wasn't looking.
/// Log level of a klog record — std.log's levels plus `raw` (untagged bytes:
/// kernel prints and legacy runtime.system.write output).
pub const KlogLevel = enum(u8) { err = 0, warn = 1, info = 2, debug = 3, raw = 4 };
/// "RK" — leads every record; a parser's resync/corruption guard.
pub const klog_record_magic: u16 = 0x4B52;
/// KlogRecordHeader.flags bit: the emitter truncated the payload to fit.
pub const klog_flag_truncated: u8 = 1;
/// Header of one ring record, followed by `name_len` bytes of task name and
/// `message_len` bytes of payload; the whole record is padded to 8 bytes.
pub const KlogRecordHeader = extern struct {
magic: u16, // klog_record_magic
level: KlogLevel,
name_len: u8, // 0..maximum_process_name
pid: u32, // sender process id; 0 = the kernel itself
sequence: u64, // per-boot monotonic record number (gaps = lost records)
timestamp_ns: u64, // monotonic ns since boot (the `clock` timebase)
message_len: u16, // payload bytes (excludes the record's trailing pad)
flags: u8, // klog_flag_* bits
_reserved: [5]u8,
};
pub const klog_record_header_size: usize = 32; // @sizeOf(KlogRecordHeader), pinned by a test
pub const klog_record_alignment: usize = 8;
/// Per-record payload cap (one line; longer emitter lines are truncated).
pub const klog_maximum_message: usize = 256;
// --- the kernel VFS root (resolve + redirect) --------------------------------
// fs_resolve routes a path through the kernel mount table. Kernel-backed mounts
// (the initrd at /system) resolve to a permanent node TOKEN served by fs_node;
// userspace mounts resolve to the backend's endpoint handle (installed in the
// caller's table, deduplicated) plus the rewritten mount-relative path — the
// caller then speaks the vfs-protocol to the backend directly. The kernel never
// blocks on a userspace filesystem.
/// fs_resolve result tags (rax).
pub const fs_route_kernel: u64 = 0; // rdx = node token; serve via fs_node
pub const fs_route_backend: u64 = 1; // rdx = endpoint handle; speak vfs-protocol
/// fs_node operations — the same numbers as the vfs-protocol Operation enum, so
/// client code shares one vocabulary.
pub const fs_node_read: u64 = 2;
pub const fs_node_status: u64 = 4;
pub const fs_node_readdir: u64 = 5;
/// fs_resolve flags (same values as the vfs-protocol open flags).
pub const fs_flag_create: u64 = 1;
/// FileStatus-shaped node metadata (matches the vfs-protocol payload layout).
pub const file_kind_regular: u32 = 0;
pub const file_kind_directory: u32 = 1;
pub const FileAttributes = extern struct {
size: u64,
kind: u32,
_pad: u32 = 0,
mtime: u64 = 0,
};
/// One fs_node readdir result: the header, followed by `name_len` name bytes in
/// the caller's buffer (matches the vfs-protocol DirectoryEntry layout).
pub const DirectoryEntryHeader = extern struct {
kind: u32,
name_len: u32,
size: u64,
};
/// The klog_status copy-out: the ring's live cursors plus the wall-clock time
/// of boot — the anchor a log persister names its per-boot directory with and
/// combines with record timestamps for wall-clock line stamps.
pub const KlogStatus = extern struct {
tail: u64, // oldest retained stream offset — always a record boundary
head: u64, // next byte to be written (end of stream)
next_sequence: u64, // the sequence the next record will get
boot_unix_seconds: u64, // wall-clock time of boot (RTC anchor)
};
/// Well-known IPC service ids for the bootstrap name registry (create_ipc_endpoint + /// Well-known IPC service ids for the bootstrap name registry (create_ipc_endpoint +
/// ipc_register/ipc_lookup). Small integers, so no string interning is needed /// ipc_register/ipc_lookup). Small integers, so no string interning is needed
/// during bring-up. The VFS server registers under `vfs`; clients look it up. /// during bring-up. The VFS server registers under `vfs`; clients look it up.
pub const ServiceId = enum(u32) { pub const ServiceId = enum(u32) {
vfs = 1, vfs = 1, // RETIRED: the router moved into the kernel (fs_resolve); the slot stays reserved
input = 2, input = 2,
ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes
device_manager = 4, // the tree, the matcher, the supervisor (docs/device-manager.md) device_manager = 4, // the tree, the matcher, the supervisor (docs/device-manager.md)
@@ -200,7 +291,7 @@ pub const ServiceId = enum(u32) {
block = 7, // a block-device driver (USB mass storage today): read/write of fixed-size blocks, the storage a filesystem sits on block = 7, // a block-device driver (USB mass storage today): read/write of fixed-size blocks, the storage a filesystem sits on
fat = 8, // the FAT filesystem server; the VFS mounts it and forwards paths under its mount point (/mnt/usb) to it fat = 8, // the FAT filesystem server; the VFS mounts it and forwards paths under its mount point (/mnt/usb) to it
display = 9, // the display service: owns the framebuffer, composites a layer stack, presents frames (docs/display.md) display = 9, // the display service: owns the framebuffer, composites a layer stack, presents frames (docs/display.md)
shm_test = 10, // the shm test server (V2): a client passes it a shared-memory capability, it maps + verifies (docs/display-v2.md) shared_memory_test = 10, // the shared-memory test server (V2): a client passes it a shared-memory capability, it maps + verifies (docs/display-v2.md)
scanout = 11, // a native scanout driver (virtio-gpu): the compositor finds it here to upgrade off the GOP framebuffer (docs/display-v2.md) scanout = 11, // a native scanout driver (virtio-gpu): the compositor finds it here to upgrade off the GOP framebuffer (docs/display-v2.md)
_, _,
}; };
+10 -9
View File
@@ -37,6 +37,11 @@ pub const Framebuffer = extern struct {
height: u32, // visible rows (e.g. 1080) height: u32, // visible rows (e.g. 1080)
pitch: u32, // bytes from the start of one row to the start of the next pitch: u32, // bytes from the start of one row to the start of the next
format: PixelFormat, format: PixelFormat,
/// The panel's refresh rate in Hz, computed from its EDID preferred timing (pixel
/// clock / total pixels per frame) while GOP was still alive — the one moment it is
/// readable (docs/gop.md). 0 = unknown (no EDID). The display service paces its
/// frame clock by it; without vblank this fixes the *rate*, never the *phase*.
refresh_hz: u32 = 0,
/// Whether a usable framebuffer was handed over. /// Whether a usable framebuffer was handed over.
pub fn present(self: Framebuffer) bool { pub fn present(self: Framebuffer) bool {
@@ -142,15 +147,11 @@ pub const BootInformation = extern struct {
/// A device-tree boot path leaves this 0 and (later) fills a `device_tree_blob` /// A device-tree boot path leaves this 0 and (later) fills a `device_tree_blob`
/// field instead, so the kernel discovers devices without knowing what booted it. /// field instead, so the kernel discovers devices without knowing what booted it.
acpi_rsdp: u64 = 0, acpi_rsdp: u64 = 0,
/// The raw `/system/services/init` ELF image, read off the boot volume by the loader /// The initial_ramdisk image: every user binary from the boot volume's /system
/// into memory that survives the handoff (classified reserved, so the kernel /// tree (init included), packed by the loader into memory that survives the
/// identity-maps it and never allocates over it). 0/0 = no init found — the /// handoff (classified reserved, so the kernel identity-maps it and never
/// kernel boots without user space. Grows into a full initial_ramdisk handoff later. /// allocates over it). Entries are named by full FHS path. 0/0 = no binaries
init_base: u64 = 0, /// found — the kernel boots without user space. See system/initial-ramdisk.zig.
init_len: u64 = 0,
/// The initial_ramdisk image (a bundle of extra user binaries — the VFS server and
/// device drivers), read off the boot volume into memory that survives the
/// handoff, same as `init` above. 0/0 = no initial_ramdisk. See system/initial-ramdisk.zig.
initial_ramdisk_base: u64 = 0, initial_ramdisk_base: u64 = 0,
initial_ramdisk_len: u64 = 0, initial_ramdisk_len: u64 = 0,
}; };
+1
View File
@@ -97,6 +97,7 @@ pub const DisplayInfo = extern struct {
height: u32 = 0, // visible rows height: u32 = 0, // visible rows
pitch: u32 = 0, // bytes from one row's start to the next pitch: u32 = 0, // bytes from one row's start to the next
format: u32 = 0, // a DisplayFormat value format: u32 = 0, // a DisplayFormat value
refresh_hz: u32 = 0, // panel refresh rate from EDID (0 = unknown); see boot-handoff
}; };
/// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree. /// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree.
+235
View File
@@ -70,6 +70,104 @@ pub const Class = enum(u8) {
_, _,
}; };
/// A human-readable name for a device/interface class code, for logs. Unknown
/// codes fall through to "class 0xNN".
pub fn className(class: u8) []const u8 {
return switch (@as(Class, @enumFromInt(class))) {
.per_interface => "per-interface",
.audio => "Audio",
.communications => "Communications",
.hid => "HID",
.physical => "Physical",
.image => "Image",
.printer => "Printer",
.mass_storage => "Mass Storage",
.hub => "Hub",
.cdc_data => "CDC Data",
.smart_card => "Smart Card",
.content_security => "Content Security",
.video => "Video",
.personal_healthcare => "Personal Healthcare",
.audio_video => "Audio/Video",
.billboard => "Billboard",
.type_c_bridge => "Type-C Bridge",
.bulk_display => "Bulk Display",
.mctp => "MCTP",
.i3c => "I3C",
.diagnostic => "Diagnostic",
.wireless_controller => "Wireless Controller",
.miscellaneous => "Miscellaneous",
.application_specific => "Application-specific",
.vendor_specific => "Vendor-specific",
_ => "Unknown",
};
}
/// The USB speed class (as xHCI reports it in PORTSC/slot contexts) named.
pub fn speedName(speed: u32) []const u8 {
return switch (speed) {
1 => "Full-speed",
2 => "Low-speed",
3 => "High-speed",
4 => "SuperSpeed",
5 => "SuperSpeedPlus",
else => "unknown-speed",
};
}
/// A USB3 Port Link State (xHCI PORTSC PLS field) named.
pub fn linkStateName(pls: u32) []const u8 {
return switch (pls) {
0 => "U0",
1 => "U1",
2 => "U2",
3 => "U3-suspended",
4 => "Disabled",
5 => "RxDetect",
6 => "Inactive",
7 => "Polling",
8 => "Recovery",
9 => "HotReset",
10 => "Compliance",
11 => "Test",
15 => "Resume",
else => "reserved",
};
}
/// The most useful readable name for an interface's (class, subclass, protocol)
/// triple, decoding the well-known combinations recognizable in a log — e.g.
/// "HID boot keyboard", "Mass Storage SCSI Bulk-Only", "Bluetooth". Falls back
/// to the class name (and then "Unknown") for codes without a spelled-out combo.
pub fn interfaceName(class: u8, subclass: u8, protocol: u8) []const u8 {
return switch (@as(Class, @enumFromInt(class))) {
.hid => if (subclass == @intFromEnum(hid.SubClass.boot)) switch (@as(hid.Protocol, @enumFromInt(protocol))) {
.keyboard => "HID boot keyboard",
.mouse => "HID boot mouse",
else => "HID boot device",
} else "HID",
.mass_storage => switch (@as(mass_storage.Protocol, @enumFromInt(protocol))) {
.bulk_only => "Mass Storage (Bulk-Only)",
.uas => "Mass Storage (UAS)",
else => "Mass Storage",
},
.hub => switch (@as(hub.Protocol, @enumFromInt(protocol))) {
.super_speed => "Hub (SuperSpeed)",
.hi_speed_multi_tt => "Hub (Hi-Speed multi-TT)",
.hi_speed_single_tt => "Hub (Hi-Speed single-TT)",
else => "Hub",
},
.wireless_controller => if (subclass == @intFromEnum(wireless_controller.SubClass.radio_frequency))
wireless_controller.protocolName(protocol)
else
"Wireless Controller",
.communications => communications.subclassName(subclass),
.application_specific => application_specific.subclassName(subclass),
.miscellaneous => "Miscellaneous",
else => className(class),
};
}
// Subclass and protocol codes qualified by Class.hub. Hubs have no subclass codes; the // Subclass and protocol codes qualified by Class.hub. Hubs have no subclass codes; the
// protocol distinguishes the hub's transaction-translator arrangement. // protocol distinguishes the hub's transaction-translator arrangement.
pub const hub = struct { pub const hub = struct {
@@ -84,6 +182,16 @@ pub const hub = struct {
super_speed = 0x03, super_speed = 0x03,
_, _,
}; };
pub fn protocolName(protocol: u8) []const u8 {
return switch (@as(Protocol, @enumFromInt(protocol))) {
.full_speed => "full-speed",
.hi_speed_single_tt => "Hi-Speed single-TT",
.hi_speed_multi_tt => "Hi-Speed multi-TT",
.super_speed => "SuperSpeed",
_ => "unknown",
};
}
}; };
// Subclass and protocol codes qualified by Class.hid. // Subclass and protocol codes qualified by Class.hid.
@@ -104,6 +212,23 @@ pub const hid = struct {
mouse = 0x02, mouse = 0x02,
_, _,
}; };
pub fn subclassName(subclass: u8) []const u8 {
return switch (@as(SubClass, @enumFromInt(subclass))) {
.none => "none",
.boot => "boot",
_ => "unknown",
};
}
pub fn protocolName(protocol: u8) []const u8 {
return switch (@as(Protocol, @enumFromInt(protocol))) {
.none => "none",
.keyboard => "keyboard",
.mouse => "mouse",
_ => "unknown",
};
}
}; };
// Subclass and protocol codes qualified by Class.mass_storage. The subclass identifies the // Subclass and protocol codes qualified by Class.mass_storage. The subclass identifies the
@@ -147,6 +272,33 @@ pub const mass_storage = struct {
vendor_specific = 0xFF, vendor_specific = 0xFF,
_, _,
}; };
pub fn subclassName(subclass: u8) []const u8 {
return switch (@as(SubClass, @enumFromInt(subclass))) {
.not_reported => "SCSI (not reported)",
.rbc => "RBC",
.atapi => "ATAPI",
.qic_157 => "QIC-157",
.ufi => "UFI",
.sff_8070i => "SFF-8070i",
.scsi => "SCSI",
.lsd_fs => "LSD FS",
.ieee_1667 => "IEEE 1667",
.vendor_specific => "vendor-specific",
_ => "unknown",
};
}
pub fn protocolName(protocol: u8) []const u8 {
return switch (@as(Protocol, @enumFromInt(protocol))) {
.cbi_completion_interrupt => "CBI",
.cbi => "CBI (no completion IRQ)",
.bulk_only => "Bulk-Only",
.uas => "UAS",
.vendor_specific => "vendor-specific",
_ => "unknown",
};
}
}; };
// Subclass and protocol codes qualified by Class.communications (CDC). The protocol codes // Subclass and protocol codes qualified by Class.communications (CDC). The protocol codes
@@ -182,6 +334,25 @@ pub const communications = struct {
network_control = 0x0D, network_control = 0x0D,
_, _,
}; };
pub fn subclassName(subclass: u8) []const u8 {
return switch (@as(SubClass, @enumFromInt(subclass))) {
.direct_line => "Direct Line",
.abstract_control => "Abstract Control (modem/serial)",
.telephone => "Telephone",
.multi_channel => "Multi-Channel",
.capi => "CAPI",
.ethernet => "Ethernet",
.atm => "ATM",
.wireless_handset => "Wireless Handset",
.device_management => "Device Management",
.mobile_direct_line => "Mobile Direct Line",
.obex => "OBEX",
.ethernet_emulation => "Ethernet Emulation",
.network_control => "Network Control",
_ => "unknown",
};
}
}; };
// Subclass and protocol codes qualified by Class.wireless_controller. // Subclass and protocol codes qualified by Class.wireless_controller.
@@ -204,6 +375,23 @@ pub const wireless_controller = struct {
bluetooth_amp = 0x04, bluetooth_amp = 0x04,
_, _,
}; };
pub fn subclassName(subclass: u8) []const u8 {
return switch (@as(SubClass, @enumFromInt(subclass))) {
.radio_frequency => "RF",
_ => "unknown",
};
}
pub fn protocolName(protocol: u8) []const u8 {
return switch (@as(Protocol, @enumFromInt(protocol))) {
.bluetooth => "Bluetooth",
.ultra_wideband => "Ultra-Wideband",
.remote_ndis => "Remote NDIS",
.bluetooth_amp => "Bluetooth AMP",
_ => "unknown",
};
}
}; };
// Subclass and protocol codes qualified by Class.miscellaneous. // Subclass and protocol codes qualified by Class.miscellaneous.
@@ -221,6 +409,20 @@ pub const miscellaneous = struct {
interface_association = 0x01, interface_association = 0x01,
_, _,
}; };
pub fn subclassName(subclass: u8) []const u8 {
return switch (@as(SubClass, @enumFromInt(subclass))) {
.common => "common",
_ => "unknown",
};
}
pub fn protocolName(protocol: u8) []const u8 {
return switch (@as(Protocol, @enumFromInt(protocol))) {
.interface_association => "Interface Association",
_ => "unknown",
};
}
}; };
// Subclass and protocol codes qualified by Class.application_specific. // Subclass and protocol codes qualified by Class.application_specific.
@@ -234,6 +436,15 @@ pub const application_specific = struct {
test_and_measurement = 0x03, test_and_measurement = 0x03,
_, _,
}; };
pub fn subclassName(subclass: u8) []const u8 {
return switch (@as(SubClass, @enumFromInt(subclass))) {
.firmware_upgrade => "Device Firmware Upgrade",
.irda_bridge => "IrDA Bridge",
.test_and_measurement => "Test & Measurement",
_ => "unknown",
};
}
}; };
/// Pack a (class, subclass, protocol) triple into one 0xCCSSPP value — the /// Pack a (class, subclass, protocol) triple into one 0xCCSSPP value — the
@@ -280,6 +491,30 @@ test "class codes match the USB-IF assignments" {
_ = application_specific.SubClass.firmware_upgrade; _ = application_specific.SubClass.firmware_upgrade;
} }
test "readable names decode the well-known triples" {
const std = @import("std");
const eql = std.testing.expectEqualStrings;
try eql("Hub", className(0x09));
try eql("Unknown", className(0x42));
// interfaceName decodes the combos we log.
try eql("HID boot keyboard", interfaceName(0x03, 0x01, 0x01));
try eql("HID boot mouse", interfaceName(0x03, 0x01, 0x02));
try eql("Mass Storage (Bulk-Only)", interfaceName(0x08, 0x06, 0x50));
try eql("Hub (SuperSpeed)", interfaceName(0x09, 0x00, 0x03));
try eql("Bluetooth", interfaceName(0xE0, 0x01, 0x01));
// The per-enum name functions.
try eql("Bulk-Only", mass_storage.protocolName(0x50));
try eql("SCSI", mass_storage.subclassName(0x06));
try eql("Bluetooth", wireless_controller.protocolName(0x01));
try eql("SuperSpeed", hub.protocolName(0x03));
try eql("keyboard", hid.protocolName(0x01));
try eql("SuperSpeed", speedName(4));
try eql("Polling", linkStateName(7));
}
test "packTriple / unpackTriple round-trip the identity a bus driver reports" { test "packTriple / unpackTriple round-trip the identity a bus driver reports" {
const std = @import("std"); const std = @import("std");
const expectEqual = std.testing.expectEqual; const expectEqual = std.testing.expectEqual;
+6 -11
View File
@@ -16,11 +16,6 @@ const protocol = runtime.device_manager_protocol;
const device = runtime.device; const device = runtime.device;
const pci_class = @import("pci-class"); const pci_class = @import("pci-class");
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Log a discovered function with its (class / subclass / prog-IF) triple decoded /// Log a discovered function with its (class / subclass / prog-IF) triple decoded
/// to human names — the boot-log breadcrumb that says *what* the hardware is, so /// to human names — the boot-log breadcrumb that says *what* the hardware is, so
/// "class 0x01 (Mass Storage Controller) subclass 0x06 (Serial ATA Controller) /// "class 0x01 (Mass Storage Controller) subclass 0x06 (Serial ATA Controller)
@@ -74,7 +69,7 @@ fn configWrite16(bus: u64, dev: u64, function: u64, offset: u64, value: u16) voi
fn initialise(endpoint: runtime.ipc.Handle) bool { fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint; _ = endpoint;
if (!device.claim(bridge_id)) { if (!device.claim(bridge_id)) {
writeLine("/system/drivers/pci-bus: unable to claim bridge device {d}\n", .{bridge_id}); std.log.info("unable to claim bridge device {d}", .{bridge_id});
return false; return false;
} }
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
@@ -85,7 +80,7 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| { const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == bridge_id) break d; if (d.id == bridge_id) break d;
} else { } else {
writeLine("/system/drivers/pci-bus: device {d} not in the device tree\n", .{bridge_id}); std.log.info("device {d} not in the device tree", .{bridge_id});
return false; return false;
}; };
// Resource 0 is the ECAM window (1 MiB of config space per bus); the bus // Resource 0 is the ECAM window (1 MiB of config space per bus); the bus
@@ -159,7 +154,7 @@ fn scan() void {
} }
} }
} }
writeLine("/system/drivers/pci-bus: {d} functions found\n", .{found}); std.log.info("{d} functions found", .{found});
} }
/// Register one function under the bridge and report it to the manager. The /// Register one function under the bridge and report it to the manager. The
@@ -229,7 +224,7 @@ fn registerAndReport(bus: u64, dev: u64, function: u64, class_triple: u32) void
} }
const registered = device.register(bridge_id, &descriptor) orelse { const registered = device.register(bridge_id, &descriptor) orelse {
writeLine("/system/drivers/pci-bus: register refused for {d}:{d}.{d}\n", .{ bus, dev, function }); std.log.info("register refused for {d}:{d}.{d}", .{ bus, dev, function });
return; return;
}; };
const report = protocol.ChildAdded{ const report = protocol.ChildAdded{
@@ -240,7 +235,7 @@ fn registerAndReport(bus: u64, dev: u64, function: u64, class_triple: u32) void
}; };
var reply: [protocol.message_maximum]u8 = undefined; var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager_handle, std.mem.asBytes(&report), &reply) catch { _ = runtime.ipc.call(manager_handle, std.mem.asBytes(&report), &reply) catch {
writeLine("/system/drivers/pci-bus: child report for {d}:{d}.{d} failed\n", .{ bus, dev, function }); std.log.info("child report for {d}:{d}.{d} failed", .{ bus, dev, function });
}; };
} }
@@ -255,7 +250,7 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse return; // bare (ramdisk sweep): stay silent const argument = init.arguments.get(1) orelse return; // bare (ramdisk sweep): stay silent
bridge_id = std.fmt.parseInt(u64, argument, 10) catch { bridge_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/pci-bus: malformed bridge device id '{s}'\n", .{argument}); std.log.info("malformed bridge device id '{s}'", .{argument});
return; return;
}; };
runtime.service.run(protocol.message_maximum, .{ runtime.service.run(protocol.message_maximum, .{
+3 -8
View File
@@ -23,11 +23,6 @@ const device = runtime.device;
const ipc = runtime.ipc; const ipc = runtime.ipc;
const protocol = runtime.input_protocol; const protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Look up the ps2-bus service, retrying while the bus (which spawned us before /// Look up the ps2-bus service, retrying while the bus (which spawned us before
/// registering) is still coming up. /// registering) is still coming up.
fn lookupBus() ?ipc.Handle { fn lookupBus() ?ipc.Handle {
@@ -75,14 +70,14 @@ pub fn main(init: runtime.process.Init) void {
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: no HID argument\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: no HID argument\n");
return; return;
} }
writeLine("/system/drivers/ps2-bus/keyboard: starting for hid {s}\n", .{hid}); std.log.info("starting for hid {s}", .{hid});
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: out of memory\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: out of memory\n");
return; return;
}; };
if (device.findDeviceDescriptorByHid(buffer, hid) == null) { if (device.findDeviceDescriptorByHid(buffer, hid) == null) {
writeLine("/system/drivers/ps2-bus/keyboard: no device for hid {s}\n", .{hid}); std.log.info("no device for hid {s}", .{hid});
return; return;
} }
@@ -90,7 +85,7 @@ pub fn main(init: runtime.process.Init) void {
// absent (as today) it defaults to us. // absent (as today) it defaults to us.
const layout_name = init.arguments.get(2) orelse "us"; const layout_name = init.arguments.get(2) orelse "us";
const layout = xkb.byName(layout_name) orelse xkb.us; const layout = xkb.byName(layout_name) orelse xkb.us;
writeLine("/system/drivers/ps2-bus/keyboard: layout {s}\n", .{layout.name}); std.log.info("layout {s}", .{layout.name});
// Attach to the bus: hand it our endpoint, and it forwards every byte the // Attach to the bus: hand it our endpoint, and it forwards every byte the
// keyboard sends (it owns the controller; we own the decoding). // keyboard sends (it owns the controller; we own the decoding).
+2 -7
View File
@@ -22,11 +22,6 @@ const device = runtime.device;
const ipc = runtime.ipc; const ipc = runtime.ipc;
const protocol = runtime.input_protocol; const protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Look up the ps2-bus service, retrying while the bus (which spawned us before /// Look up the ps2-bus service, retrying while the bus (which spawned us before
/// registering) is still coming up. /// registering) is still coming up.
fn lookupBus() ?ipc.Handle { fn lookupBus() ?ipc.Handle {
@@ -54,14 +49,14 @@ pub fn main(init: runtime.process.Init) void {
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: no HID argument\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: no HID argument\n");
return; return;
} }
writeLine("/system/drivers/ps2-bus/mouse: starting for hid {s}\n", .{hid}); std.log.info("starting for hid {s}", .{hid});
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: out of memory\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: out of memory\n");
return; return;
}; };
if (ps2.findMouseDescriptor(buffer) == null) { if (ps2.findMouseDescriptor(buffer) == null) {
writeLine("/system/drivers/ps2-bus/mouse: no device for hid {s}\n", .{hid}); std.log.info("no device for hid {s}", .{hid});
return; return;
} }
+5 -12
View File
@@ -16,13 +16,6 @@ const ps2 = @import("ps2-library.zig");
const device = runtime.device; const device = runtime.device;
const ipc = runtime.ipc; const ipc = runtime.ipc;
/// Format one whole log line and emit it in a single `debug_write`, so output
/// from the child drivers (which run concurrently) can never interleave with it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Ask the device on `port` what it is, then spawn the matching driver from the /// Ask the device on `port` what it is, then spawn the matching driver from the
/// initial-ramdisk, handing it the device's HID as argv[1]. The driver is chosen /// initial-ramdisk, handing it the device's HID as argv[1]. The driver is chosen
/// from what the device reports, not from the port number. Returns the identified /// from what the device reports, not from the port number. Returns the identified
@@ -30,19 +23,19 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
/// attaches, or null if nothing was spawned. /// attaches, or null if nothing was spawned.
fn spawnIdentifiedDriver(controller: ps2.Controller, port: ps2.Port) ?ps2.DeviceType { fn spawnIdentifiedDriver(controller: ps2.Controller, port: ps2.Port) ?ps2.DeviceType {
const device_type = controller.identifyDevice(port) orelse { const device_type = controller.identifyDevice(port) orelse {
writeLine("/system/drivers/ps2-bus: identify timed out on port {s}\n", .{@tagName(port)}); std.log.info("identify timed out on port {s}", .{@tagName(port)});
return null; return null;
}; };
const driver_name = device_type.driverName() orelse { const driver_name = device_type.driverName() orelse {
writeLine("/system/drivers/ps2-bus: unrecognized device on port {s}\n", .{@tagName(port)}); std.log.info("unrecognized device on port {s}", .{@tagName(port)});
return null; return null;
}; };
const hid = device_type.hid() orelse ""; const hid = device_type.hid() orelse "";
if (runtime.system.spawnWithArguments(driver_name, &.{hid}) != null) { if (runtime.system.spawnWithArguments(driver_name, &.{hid}) != null) {
writeLine("/system/drivers/ps2-bus: port {s} is a {s}, spawned {s}\n", .{ @tagName(port), hid, driver_name }); std.log.info("port {s} is a {s}, spawned {s}", .{ @tagName(port), hid, driver_name });
return device_type; return device_type;
} }
writeLine("/system/drivers/ps2-bus: failed to spawn {s}\n", .{driver_name}); std.log.info("failed to spawn {s}", .{driver_name});
return null; return null;
} }
@@ -83,7 +76,7 @@ fn handleAttach(message: []const u8, got: ipc.Received, out: []u8) usize {
const device_type = maybe_type orelse continue; const device_type = maybe_type orelse continue;
if (@intFromEnum(device_type) != request.device_type) continue; if (@intFromEnum(device_type) != request.device_type) continue;
port_endpoints[port_index] = endpoint; port_endpoints[port_index] = endpoint;
writeLine("/system/drivers/ps2-bus: {s} driver attached\n", .{@tagName(device_type)}); std.log.info("{s} driver attached", .{@tagName(device_type)});
return reply.write(out, .ok); return reply.write(out, .ok);
} }
return reply.write(out, .no_such_device); return reply.write(out, .no_such_device);
+12 -8
View File
@@ -23,11 +23,6 @@ const ipc = runtime.ipc;
const process = runtime.process; const process = runtime.process;
const input_protocol = runtime.input_protocol; const input_protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// The modifier state a character lookup needs — derived from the report's // The modifier state a character lookup needs — derived from the report's
// modifier byte, plus the driver-tracked caps-lock toggle. // modifier byte, plus the driver-tracked caps-lock toggle.
const ModifierSnapshot = struct { const ModifierSnapshot = struct {
@@ -71,7 +66,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
}; };
const device_id = std.fmt.parseInt(u64, argument, 10) catch { const device_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/usb-hid/keyboard: malformed device id '{s}'\n", .{argument}); std.log.info("malformed device id '{s}'", .{argument});
return; return;
}; };
const layout = xkb.byName(init.arguments.get(2) orelse "us") orelse xkb.us; const layout = xkb.byName(init.arguments.get(2) orelse "us") orelse xkb.us;
@@ -82,7 +77,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
} }
var device = runtime.usb.open(device_id) orelse { var device = runtime.usb.open(device_id) orelse {
writeLine("/system/drivers/usb-hid/keyboard: could not open device {d}\n", .{device_id}); std.log.info("could not open device {d}", .{device_id});
return; return;
}; };
const endpoint = device.findEndpoint(runtime.usb.transfer_type_interrupt, true) orelse { const endpoint = device.findEndpoint(runtime.usb.transfer_type_interrupt, true) orelse {
@@ -104,7 +99,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
}; };
_ = process.bindSignals(device.endpoint); _ = process.bindSignals(device.endpoint);
writeLine("/system/drivers/usb-hid/keyboard: ok (device {d}, interface {d}, layout {s})\n", .{ device_id, device.interface_number, layout.name }); std.log.info("ok (device {d}, interface {d}, layout {s})", .{ device_id, device.interface_number, layout.name });
var decoder = hid.KeyboardDecoder{}; var decoder = hid.KeyboardDecoder{};
var caps_lock = false; var caps_lock = false;
@@ -153,6 +148,15 @@ pub fn main(init: runtime.process.Init) void {
.character = character, .character = character,
.modifiers = modifier_word, .modifiers = modifier_word,
}); });
// Echo the character to the log — a simple end-to-end
// keyboard check on real hardware: type a known phrase,
// then read it back from usb-hid-keyboard.log (or watch
// it appear live on screen in a -Ddiagnose boot, where
// the kernel console is a log sink). Printable ASCII and
// newline only; other keys are left to the input service.
if (character == '\n' or (character >= 0x20 and character < 0x7F)) {
_ = runtime.system.write(&[1]u8{@intCast(character)});
}
} }
}, },
.released => { .released => {
+3 -8
View File
@@ -18,11 +18,6 @@ const ipc = runtime.ipc;
const process = runtime.process; const process = runtime.process;
const input_protocol = runtime.input_protocol; const input_protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// The current pressed-button bitmask in input-protocol terms. // The current pressed-button bitmask in input-protocol terms.
fn buttonMask(buttons: u8) u32 { fn buttonMask(buttons: u8) u32 {
var mask: u32 = 0; var mask: u32 = 0;
@@ -38,7 +33,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
}; };
const device_id = std.fmt.parseInt(u64, argument, 10) catch { const device_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/usb-hid/mouse: malformed device id '{s}'\n", .{argument}); std.log.info("malformed device id '{s}'", .{argument});
return; return;
}; };
@@ -47,7 +42,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
} }
var device = runtime.usb.open(device_id) orelse { var device = runtime.usb.open(device_id) orelse {
writeLine("/system/drivers/usb-hid/mouse: could not open device {d}\n", .{device_id}); std.log.info("could not open device {d}", .{device_id});
return; return;
}; };
const endpoint = device.findEndpoint(runtime.usb.transfer_type_interrupt, true) orelse { const endpoint = device.findEndpoint(runtime.usb.transfer_type_interrupt, true) orelse {
@@ -67,7 +62,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
}; };
_ = process.bindSignals(device.endpoint); _ = process.bindSignals(device.endpoint);
writeLine("/system/drivers/usb-hid/mouse: ok (device {d}, interface {d})\n", .{ device_id, device.interface_number }); std.log.info("ok (device {d}, interface {d})", .{ device_id, device.interface_number });
var previous_buttons: u8 = 0; var previous_buttons: u8 = 0;
var receive: [64]u8 = undefined; var receive: [64]u8 = undefined;
+16 -9
View File
@@ -18,11 +18,6 @@ const bot = @import("bulk-only-transport.zig");
const block_protocol = @import("block-protocol"); const block_protocol = @import("block-protocol");
const dma = runtime.dma; const dma = runtime.dma;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
var device_id: u64 = 0; var device_id: u64 = 0;
var device: runtime.usb.Device = undefined; var device: runtime.usb.Device = undefined;
var bulk_in: runtime.usb.Endpoint = undefined; var bulk_in: runtime.usb.Endpoint = undefined;
@@ -66,6 +61,12 @@ fn transact(cdb: []const u8, direction_in: bool, data_physical: u64, data_length
return status.status == @intFromEnum(bot.CommandStatus.passed); return status.status == @intFromEnum(bot.CommandStatus.passed);
} }
/// Set when bring-up failed with the device PRESENT (an opened device that then
/// failed a step): main exits nonzero, and the device manager restarts us with
/// backoff — a transient failure heals instead of leaving storage down forever.
/// Device-absent paths stay clean exits: nothing to serve, nothing to retry.
var bring_up_failed = false;
fn initialise(endpoint: runtime.ipc.Handle) bool { fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint; _ = endpoint;
if (!runtime.usb.helloManager(device_id)) { if (!runtime.usb.helloManager(device_id)) {
@@ -73,15 +74,17 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
return false; return false;
} }
device = runtime.usb.open(device_id) orelse { device = runtime.usb.open(device_id) orelse {
writeLine("/system/drivers/usb-storage: could not open device {d}\n", .{device_id}); std.log.info("could not open device {d}", .{device_id});
return false; return false;
}; };
bulk_in = device.findEndpoint(runtime.usb.transfer_type_bulk, true) orelse { bulk_in = device.findEndpoint(runtime.usb.transfer_type_bulk, true) orelse {
_ = runtime.system.write("/system/drivers/usb-storage: no bulk-IN endpoint\n"); _ = runtime.system.write("/system/drivers/usb-storage: no bulk-IN endpoint\n");
bring_up_failed = true;
return false; return false;
}; };
bulk_out = device.findEndpoint(runtime.usb.transfer_type_bulk, false) orelse { bulk_out = device.findEndpoint(runtime.usb.transfer_type_bulk, false) orelse {
_ = runtime.system.write("/system/drivers/usb-storage: no bulk-OUT endpoint\n"); _ = runtime.system.write("/system/drivers/usb-storage: no bulk-OUT endpoint\n");
bring_up_failed = true;
return false; return false;
}; };
command_wrapper = dma.alloc(4096, dma.coherent) orelse return false; command_wrapper = dma.alloc(4096, dma.coherent) orelse return false;
@@ -104,6 +107,7 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
const capacity_command = scsi.readCapacity10(); const capacity_command = scsi.readCapacity10();
if (!transact(&capacity_command, true, command_data.physical, 8)) { if (!transact(&capacity_command, true, command_data.physical, 8)) {
_ = runtime.system.write("/system/drivers/usb-storage: READ CAPACITY failed\n"); _ = runtime.system.write("/system/drivers/usb-storage: READ CAPACITY failed\n");
bring_up_failed = true;
return false; return false;
} }
var capacity_bytes: [8]u8 = undefined; var capacity_bytes: [8]u8 = undefined;
@@ -112,14 +116,14 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
const capacity = scsi.parseCapacity(capacity_bytes); const capacity = scsi.parseCapacity(capacity_bytes);
block_size = capacity.block_size; block_size = capacity.block_size;
block_count = @as(u64, capacity.last_lba) + 1; block_count = @as(u64, capacity.last_lba) + 1;
writeLine("/system/drivers/usb-storage: ready ({d} blocks x {d} bytes)\n", .{ block_count, block_size }); std.log.info("ready ({d} blocks x {d} bytes)", .{ block_count, block_size });
// Self-check: read block 0 and log its trailing signature (0x55AA for a boot // Self-check: read block 0 and log its trailing signature (0x55AA for a boot
// sector) — proof READ(10) works end to end over the bulk path. // sector) — proof READ(10) works end to end over the bulk path.
const read0 = scsi.read10(0, 1); const read0 = scsi.read10(0, 1);
if (block_size <= 4096 and transact(&read0, true, command_data.physical, block_size)) { if (block_size <= 4096 and transact(&read0, true, command_data.physical, block_size)) {
const sector: [*]const u8 = @ptrFromInt(command_data.virtual); const sector: [*]const u8 = @ptrFromInt(command_data.virtual);
writeLine("/system/drivers/usb-storage: block 0 signature 0x{x:0>2}{x:0>2}\n", .{ sector[510], sector[511] }); std.log.info("block 0 signature 0x{x:0>2}{x:0>2}", .{ sector[510], sector[511] });
} }
return true; return true;
} }
@@ -171,7 +175,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
}; };
device_id = std.fmt.parseInt(u64, argument, 10) catch { device_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/usb-storage: malformed device id '{s}'\n", .{argument}); std.log.info("malformed device id '{s}'", .{argument});
return; return;
}; };
runtime.service.run(block_protocol.message_maximum, .{ runtime.service.run(block_protocol.message_maximum, .{
@@ -179,4 +183,7 @@ pub fn main(init: runtime.process.Init) void {
.init = initialise, .init = initialise,
.on_message = onMessage, .on_message = onMessage,
}); });
// A failure exit (nonzero -> .aborted) tells the device manager to restart
// us with backoff; a clean return means there was nothing to serve.
if (bring_up_failed) runtime.system.exit(1);
} }
+216 -42
View File
@@ -64,13 +64,6 @@ fn reportEndpointFor(device_token: u64) ?usize {
return null; return null;
} }
/// Format one whole log line and emit it in a single `debug_write`, so
/// concurrent instances (one per controller) can never interleave mid-line.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
var controller_id: u64 = protocol.no_device; var controller_id: u64 = protocol.no_device;
/// Claim the assigned controller, find its register window, and hello the /// Claim the assigned controller, find its register window, and hello the
@@ -79,7 +72,7 @@ var controller_id: u64 = protocol.no_device;
fn initialise(endpoint: runtime.ipc.Handle) bool { fn initialise(endpoint: runtime.ipc.Handle) bool {
service_endpoint = endpoint; service_endpoint = endpoint;
if (!device.claim(controller_id)) { if (!device.claim(controller_id)) {
writeLine("/system/drivers/usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id}); std.log.info("unable to claim controller device {d}", .{controller_id});
return false; return false;
} }
@@ -92,7 +85,7 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| { const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == controller_id) break d; if (d.id == controller_id) break d;
} else { } else {
writeLine("/system/drivers/usb-xhci-bus: device {d} not in the device tree\n", .{controller_id}); std.log.info("device {d} not in the device tree", .{controller_id});
return false; return false;
}; };
@@ -105,10 +98,10 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
break resource; break resource;
} }
} else { } else {
writeLine("/system/drivers/usb-xhci-bus: controller device {d} has no register BAR\n", .{controller_id}); std.log.info("controller device {d} has no register BAR", .{controller_id});
return false; return false;
}; };
writeLine("/system/drivers/usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{ std.log.info("claimed controller device {d} (registers at 0x{x}, {d} bytes)", .{
controller_id, controller_id,
register_window.start, register_window.start,
register_window.len, register_window.len,
@@ -124,7 +117,7 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: controller reset/bring-up failed\n"); _ = runtime.system.write("/system/drivers/usb-xhci-bus: controller reset/bring-up failed\n");
return false; return false;
}; };
writeLine("/system/drivers/usb-xhci-bus: controller running ({d} slots, {d}-byte contexts)\n", .{ std.log.info("controller running ({d} slots, {d}-byte contexts)", .{
controller.?.max_slots, controller.?.max_slots,
controller.?.context_size, controller.?.context_size,
}); });
@@ -150,6 +143,7 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: no device manager to hello\n"); _ = runtime.system.write("/system/drivers/usb-xhci-bus: no device manager to hello\n");
return false; return false;
}; };
manager_handle = h; // the tick's hot-plug dispatch reports through this
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.bus), .device_id = controller_id }; const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.bus), .device_id = controller_id };
var reply: [protocol.message_maximum]u8 = undefined; var reply: [protocol.message_maximum]u8 = undefined;
const n = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch { const n = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch {
@@ -193,46 +187,179 @@ fn speedName(speed: u32) []const u8 {
/// report one child per interface — carrying the interface's (class, subclass, /// report one child per interface — carrying the interface's (class, subclass,
/// protocol) triple as identity, which is what the device manager matches a /// protocol) triple as identity, which is what the device manager matches a
/// class driver against. /// class driver against.
var manager_handle: ?runtime.ipc.Handle = null;
// Per-root-port connected state from the previous tick, so the poll acts on
// empty->connected transitions (edge), never re-attempting a level every tick.
var prev_connected: [64]bool = [_]bool{false} ** 64;
fn scanPorts(manager: runtime.ipc.Handle) void { fn scanPorts(manager: runtime.ipc.Handle) void {
const engine = if (controller) |*c| c else { const engine = if (controller) |*c| c else {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: controller not initialised\n"); _ = runtime.system.write("/system/drivers/usb-xhci-bus: controller not initialised\n");
return; return;
}; };
writeLine("/system/drivers/usb-xhci-bus: {d} root-hub ports\n", .{engine.max_ports}); std.log.info("{d} root-hub ports", .{engine.max_ports});
var port: u32 = 1; var port: u32 = 1;
var connected: u32 = 0; var connected: u32 = 0;
while (port <= engine.max_ports) : (port += 1) { while (port <= engine.max_ports) : (port += 1) {
const port_status = engine.portStatus(port); if (!engine.portConnected(port)) continue;
if (port_status & 1 == 0) continue; // CCS: nothing connected if (port < prev_connected.len) prev_connected[port] = true; // don't re-fire the poll for these
connected += 1; connected += 1;
const speed = (port_status >> 10) & 0xF; // the PORTSC port-speed class bringUpPort(manager, engine, port);
writeLine("/system/drivers/usb-xhci-bus: port {d} connected — {s} (speed class {d})\n", .{ port, speedName(speed), speed }); }
if (connected == 0) {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: no devices connected\n");
engine.dumpPortTopology(); // help diagnose an empty scan: the xECP map + raw PORTSC
}
}
const usb_device = engine.setupDevice(port, speed) orelse { /// Bring up whatever is on `port`: setup + enumerate + register/report one child
writeLine("/system/drivers/usb-xhci-bus: port {d} device setup failed\n", .{port}); /// per interface. Shared by the boot scan and hot-plug (a port-change event with
continue; /// the port now connected).
}; fn bringUpPort(manager: runtime.ipc.Handle, engine: *library.Controller, port: u32) void {
if (!engine.enumerate(usb_device)) { const speed = (engine.portStatus(port) >> 10) & 0xF; // the PORTSC port-speed class
writeLine("/system/drivers/usb-xhci-bus: port {d} enumeration failed\n", .{port}); std.log.info("port {d} connected — {s} (speed class {d})", .{ port, speedName(speed), speed });
continue;
}
writeLine("/system/drivers/usb-xhci-bus: port {d} device vendor 0x{x:0>4} product 0x{x:0>4}, {d} interface(s)\n", .{
port,
usb_device.device_descriptor.vendor_id,
usb_device.device_descriptor.product_id,
usb_device.interface_count,
});
for (usb_device.interfaces[0..usb_device.interface_count]) |*interface| { const usb_device = engine.setupDevice(port, speed) orelse {
// Record the id each interface was registered as, so a class driver std.log.info("port {d} device setup failed", .{port});
// opening the interface (by that id) resolves to it. return;
if (reportInterface(manager, port, interface.*)) |registered| { };
interface.registered_device_id = registered; if (!engine.enumerate(usb_device)) {
} std.log.info("port {d} enumeration failed", .{port});
return;
}
var maker_buffer: [64]u8 = undefined;
var product_buffer: [64]u8 = undefined;
const maker = engine.readString(usb_device, @intFromEnum(usb_device.device_descriptor.manufacturer_index), &maker_buffer) orelse "?";
const product = engine.readString(usb_device, @intFromEnum(usb_device.device_descriptor.product_index), &product_buffer) orelse "?";
std.log.info("port {d} device: {s} \"{s} {s}\" (0x{x:0>4}:0x{x:0>4}), {d} interface(s)", .{
port,
usb_ids.className(usb_device.device_descriptor.device_class),
maker,
product,
usb_device.device_descriptor.vendor_id,
usb_device.device_descriptor.product_id,
usb_device.interface_count,
});
for (usb_device.interfaces[0..usb_device.interface_count]) |*interface| {
// Record the id each interface was registered as, so a class driver
// opening the interface (by that id) resolves to it.
if (reportInterface(manager, port, interface.*)) |registered| {
interface.registered_device_id = registered;
} }
} }
if (connected == 0) _ = runtime.system.write("/system/drivers/usb-xhci-bus: no devices connected\n");
// A hub (class 9) is bus infrastructure the bus drives itself: configure it
// and power its downstream ports (docs/usb-hub.md). Its interfaces are still
// reported above, but no external class driver binds it.
if (deviceIsHub(usb_device)) _ = engine.setupHub(usb_device);
}
// A compact topology-unique port key for a hub downstream port: 1000 + slot*100
// + port. Stays a few digits (the id tag "P<key>I<iface>" has an 8-byte cap)
// while never colliding with a root port (1..N) or another (hub, port).
fn hubPortKey(hub_slot: u8, port: u16) u32 {
return 1000 + @as(u32, hub_slot) * 100 + port;
}
/// Service a change on hub downstream `port`: dispatch a connect (enumerate the
/// new device) or a disconnect (tear the old one down). Recurses for a hub
/// behind a hub — a nested hub is set up on connect and its downstream devices
/// torn down first on disconnect.
fn bringUpBehindHub(manager: runtime.ipc.Handle, engine: *library.Controller, hub: *library.Device, port: u16) void {
const status = engine.hubPortStatusAck(hub, port) orelse return;
const connected = library.Controller.hubPortConnected(status);
const existing = engine.deviceOnHubPort(hub, port);
if (connected != (existing != null)) // only when a device appears or leaves — not empty seed-sweep ports
std.log.info("hub slot {d} port {d}: {s} (status 0x{x:0>4})", .{ hub.slot_id, port, if (connected) "device connected" else "device removed", status & 0xFFFF });
if (!connected) {
if (existing) |dev| tearDownHubDevice(manager, engine, dev);
return;
}
if (existing != null) return; // already up
const usb_device = engine.serviceHubPort(hub, port) orelse return;
if (!engine.enumerate(usb_device)) {
std.log.info("hub slot {d} port {d}: enumeration failed", .{ hub.slot_id, port });
return;
}
var maker_buffer: [64]u8 = undefined;
var product_buffer: [64]u8 = undefined;
const maker = engine.readString(usb_device, @intFromEnum(usb_device.device_descriptor.manufacturer_index), &maker_buffer) orelse "?";
const product = engine.readString(usb_device, @intFromEnum(usb_device.device_descriptor.product_index), &product_buffer) orelse "?";
std.log.info("hub slot {d} port {d} device: {s} \"{s} {s}\" (0x{x:0>4}:0x{x:0>4}), {d} interface(s)", .{
hub.slot_id, port,
usb_ids.className(usb_device.device_descriptor.device_class),
maker,
product,
usb_device.device_descriptor.vendor_id,
usb_device.device_descriptor.product_id,
usb_device.interface_count,
});
for (usb_device.interfaces[0..usb_device.interface_count]) |*interface| {
if (reportInterface(manager, hubPortKey(hub.slot_id, port), interface.*)) |registered| {
interface.registered_device_id = registered;
}
}
if (deviceIsHub(usb_device)) _ = engine.setupHub(usb_device);
}
/// Tear down a device that disconnected from a hub: recursively tear down its
/// own downstream devices first if it is a hub, report each interface removed,
/// then Disable Slot. Mirrors tearDownPort for a hub-attached device.
fn tearDownHubDevice(manager: runtime.ipc.Handle, engine: *library.Controller, dev: *library.Device) void {
// A hub that left takes its whole subtree with it — tear children down first.
if (dev.is_hub) {
while (engine.nextChildOf(dev.slot_id, 0)) |child| tearDownHubDevice(manager, engine, child);
}
std.log.info("hub device slot {d} disconnected", .{dev.slot_id});
const key = hubPortKey(dev.parent_slot, dev.parent_port);
for (dev.interfaces[0..dev.interface_count]) |*interface| {
if (interface.registered_device_id == 0) continue;
const event = protocol.ChildRemoved{
.parent = controller_id,
.bus_address = (@as(u64, key) << 8) | interface.number,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&event), &reply) catch {};
interface.registered_device_id = 0;
}
engine.tearDownDevice(dev);
}
/// Whether an enumerated device is a hub — class 9 at the device or the
/// interface level (a hub's single interface is class 9/0/0).
fn deviceIsHub(usb_device: *const library.Device) bool {
if (usb_device.device_descriptor.device_class == @intFromEnum(usb_ids.Class.hub)) return true;
for (usb_device.interfaces[0..usb_device.interface_count]) |interface| {
if (interface.class == @intFromEnum(usb_ids.Class.hub)) return true;
}
return false;
}
/// Tear down whatever was on `port` after an unplug: report each registered
/// interface as removed (the manager prunes the node, notifies watchers, and
/// stops the class driver's world honestly), then release the controller-side
/// device state (Disable Slot).
fn tearDownPort(manager: runtime.ipc.Handle, engine: *library.Controller, port: u32) void {
const usb_device = engine.deviceOnPort(port) orelse return;
std.log.info("port {d} disconnected", .{port});
for (usb_device.interfaces[0..usb_device.interface_count]) |*interface| {
if (interface.registered_device_id == 0) continue;
const event = protocol.ChildRemoved{
.parent = controller_id,
.bus_address = (@as(u64, port) << 8) | interface.number,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&event), &reply) catch {
std.log.info("child-removed report for port {d} interface {d} failed", .{ port, interface.number });
};
interface.registered_device_id = 0;
}
engine.tearDownDevice(usb_device);
} }
/// Register one interface as a resource-less child of the controller and report /// Register one interface as a resource-less child of the controller and report
@@ -259,7 +386,7 @@ fn reportInterface(manager: runtime.ipc.Handle, port: u32, interface: library.In
descriptor.hid_len = hid_text.len; descriptor.hid_len = hid_text.len;
@memcpy(descriptor.hid[0..hid_text.len], hid_text); @memcpy(descriptor.hid[0..hid_text.len], hid_text);
const registered = device.register(controller_id, &descriptor) orelse { const registered = device.register(controller_id, &descriptor) orelse {
writeLine("/system/drivers/usb-xhci-bus: register refused for port {d} interface {d}\n", .{ port, interface.number }); std.log.info("register refused for port {d} interface {d}", .{ port, interface.number });
return null; return null;
}; };
@@ -271,12 +398,13 @@ fn reportInterface(manager: runtime.ipc.Handle, port: u32, interface: library.In
}; };
var reply: [protocol.message_maximum]u8 = undefined; var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&report), &reply) catch { _ = runtime.ipc.call(manager, std.mem.asBytes(&report), &reply) catch {
writeLine("/system/drivers/usb-xhci-bus: child report for port {d} interface {d} failed\n", .{ port, interface.number }); std.log.info("child report for port {d} interface {d} failed", .{ port, interface.number });
return null; return null;
}; };
writeLine("/system/drivers/usb-xhci-bus: port {d} interface {d} class {d}/{d}/{d} registered as device {d}\n", .{ std.log.info("port {d} interface {d}: {s} ({d}/{d}/{d}) registered as device {d}", .{
port, port,
interface.number, interface.number,
usb_ids.interfaceName(interface.class, interface.subclass, interface.protocol),
interface.class, interface.class,
interface.subclass, interface.subclass,
interface.protocol, interface.protocol,
@@ -387,6 +515,52 @@ fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_timer_bit == 0) return; if (badge & runtime.ipc.notify_timer_bit == 0) return;
if (controller) |*engine| { if (controller) |*engine| {
engine.pump(); engine.pump();
// Poll every root port and reconcile — a device present but not yet
// enumerated is brought up; a device gone is torn down. This does NOT
// depend on a Port Status Change EVENT firing: the boot scan runs ~3 ms
// after the controller reset, far too early for a USB2 connection to
// debounce (~100 ms), and the SuperSpeed devices that DO show up early
// proved the event path unreliable for the late USB2 companion hub on
// real hardware. Polling catches it on the next tick regardless.
if (manager_handle) |manager| {
var port: u32 = 1;
while (port <= engine.max_ports and port <= prev_connected.len) : (port += 1) {
const connected = engine.portConnected(port);
const was = prev_connected[port];
prev_connected[port] = connected;
if (connected and !was and engine.deviceOnPort(port) == null) {
// Rising edge the boot scan missed (it ran before the USB2
// connection debounced): bring the device up now.
std.log.info("root port {d}: device appeared (PORTSC 0x{x:0>8})", .{ port, engine.portStatus(port) });
bringUpPort(manager, engine, port);
} else if (!connected and was and engine.deviceOnPort(port) != null) {
tearDownPort(manager, engine, port);
}
}
}
while (engine.takePortChange()) |port| {
const manager = manager_handle orelse break;
const connected = engine.portConnected(port);
std.log.info("root port {d} change: {s} (PORTSC 0x{x:0>8})", .{ port, if (connected) "connected" else "empty", engine.portStatus(port) });
if (connected) {
if (engine.deviceOnPort(port) == null) bringUpPort(manager, engine, port);
} else {
tearDownPort(manager, engine, port);
}
}
// Downstream hub-port changes (docs/usb-hub.md): a device connected on a
// hub's downstream port is enumerated and registered here, so a keyboard
// behind a hub reaches its class driver like one on a root port.
// Cap per tick: even if a hub's change bits refuse to clear, the driver
// must not spin here — it services a bounded batch and yields to the
// next tick (and to storage, input, everything else).
var serviced: u32 = 0;
while (engine.takeHubChange()) |change| {
const manager = manager_handle orelse break;
bringUpBehindHub(manager, engine, change.hub, change.port);
serviced += 1;
if (serviced >= 32) break;
}
while (engine.takeReport()) |report| { while (engine.takeReport()) |report| {
var message = transfer.InterruptReport{ var message = transfer.InterruptReport{
.device_token = report.device_token, .device_token = report.device_token,
@@ -407,7 +581,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
}; };
controller_id = std.fmt.parseInt(u64, argument, 10) catch { controller_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/usb-xhci-bus: malformed controller device id '{s}'\n", .{argument}); std.log.info("malformed controller device id '{s}'", .{argument});
return; return;
}; };
runtime.service.run(transfer.message_maximum, .{ runtime.service.run(transfer.message_maximum, .{
+673 -33
View File
@@ -22,6 +22,7 @@ const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const mmio = @import("mmio"); const mmio = @import("mmio");
const usb_abi = @import("usb-abi"); const usb_abi = @import("usb-abi");
const usb_ids = @import("usb-ids");
const dma = runtime.dma; const dma = runtime.dma;
const system = runtime.system; const system = runtime.system;
@@ -85,6 +86,7 @@ pub const TrbType = enum(u6) {
status_stage = 4, status_stage = 4,
link = 6, link = 6,
enable_slot = 9, enable_slot = 9,
disable_slot = 10,
address_device = 11, address_device = 11,
configure_endpoint = 12, configure_endpoint = 12,
evaluate_context = 13, evaluate_context = 13,
@@ -99,6 +101,7 @@ pub const TrbType = enum(u6) {
pub const CompletionCode = enum(u8) { pub const CompletionCode = enum(u8) {
invalid = 0, invalid = 0,
success = 1, success = 1,
usb_transaction_error = 4,
short_packet = 13, short_packet = 13,
_, _,
}; };
@@ -279,6 +282,21 @@ pub const Device = struct {
// Transfer rings configured for this device's interrupt/bulk endpoints. // Transfer rings configured for this device's interrupt/bulk endpoints.
endpoint_ring_count: u8 = 0, endpoint_ring_count: u8 = 0,
endpoint_rings: [max_configured_endpoints]ConfiguredEndpoint = [_]ConfiguredEndpoint{.{}} ** max_configured_endpoints, endpoint_rings: [max_configured_endpoints]ConfiguredEndpoint = [_]ConfiguredEndpoint{.{}} ** max_configured_endpoints,
// Hub topology (docs/usb-hub.md). A device behind a hub is addressed with a
// route string; these carry the fields buildAddressInputContext needs.
route: u32 = 0, // Slot Context route string (5 tiers x 4 bits); 0 = on a root port
root_port: u32 = 0, // the ROOT-hub port the chain hangs off (inherited down a chain)
parent_slot: u8 = 0, // the parent hub's slot id (0 = on a root port) — the TT hub
parent_port: u8 = 0, // the parent hub's downstream port this device sits on
// Set when this device IS a hub, after setupHub configures it.
is_hub: bool = false,
hub_ports: u8 = 0, // downstream port count from the hub descriptor
hub_multi_tt: bool = false,
// Downstream ports with a pending change to service (bit P = port P), set
// by the status-change endpoint (and by an initial sweep in setupHub).
hub_change_mask: u32 = 0,
}; };
// A standing interrupt-IN subscription: the endpoint's ring is kept armed with a // A standing interrupt-IN subscription: the endpoint's ring is kept armed with a
@@ -297,6 +315,9 @@ const Subscription = struct {
// device token and the endpoint handle its reports are sent to. // device token and the endpoint handle its reports are sent to.
device_token: u64 = 0, device_token: u64 = 0,
report_endpoint: usize = 0, report_endpoint: usize = 0,
// When set, this is an IN-PROCESS hub status-change subscription: completions
// set the hub's pending-change mask instead of queuing a class-driver report.
hub: ?*Device = null,
}; };
// One interrupt report waiting for the bus layer to push it to a subscriber. // One interrupt report waiting for the bus layer to push it to a subscriber.
@@ -314,6 +335,109 @@ const max_devices = 8;
const max_subscriptions = 8; const max_subscriptions = 8;
const report_queue_capacity = 16; const report_queue_capacity = 16;
// --- USB hub class requests + constants (docs/usb-hub.md) ------------------
//
// A hub is bus infrastructure the CONTROLLER driver handles in-process: the
// route strings and slot contexts a downstream device needs only exist here.
// These are the class-specific control requests to a hub device.
// A USB2 hub port's wPortStatus speed bits (bit 9 = low-speed, bit 10 =
// high-speed; neither = full-speed) mapped to the xHCI speed id.
fn mapHubPortSpeed(port_speed_bits: u32) u32 {
if (port_speed_bits & 0x1 != 0) return 2; // low-speed (wPortStatus bit 9)
if (port_speed_bits & 0x2 != 0) return 3; // high-speed
return 1; // full-speed
}
fn routeDepth(route: u32) u16 {
// Tiers used by a route string: each nonzero 4-bit nibble is one tier.
var depth: u16 = 0;
var r = route;
while (r != 0) : (r >>= 4) {
if (r & 0xF != 0) depth += 1;
}
return depth;
}
const hubreq = struct {
// Hub descriptor types (GET_DESCRIPTOR value high byte).
const descriptor_usb2: u8 = 0x29;
const descriptor_usb3: u8 = 0x2A;
// Hub/port feature selectors (SET_FEATURE / CLEAR_FEATURE value).
const feature_port_reset: u16 = 4;
const feature_port_power: u16 = 8;
const feature_c_port_connection: u16 = 16;
const feature_c_port_reset: u16 = 20;
const feature_c_port_link_state: u16 = 25; // SS
const feature_bh_port_reset: u16 = 28; // SS
const feature_c_bh_port_reset: u16 = 29; // SS
// Port status (wPortStatus, first 16 bits of the 4-byte GET_STATUS result).
const status_connection: u16 = 1 << 0;
const status_enable: u16 = 1 << 1;
const status_reset: u16 = 1 << 4;
// Port-status change bits (wPortChange, the high 16 bits).
const change_connection: u16 = 1 << 0;
const change_reset: u16 = 1 << 4;
// wHubCharacteristics bit 7: multiple transaction translators.
const characteristics_multi_tt: u16 = 1 << 7;
// SET_HUB_DEPTH (SuperSpeed hubs, so they can compose route strings).
const request_set_hub_depth: u8 = 12;
fn getDescriptor(kind: u8, length: u16) usb_abi.Request {
return .{
.request_type = .{ .recipient = .device, .kind = .class, .direction = .device_to_host },
.request_code = .get_descriptor,
.value = @as(u16, kind) << 8,
.index = 0,
.length = length,
};
}
fn setPortFeature(feature: u16, port: u16) usb_abi.Request {
return .{
.request_type = .{ .recipient = .other, .kind = .class, .direction = .host_to_device },
.request_code = .set_feature,
.value = feature,
.index = port,
.length = 0,
};
}
fn clearPortFeature(feature: u16, port: u16) usb_abi.Request {
return .{
.request_type = .{ .recipient = .other, .kind = .class, .direction = .host_to_device },
.request_code = .clear_feature,
.value = feature,
.index = port,
.length = 0,
};
}
fn getPortStatus(port: u16) usb_abi.Request {
return .{
.request_type = .{ .recipient = .other, .kind = .class, .direction = .device_to_host },
.request_code = .get_status,
.value = 0,
.index = port,
.length = 4,
};
}
fn setHubDepth(depth: u16) usb_abi.Request {
return .{
.request_type = .{ .recipient = .device, .kind = .class, .direction = .host_to_device },
.request_code = @enumFromInt(request_set_hub_depth),
.value = depth,
.index = 0,
.length = 0,
};
}
};
pub const Controller = struct { pub const Controller = struct {
register_base: usize, register_base: usize,
op_base: usize, op_base: usize,
@@ -330,6 +454,8 @@ pub const Controller = struct {
subscriptions: [max_subscriptions]Subscription = [_]Subscription{.{}} ** max_subscriptions, subscriptions: [max_subscriptions]Subscription = [_]Subscription{.{}} ** max_subscriptions,
report_queue: [report_queue_capacity]Report = [_]Report{.{}} ** report_queue_capacity, report_queue: [report_queue_capacity]Report = [_]Report{.{}} ** report_queue_capacity,
report_count: usize = 0, report_count: usize = 0,
port_changes: [16]u32 = undefined,
port_change_count: usize = 0,
// Transferred length of the most recent awaited transfer (requested minus the // Transferred length of the most recent awaited transfer (requested minus the
// event residual); read right after a control or bulk transfer returns true. // event residual); read right after a control or bulk transfer returns true.
last_transfer_length: u32 = 0, last_transfer_length: u32 = 0,
@@ -360,6 +486,73 @@ pub const Controller = struct {
} }
// PORTSC for 1-based port `port`. // PORTSC for 1-based port `port`.
/// Diagnostic: walk the xECP list and log each Supported Protocol capability
/// (USB 2.0 vs 3.x, the compatible root-port range), then dump every port's
/// raw PORTSC. Reveals where the USB2 root ports are and their state — for
/// finding a USB2 companion hub that isn't presenting a connection.
pub fn dumpPortTopology(self: *const Controller) void {
const hccparams1 = read32(self.register_base + cap_hccparams1);
var offset: usize = (hccparams1 >> 16) & 0xFFFF; // xECP: dword offset from register_base
var guard: u32 = 0;
while (offset != 0 and guard < 64) : (guard += 1) {
const cap_base = self.register_base + offset * 4;
const dw0 = read32(cap_base);
const id = dw0 & 0xFF;
if (id == 2) { // Supported Protocol
const dw2 = read32(cap_base + 8);
const major = (dw0 >> 24) & 0xFF;
const minor = (dw0 >> 16) & 0xFF;
const port_offset = dw2 & 0xFF;
const port_count = (dw2 >> 8) & 0xFF;
std.log.info("xECP USB {d}.{d}: root ports {d}..{d}", .{ major, minor, port_offset, port_offset + port_count - 1 });
}
const next = (dw0 >> 8) & 0xFF;
if (next == 0) break;
offset += next;
}
var port: u32 = 1;
while (port <= self.max_ports) : (port += 1) {
const portsc = self.portStatus(port);
std.log.info("PORTSC[{d}] 0x{x:0>8}: {s}, {s}, link={s}, power={s}, {s}", .{
port,
portsc,
if (portsc & 1 != 0) "connected" else "empty",
if (portsc & 2 != 0) "enabled" else "disabled",
usb_ids.linkStateName((portsc >> 5) & 0xF),
if (portsc & (1 << 9) != 0) "on" else "off",
usb_ids.speedName((portsc >> 10) & 0xF),
});
}
}
/// Read USB STRING descriptor `index` (English, langid 0x0409) into `out` as
/// ASCII, returning the slice — for logging manufacturer/product names.
/// Null for index 0 (no string) or a failed transfer. Non-ASCII code units
/// become '?'.
pub fn readString(self: *Controller, device: *Device, index: u8, out: []u8) ?[]const u8 {
if (index == 0) return null;
var raw: [256]u8 = undefined;
const request = usb_abi.Request{
.request_type = .{ .recipient = .device, .kind = .standard, .direction = .device_to_host },
.request_code = .get_descriptor,
.value = (@as(u16, 3) << 8) | index, // STRING descriptor
.index = 0x0409, // English (US)
.length = raw.len,
};
if (!self.controlTransfer(device, request, raw[0..], true)) return null;
const length = raw[0]; // bLength; the UTF-16LE payload is bytes 2..length
if (length < 2) return null;
const chars = (@min(length, raw.len) - 2) / 2;
var n: usize = 0;
var i: usize = 0;
while (i < chars and n < out.len) : (i += 1) {
const unit = @as(u16, raw[2 + i * 2]) | (@as(u16, raw[2 + i * 2 + 1]) << 8);
out[n] = if (unit >= 0x20 and unit < 0x7F) @intCast(unit) else '?';
n += 1;
}
return out[0..n];
}
pub fn portStatus(self: *const Controller, port: u32) u32 { pub fn portStatus(self: *const Controller, port: u32) u32 {
return read32(self.op_base + op_portsc_base + op_portsc_stride * (port - 1)); return read32(self.op_base + op_portsc_base + op_portsc_stride * (port - 1));
} }
@@ -429,11 +622,30 @@ pub const Controller = struct {
write64(self.interrupter(event_ring_dequeue_pointer), self.event_ring.segment.physical); write64(self.interrupter(event_ring_dequeue_pointer), self.event_ring.segment.physical);
write64(self.interrupter(event_ring_segment_table_base), self.event_ring.table.physical); write64(self.interrupter(event_ring_segment_table_base), self.event_ring.table.physical);
write32(self.interrupter(interrupter_moderation), 0); write32(self.interrupter(interrupter_moderation), 0);
// Enable the interrupter (IMAN.IE) and USBCMD.INTE. We still POLL the
// Run. (Interrupts are left disabled — the event ring is polled.) // event ring — no interrupt is wired — but some controllers (QEMU's
// qemu-xhci among them) only WRITE runtime events to the ring when the
// interrupter is enabled, so a hot-plug port-change event is silently
// dropped otherwise. Enabling it is harmless to a polling driver.
write32(self.interrupter(interrupter_management), 1 << 1); // IE
mmio.wmb(); mmio.wmb();
write32(self.operational(op_usbcmd), read32(self.operational(op_usbcmd)) | usbcmd_run);
// Run.
mmio.wmb();
write32(self.operational(op_usbcmd), read32(self.operational(op_usbcmd)) | usbcmd_run | usbcmd_interrupter_enable);
if (!waitClear(self.operational(op_usbsts), usbsts_halted)) return null; if (!waitClear(self.operational(op_usbsts), usbsts_halted)) return null;
// Power EVERY port — including empty ones — so a later hot-plug can
// signal a connect (an unpowered port reports nothing: PP=0 is why a
// device added after boot never raised a port-change event). Boot-time
// devices are on already-powered ports; this just extends power to the
// rest. Write PP without disturbing the write-1-to-clear bits.
var port: u32 = 1;
while (port <= self.max_ports) : (port += 1) {
const status = self.portStatus(port);
if (status & portsc_power == 0)
self.writePortStatus(port, (status & ~portsc_write_1_to_clear) | portsc_power);
}
return self; return self;
} }
@@ -601,10 +813,19 @@ pub const Controller = struct {
const base = device.input_context.virtual; const base = device.input_context.virtual;
// Input Control Context (index 0): Add flags in dword 1 = A0 | A1. // Input Control Context (index 0): Add flags in dword 1 = A0 | A1.
contextDword(base, 0, 1, cs).* = 0b11; contextDword(base, 0, 1, cs).* = 0b11;
// Slot Context (index 1): speed[23:20], Context Entries[31:27] = 1. // Slot Context dword 0: Route String[19:0], Speed[23:20], Context
contextDword(base, 1, 0, cs).* = (device.speed << 20) | (@as(u32, 1) << 27); // Entries[31:27] = 1.
// Root Hub Port Number[23:16]. contextDword(base, 1, 0, cs).* = (device.route & 0xFFFFF) | (device.speed << 20) | (@as(u32, 1) << 27);
contextDword(base, 1, 1, cs).* = device.port << 16; // Slot Context dword 1: Root Hub Port Number[23:16] — the ROOT port the
// hub chain hangs off (inherited down a chain), not the device's own
// downstream hub port.
contextDword(base, 1, 1, cs).* = device.root_port << 16;
// Slot Context dword 2: the transaction translator — a full/low-speed
// device behind a high-speed hub routes split transactions through the
// parent hub's TT. Parent Hub Slot ID[7:0], Parent Port Number[13:8].
if (device.parent_slot != 0 and device.speed < 3) {
contextDword(base, 1, 2, cs).* = @as(u32, device.parent_slot) | (@as(u32, device.parent_port) << 8);
}
// EP0 Context (index 2): CErr[2:1]=3, EP Type[5:3]=Control(4), MPS[31:16]. // EP0 Context (index 2): CErr[2:1]=3, EP Type[5:3]=Control(4), MPS[31:16].
contextDword(base, 2, 1, cs).* = (@as(u32, 3) << 1) | (@as(u32, 4) << 3) | (device.max_packet_size_0 << 16); contextDword(base, 2, 1, cs).* = (@as(u32, 3) << 1) | (@as(u32, 4) << 3) | (device.max_packet_size_0 << 16);
// TR Dequeue Pointer (dwords 2:3) with Dequeue Cycle State = 1. // TR Dequeue Pointer (dwords 2:3) with Dequeue Cycle State = 1.
@@ -616,12 +837,31 @@ pub const Controller = struct {
} }
fn addressDeviceCommand(self: *Controller, device: *Device) bool { fn addressDeviceCommand(self: *Controller, device: *Device) bool {
const physical = self.submitCommand(.{ // Retry on a USB Transaction Error (code 4): a freshly-reset device can
.parameter = device.input_context.physical, // miss the first SET_ADDRESS; re-reset the port and try again (xHCI
.control = trbControl(.address_device, @as(u32, device.slot_id) << 24), // 4.6.5). Up to 3 attempts.
}); var attempt: u32 = 0;
const code = self.awaitCommand(physical) orelse return false; while (attempt < 3) : (attempt += 1) {
return code == @intFromEnum(CompletionCode.success); const physical = self.submitCommand(.{
.parameter = device.input_context.physical,
.control = trbControl(.address_device, @as(u32, device.slot_id) << 24),
});
const code = self.awaitCommand(physical) orelse {
std.log.info("port {d} setup: Address Device timed out (attempt {d})", .{ device.port, attempt + 1 });
return false;
};
if (code == @intFromEnum(CompletionCode.success)) return true;
std.log.info("port {d} setup: Address Device completion code {d} (attempt {d})", .{ device.port, code, attempt + 1 });
if (code != @intFromEnum(CompletionCode.usb_transaction_error)) return false;
// Re-reset a root-port device and wait the recovery interval before
// retrying. (A device behind a hub is reset through the hub — not
// retried here; its port was reset in serviceHubPort.)
if (device.parent_slot == 0) {
if (!self.resetPort(device.port)) return false;
system.sleep(10);
} else return false;
}
return false;
} }
/// Reset the port, enable a slot, and address the device on it: after this the /// Reset the port, enable a slot, and address the device on it: after this the
@@ -629,15 +869,45 @@ pub const Controller = struct {
/// or null on any failure. The EP0 MPS is taken from the speed default and /// or null on any failure. The EP0 MPS is taken from the speed default and
/// corrected from the device descriptor by `refreshMaxPacketSize0` if needed. /// corrected from the device descriptor by `refreshMaxPacketSize0` if needed.
pub fn setupDevice(self: *Controller, port: u32, speed: u32) ?*Device { pub fn setupDevice(self: *Controller, port: u32, speed: u32) ?*Device {
if (!self.resetPort(port)) return null; // A SuperSpeed port that has trained its link is ALREADY enabled — the
const slot_id = self.enableSlot() orelse return null; // xHCI advances USB3 ports to Enabled with no reset (spec 4.3). Driving
const device = self.allocateDevice() orelse return null; // a hot reset into a live SS link drops PED mid-reset on real silicon
// (observed: "setup failed" in the same millisecond as "connected").
// Only a not-yet-enabled port — every USB2 device, or a stuck SS link —
// needs the reset to enable.
const already_enabled = speed >= 4 and self.portStatus(port) & portsc_enabled != 0;
var effective_speed = speed;
if (!already_enabled) {
if (!self.resetPort(port)) {
std.log.info("port {d} setup: port reset failed (PORTSC 0x{x:0>8})", .{ port, self.portStatus(port) });
return null;
}
// A USB2 port's PORTSC speed field is only meaningful once the port
// is enabled by the reset — sample it NOW, not at connect time
// (pre-reset reads misreport on real controllers; M20).
effective_speed = (self.portStatus(port) >> 10) & 0xF;
if (effective_speed == 0) effective_speed = speed; // defensive: keep the caller's read
// USB 2.0 spec 7.1.7.5: a device needs a reset-recovery interval
// (TRSTRCY, 10 ms) after reset before it answers SET_ADDRESS.
// Addressing immediately gives a USB Transaction Error (code 4) on
// real full-speed devices; QEMU tolerates the omission.
system.sleep(10);
}
const slot_id = self.enableSlot() orelse {
std.log.info("port {d} setup: Enable Slot failed", .{port});
return null;
};
const device = self.allocateDevice() orelse {
std.log.info("port {d} setup: no free device slot", .{port});
return null;
};
device.* = .{ device.* = .{
.used = true, .used = true,
.slot_id = slot_id, .slot_id = slot_id,
.port = port, .port = port,
.speed = speed, .speed = effective_speed,
.max_packet_size_0 = defaultMaxPacketSize0(speed), .max_packet_size_0 = defaultMaxPacketSize0(effective_speed),
.root_port = port, // a root-port device: the chain root IS this port
}; };
device.input_context = dma.alloc(page_size, dma.coherent) orelse return self.abandon(device); device.input_context = dma.alloc(page_size, dma.coherent) orelse return self.abandon(device);
device.device_context = dma.alloc(page_size, dma.coherent) orelse return self.abandon(device); device.device_context = dma.alloc(page_size, dma.coherent) orelse return self.abandon(device);
@@ -652,22 +922,262 @@ pub const Controller = struct {
return device; return device;
} }
/// Configure an enumerated class-9 device as a hub (docs/usb-hub.md): read
/// the hub descriptor for the downstream port count, tell the controller the
/// slot is a hub (so it routes downstream traffic), SET_HUB_DEPTH for a
/// SuperSpeed hub, and power every downstream port. Downstream enumeration
/// (status-change handling) is B4b. Returns false on a control-transfer
/// failure; the hub is still registered, just inert.
pub fn setupHub(self: *Controller, device: *Device) bool {
const is_usb3 = device.speed >= 4;
var descriptor: [16]u8 = undefined;
const kind: u8 = if (is_usb3) hubreq.descriptor_usb3 else hubreq.descriptor_usb2;
if (!self.controlTransfer(device, hubreq.getDescriptor(kind, descriptor.len), descriptor[0..], true)) {
std.log.info("hub slot {d}: hub descriptor read failed", .{device.slot_id});
return false;
}
device.is_hub = true;
device.hub_ports = descriptor[2]; // bNbrPorts
const characteristics = @as(u16, descriptor[3]) | (@as(u16, descriptor[4]) << 8);
device.hub_multi_tt = !is_usb3 and (characteristics & hubreq.characteristics_multi_tt != 0);
// A SuperSpeed hub needs its depth (tiers from the root) to compose the
// route strings of devices below it.
if (is_usb3) {
const depth = routeDepth(device.route);
_ = self.controlTransfer(device, hubreq.setHubDepth(depth), &.{}, false);
}
// Tell the controller the slot is a hub — Hub bit, Number of Ports, and
// (for a USB2 multi-TT hub) MTT + TT Think Time. Configure Endpoint with
// only the slot add-flag (A0) evaluates these (xHCI 4.6.6).
if (!self.configureSlotAsHub(device)) {
std.log.info("hub slot {d}: could not configure slot as a hub", .{device.slot_id});
return false;
}
// Power every downstream port.
var port: u16 = 1;
while (port <= device.hub_ports) : (port += 1) {
_ = self.controlTransfer(device, hubreq.setPortFeature(hubreq.feature_port_power, port), &.{}, false);
}
// Seed every downstream port as pending: the bus tick GET_STATUSes each
// and enumerates the connected ones. This makes a STATIC topology (a
// device present at power-on) work without relying on the initial
// status-change interrupt edge; the interrupt then handles later plugs.
device.hub_change_mask = if (device.hub_ports >= 31) 0xFFFF_FFFE else (@as(u32, 1) << @intCast(device.hub_ports + 1)) - 2;
// Arm the status-change interrupt endpoint (in-process) for hot-plug.
self.armHubStatus(device);
std.log.info("hub slot {d}: {d} downstream ports powered ({s})", .{
device.slot_id,
device.hub_ports,
if (is_usb3) "SuperSpeed" else if (device.hub_multi_tt) "USB2 multi-TT" else "USB2 single-TT",
});
return true;
}
/// Arm the hub's interrupt-IN status-change endpoint with an in-process
/// subscription: completions set the hub's pending-change mask (serviced on
/// the bus tick). Best-effort — a hub with no interrupt endpoint (shouldn't
/// happen) just relies on the initial sweep.
fn armHubStatus(self: *Controller, device: *Device) void {
for (device.interfaces[0..device.interface_count]) |interface| {
for (interface.endpoints[0..interface.endpoint_count]) |endpoint| {
const is_interrupt = endpoint.transfer_type == 3;
const is_in = endpoint.address & 0x80 != 0;
if (!is_interrupt or !is_in) continue;
const ring = self.getOrConfigureEndpoint(device, endpoint) orelse return;
const subscription = self.allocateSubscription() orelse return;
const buffer = dma.alloc(page_size, dma.coherent) orelse return;
const number: u8 = endpoint.address & 0x0F;
subscription.* = .{
.active = true,
.slot_id = device.slot_id,
.dci = doorbellContextIndex(number, true),
.endpoint_address = endpoint.address,
.ring = ring,
.buffer = buffer,
.max_length = endpoint.max_packet_size,
.hub = device,
};
self.armInterrupt(subscription);
return;
}
}
}
/// The next pending (hub, downstream-port) change to service, or null. Clears
/// the returned port's bit. Called on the bus tick.
pub fn takeHubChange(self: *Controller) ?struct { hub: *Device, port: u16 } {
for (&self.devices) |*device| {
if (!device.used or !device.is_hub or device.hub_change_mask == 0) continue;
const bit: u5 = @intCast(@ctz(device.hub_change_mask));
device.hub_change_mask &= ~(@as(u32, 1) << bit);
if (bit == 0) continue; // bit 0 is the hub itself, not a downstream port
return .{ .hub = device, .port = bit };
}
return null;
}
fn readHubPortStatus(self: *Controller, hub: *Device, port: u16) ?u32 {
var buffer: [4]u8 = undefined;
if (!self.controlTransfer(hub, hubreq.getPortStatus(port), buffer[0..], true)) return null;
return @as(u32, buffer[0]) | (@as(u32, buffer[1]) << 8) | (@as(u32, buffer[2]) << 16) | (@as(u32, buffer[3]) << 24);
}
/// Bring up (or note the disconnect of) a device on hub downstream `port`:
/// read the port status, acknowledge the change bits, and on a fresh connect
/// reset the port, read the speed, and setup+address the downstream device
/// (route string + TT). Returns the addressed device for the bus to enumerate
/// and register, or null (empty port, disconnect, or a failure).
/// Read a downstream hub port's status and acknowledge its latched change
/// bits (so it can signal again). Returns the wPortStatus word.
pub fn hubPortStatusAck(self: *Controller, hub: *Device, port: u16) ?u32 {
const status = self.readHubPortStatus(hub, port) orelse return null;
// Acknowledge every latched change bit. A SuperSpeed hub has extra ones
// (link-state, BH-reset) beyond a USB2 hub's connection/reset — leaving
// any set makes the hub's status-change endpoint re-report the same port
// forever, spinning the driver (a real SuperSpeed hub hung boot here;
// QEMU's USB2 hub has none of these). Clearing an inapplicable feature
// is harmless (the hub STALLs it and we move on).
_ = self.controlTransfer(hub, hubreq.clearPortFeature(hubreq.feature_c_port_connection, port), &.{}, false);
_ = self.controlTransfer(hub, hubreq.clearPortFeature(hubreq.feature_c_port_reset, port), &.{}, false);
if (hub.speed >= 4) {
_ = self.controlTransfer(hub, hubreq.clearPortFeature(hubreq.feature_c_port_link_state, port), &.{}, false);
_ = self.controlTransfer(hub, hubreq.clearPortFeature(hubreq.feature_c_bh_port_reset, port), &.{}, false);
}
return status;
}
pub fn hubPortConnected(status: u32) bool {
return status & hubreq.status_connection != 0;
}
/// Reset + address a device on a connected, empty downstream hub `port` (the
/// bus has confirmed connect and no existing device): reset the port, read
/// the speed, and setup+address the downstream device (route string + TT).
/// Returns the addressed device for the bus to enumerate + register.
pub fn serviceHubPort(self: *Controller, hub: *Device, port: u16) ?*Device {
const status = self.readHubPortStatus(hub, port) orelse return null;
// Reset the port if not yet enabled, then wait (bounded) for enable.
if (status & hubreq.status_enable == 0) {
_ = self.controlTransfer(hub, hubreq.setPortFeature(hubreq.feature_port_reset, port), &.{}, false);
var tries: u32 = 0;
while (tries < 200) : (tries += 1) {
system.sleep(5);
const s = self.readHubPortStatus(hub, port) orelse return null;
if (s & hubreq.status_enable != 0) break;
}
_ = self.controlTransfer(hub, hubreq.clearPortFeature(hubreq.feature_c_port_reset, port), &.{}, false);
}
const enabled = self.readHubPortStatus(hub, port) orelse return null;
if (enabled & hubreq.status_enable == 0) {
std.log.info("hub slot {d} port {d}: reset did not enable", .{ hub.slot_id, port });
return null;
}
const downstream_speed = (enabled >> 9) & 0x3; // wPortStatus: bit9 low-speed, bit10 high-speed
return self.setupDeviceBehindHub(hub, port, downstream_speed);
}
pub fn deviceOnHubPort(self: *Controller, hub: *Device, port: u16) ?*Device {
for (&self.devices) |*device| {
if (device.used and device.parent_slot == hub.slot_id and device.parent_port == port) return device;
}
return null;
}
/// Enable a slot and Address a device behind `hub` on downstream `port`, with
/// the composed route string, inherited root port, and TT fields (so the
/// controller routes split transactions through this hub's TT for a
/// full/low-speed device). Mirrors setupDevice for a root-port device.
fn setupDeviceBehindHub(self: *Controller, hub: *Device, port: u16, speed: u32) ?*Device {
const slot_id = self.enableSlot() orelse {
std.log.info("hub slot {d} port {d}: Enable Slot failed", .{ hub.slot_id, port });
return null;
};
const device = self.allocateDevice() orelse return null;
const child_speed = mapHubPortSpeed(speed);
device.* = .{
.used = true,
.slot_id = slot_id,
.port = hub.root_port,
.speed = child_speed,
.max_packet_size_0 = defaultMaxPacketSize0(child_speed),
.route = (hub.route << 4) | (port & 0xF),
.root_port = hub.root_port,
.parent_slot = hub.slot_id,
.parent_port = @intCast(port),
};
device.input_context = dma.alloc(page_size, dma.coherent) orelse return self.abandon(device);
device.device_context = dma.alloc(page_size, dma.coherent) orelse return self.abandon(device);
device.ep0_ring = .{ .region = dma.alloc(page_size, dma.coherent) orelse return self.abandon(device) };
device.ep0_ring.installLink();
device.control_buffer = dma.alloc(page_size, dma.coherent) orelse return self.abandon(device);
self.buildAddressInputContext(device);
const array: [*]volatile u64 = @ptrFromInt(self.device_context_array.virtual);
array[device.slot_id] = device.device_context.physical;
if (!self.addressDeviceCommand(device)) return self.abandon(device);
return device;
}
/// Configure Endpoint with only A0 (slot) set: rebuild the slot context with
/// the Hub bit, Number of Ports, and MTT/TT-Think-Time, so the controller
/// treats this slot as a hub.
fn configureSlotAsHub(self: *Controller, device: *Device) bool {
const cs = self.context_size;
const base = device.input_context.virtual;
@memset(@as([*]u8, @ptrFromInt(base))[0 .. 2 * cs], 0);
contextDword(base, 0, 1, cs).* = 0b1; // Input Control Context add flags: A0 (slot)
// Slot Context dword 0: route, speed, context entries, plus Hub[26] and
// (USB2 multi-TT) MTT[25].
var dword0: u32 = (device.route & 0xFFFFF) | (device.speed << 20) | (@as(u32, 1) << 27) | (@as(u32, 1) << 26);
if (device.hub_multi_tt) dword0 |= (@as(u32, 1) << 25);
contextDword(base, 1, 0, cs).* = dword0;
// Slot Context dword 1: Root Hub Port Number[23:16], Number of Ports[31:24].
contextDword(base, 1, 1, cs).* = (device.root_port << 16) | (@as(u32, device.hub_ports) << 24);
// Slot Context dword 2: TT Think Time[17:16] = 0 (8 FS bit times); the
// parent-TT fields (if this hub is itself behind a hub) carry over.
if (device.parent_slot != 0 and device.speed < 3) {
contextDword(base, 1, 2, cs).* = @as(u32, device.parent_slot) | (@as(u32, device.parent_port) << 8);
}
return self.configureEndpointCommand(device);
}
fn abandon(self: *Controller, device: *Device) ?*Device { fn abandon(self: *Controller, device: *Device) ?*Device {
_ = self; _ = self;
device.used = false; device.used = false;
return null; return null;
} }
fn awaitTransfer(self: *Controller, requested_length: u32) ?u8 { /// Await the completion of OUR transfer — identified by the event's slot id
/// (control[31:24]) and endpoint DCI (control[20:16]). Any other transfer
/// event is either a subscription's report (serviced) or foreign noise (an
/// interrupt endpoint's error/stale completion whose TRB pointer no longer
/// matches the armed one) — DROPPED, never misattributed: claiming a foreign
/// event as our completion desynchronized the mass-storage bulk protocol in
/// a way that survived every driver restart (the 1-in-3 READ CAPACITY
/// failure at boot, with a USB keyboard and mouse polling concurrently).
fn awaitTransfer(self: *Controller, slot_id: u8, dci: u32, requested_length: u32) ?u8 {
const deadline = system.clock() + 1_000_000_000; const deadline = system.clock() + 1_000_000_000;
while (true) { while (true) {
const event = self.nextEvent(deadline) orelse return null; const event = self.nextEvent(deadline) orelse return null;
if (trbType(event.control) == @intFromEnum(TrbType.transfer_event)) { if (trbType(event.control) != @intFromEnum(TrbType.transfer_event)) continue;
if (self.serviceInterruptEvent(event)) continue; // a subscription's report if (self.serviceInterruptEvent(event)) continue; // a subscription's report
const residual = event.status & 0xFFFFFF; const event_slot: u8 = @truncate(event.control >> 24);
self.last_transfer_length = if (residual >= requested_length) 0 else requested_length - residual; const event_dci: u32 = (event.control >> 16) & 0x1F;
return completionCode(event.status); // our transfer's completion (or error) if (event_slot != slot_id or event_dci != dci) {
std.log.info("dropped foreign transfer event (slot {d} dci {d}, code {d})", .{ event_slot, event_dci, completionCode(event.status) });
continue;
} }
const residual = event.status & 0xFFFFFF;
self.last_transfer_length = if (residual >= requested_length) 0 else requested_length - residual;
return completionCode(event.status); // our transfer's completion (or error)
} }
} }
@@ -708,7 +1218,7 @@ pub const Controller = struct {
mmio.wmb(); mmio.wmb();
self.ringDoorbell(device.slot_id, 1); // DCI 1 = EP0 self.ringDoorbell(device.slot_id, 1); // DCI 1 = EP0
const code = self.awaitTransfer(@intCast(data.len)) orelse return false; const code = self.awaitTransfer(device.slot_id, 1, @intCast(data.len)) orelse return false;
if (code != @intFromEnum(CompletionCode.success) and code != @intFromEnum(CompletionCode.short_packet)) return false; if (code != @intFromEnum(CompletionCode.success) and code != @intFromEnum(CompletionCode.short_packet)) return false;
if (has_data and direction_in) { if (has_data and direction_in) {
@@ -731,7 +1241,56 @@ pub const Controller = struct {
/// endpoints into `device`, and select the configuration. After this the /// endpoints into `device`, and select the configuration. After this the
/// device is in the configured state and its interfaces are ready to match a /// device is in the configured state and its interfaces are ready to match a
/// class driver. Returns false on any control-transfer failure. /// class driver. Returns false on any control-transfer failure.
/// Correct EP0's max packet size from the device itself. The context starts
/// with the SPEED-DEFAULT (full-speed: 8, but the true value may be 8/16/32/
/// 64 — byte 7 of the device descriptor). Read just the descriptor's first
/// 8 bytes (always deliverable at any legal MPS0), and when the device
/// disagrees with the context, issue Evaluate Context to fix EP0 before any
/// longer transfer. Real controllers fault the full 18-byte read on a wrong
/// MPS0; QEMU forgives it — the classic full-speed-mouse-on-real-hardware
/// failure (M20).
fn refreshMaxPacketSize0(self: *Controller, device: *Device) bool {
// SuperSpeed (and above) fix EP0's max packet size at 512, and encode
// bMaxPacketSize0 as an EXPONENT (9 = 2^9 = 512), not a literal size —
// the default is already correct and byte 7 must NOT be read as a size.
// Only full/low/high speed carry a literal 8/16/32/64 that can differ
// from the speed default and need this correction. (Reading the SS
// exponent as a size set EP0 to 9 bytes and broke every following
// transfer — a SuperSpeed hub failing to enumerate on real hardware.)
if (device.speed >= 4) return true;
var head: [8]u8 = undefined;
const request = usb_abi.getDescriptor(.device, 0, 0, 8);
if (!self.controlTransfer(device, request, head[0..], true)) return false;
const actual: u32 = head[7];
if (actual == 0 or actual == device.max_packet_size_0) return true;
// Input Control Context: add-flag A1 (EP0 only); EP0 context rebuilt
// with the corrected MPS. Fields not being changed stay zero (the
// controller evaluates only the added context).
const cs = self.context_size;
const base = device.input_context.virtual;
@memset(@as([*]u8, @ptrFromInt(base))[0 .. 3 * cs], 0);
contextDword(base, 0, 1, cs).* = 0b10; // A1
contextDword(base, 2, 1, cs).* = (@as(u32, 3) << 1) | (@as(u32, 4) << 3) | (actual << 16);
const physical = self.submitCommand(.{
.parameter = device.input_context.physical,
.control = trbControl(.evaluate_context, @as(u32, device.slot_id) << 24),
});
const code = self.awaitCommand(physical) orelse {
std.log.info("slot {d}: Evaluate Context (MPS0 {d} -> {d}) timed out", .{ device.slot_id, device.max_packet_size_0, actual });
return false;
};
if (code != @intFromEnum(CompletionCode.success)) {
std.log.info("slot {d}: Evaluate Context (MPS0 {d} -> {d}) completion code {d}", .{ device.slot_id, device.max_packet_size_0, actual, code });
return false;
}
device.max_packet_size_0 = actual;
return true;
}
pub fn enumerate(self: *Controller, device: *Device) bool { pub fn enumerate(self: *Controller, device: *Device) bool {
if (!self.refreshMaxPacketSize0(device)) return false;
device.device_descriptor = self.getDeviceDescriptor(device) orelse return false; device.device_descriptor = self.getDeviceDescriptor(device) orelse return false;
// The configuration descriptor's own 9 bytes carry the total length of // The configuration descriptor's own 9 bytes carry the total length of
@@ -901,8 +1460,9 @@ pub const Controller = struct {
mmio.wmb(); mmio.wmb();
const number: u8 = endpoint.address & 0x0F; const number: u8 = endpoint.address & 0x0F;
const direction_in = endpoint.address & 0x80 != 0; const direction_in = endpoint.address & 0x80 != 0;
self.ringDoorbell(device.slot_id, doorbellContextIndex(number, direction_in)); const dci = doorbellContextIndex(number, direction_in);
const code = self.awaitTransfer(length) orelse return null; self.ringDoorbell(device.slot_id, dci);
const code = self.awaitTransfer(device.slot_id, dci, length) orelse return null;
if (code != @intFromEnum(CompletionCode.success) and code != @intFromEnum(CompletionCode.short_packet)) return null; if (code != @intFromEnum(CompletionCode.success) and code != @intFromEnum(CompletionCode.short_packet)) return null;
return self.last_transfer_length; return self.last_transfer_length;
} }
@@ -960,9 +1520,20 @@ pub const Controller = struct {
if (!subscription.active or subscription.armed_trb_physical != trb_pointer) continue; if (!subscription.active or subscription.armed_trb_physical != trb_pointer) continue;
const code = completionCode(event.status); const code = completionCode(event.status);
if (code == @intFromEnum(CompletionCode.success) or code == @intFromEnum(CompletionCode.short_packet)) { if (code == @intFromEnum(CompletionCode.success) or code == @intFromEnum(CompletionCode.short_packet)) {
const residual = event.status & 0xFFFFFF; if (subscription.hub) |hub_device| {
const transferred: u16 = if (residual >= subscription.max_length) 0 else @intCast(subscription.max_length - residual); // Hub status-change report: OR the changed-port bitmap into
self.enqueueReport(subscription, transferred); // the hub's pending mask (bit 0 = the hub itself, ignored;
// bit P = downstream port P). The control transfers to
// service it run on the bus tick, not here.
const bytes: [*]const u8 = @ptrFromInt(subscription.buffer.virtual);
var i: usize = 0;
while (i < subscription.max_length and i < 4) : (i += 1)
hub_device.hub_change_mask |= @as(u32, bytes[i]) << @intCast(i * 8);
} else {
const residual = event.status & 0xFFFFFF;
const transferred: u16 = if (residual >= subscription.max_length) 0 else @intCast(subscription.max_length - residual);
self.enqueueReport(subscription, transferred);
}
} }
self.armInterrupt(subscription); // keep polling self.armInterrupt(subscription); // keep polling
return true; return true;
@@ -993,12 +1564,81 @@ pub const Controller = struct {
return report; return report;
} }
/// Drain any events currently on the event ring, dispatching interrupt reports /// Drain any events currently on the event ring: interrupt reports into the
/// into the queue. Non-blocking — called on the driver's timer tick. /// report queue, PORT STATUS CHANGES into the port-change queue (hot-plug —
/// these were silently dropped before M20). Non-blocking — called on the
/// driver's timer tick.
pub fn pump(self: *Controller) void { pub fn pump(self: *Controller) void {
while (true) { while (true) {
const event = self.nextEvent(system.clock()) orelse return; // deadline=now: null when empty const event = self.nextEvent(system.clock()) orelse return; // deadline=now: null when empty
if (trbType(event.control) == @intFromEnum(TrbType.transfer_event)) _ = self.serviceInterruptEvent(event); const kind = trbType(event.control);
if (kind == @intFromEnum(TrbType.transfer_event)) {
_ = self.serviceInterruptEvent(event);
} else if (kind == @intFromEnum(TrbType.port_status_change_event)) {
// Port ID rides bits 31:24 of the TRB's first dword.
const port: u32 = @intCast((event.parameter >> 24) & 0xFF);
if (port == 0 or port > self.max_ports) continue;
// Acknowledge the change bits so the port can signal again.
const status = self.portStatus(port);
self.writePortStatus(port, (status & ~portsc_write_1_to_clear) | (status & portsc_change_mask));
if (self.port_change_count < self.port_changes.len) {
self.port_changes[self.port_change_count] = port;
self.port_change_count += 1;
}
}
} }
} }
/// Dequeue the oldest pending port change (a port whose connect state may
/// have flipped), or null. The bus layer reads PORTSC to decide plug/unplug.
pub fn takePortChange(self: *Controller) ?u32 {
if (self.port_change_count == 0) return null;
const port = self.port_changes[0];
var i: usize = 1;
while (i < self.port_change_count) : (i += 1) self.port_changes[i - 1] = self.port_changes[i];
self.port_change_count -= 1;
return port;
}
/// Whether a port currently has a device connected (PORTSC.CCS).
pub fn portConnected(self: *const Controller, port: u32) bool {
return self.portStatus(port) & portsc_connected != 0;
}
/// The tracked device on `port`, or null.
pub fn deviceOnPort(self: *Controller, port: u32) ?*Device {
for (&self.devices) |*device| {
if (device.used and device.port == port) return device;
}
return null;
}
/// The next used device whose parent hub is `hub_slot` and slot id > `after`
/// (for recursive teardown when a hub itself disconnects), or null.
pub fn nextChildOf(self: *Controller, hub_slot: u8, after: u8) ?*Device {
for (&self.devices) |*device| {
if (device.used and device.parent_slot == hub_slot and device.slot_id > after) return device;
}
return null;
}
/// Tear a device down after unplug: cancel its interrupt subscriptions,
/// Disable Slot (frees the controller's slot state), clear its context-array
/// entry, and release the tracking slot. DMA regions leak (as elsewhere) —
/// bounded by the device-slot count.
pub fn tearDownDevice(self: *Controller, device: *Device) void {
for (&self.subscriptions) |*subscription| {
if (subscription.active and subscription.slot_id == device.slot_id) subscription.active = false;
}
const physical = self.submitCommand(.{
.control = trbControl(.disable_slot, @as(u32, device.slot_id) << 24),
});
if (self.awaitCommand(physical)) |code| {
if (code != @intFromEnum(CompletionCode.success))
std.log.info("slot {d}: Disable Slot completion code {d}", .{ device.slot_id, code });
} else std.log.info("slot {d}: Disable Slot timed out", .{device.slot_id});
const array: [*]volatile u64 = @ptrFromInt(self.device_context_array.virtual);
array[device.slot_id] = 0;
device.used = false;
}
}; };
+58 -50
View File
@@ -18,7 +18,7 @@ const runtime = @import("runtime");
const mmio = @import("mmio"); const mmio = @import("mmio");
const device = runtime.device; const device = runtime.device;
const dma = runtime.dma; const dma = runtime.dma;
const shm = runtime.shm; const shared_memory = runtime.shared_memory;
const system = runtime.system; const system = runtime.system;
const ipc = runtime.ipc; const ipc = runtime.ipc;
const dp = runtime.display_protocol; const dp = runtime.display_protocol;
@@ -53,13 +53,18 @@ const offered_modes = [_]Mode{ .{ .width = 640, .height = 480 }, .{ .width = 800
var current_width: u32 = offered_modes[0].width; var current_width: u32 = offered_modes[0].width;
var current_height: u32 = offered_modes[0].height; var current_height: u32 = offered_modes[0].height;
/// Monotonic fence id for fenced (vsync) flushes; the device signals the fence when the flush /// Monotonic fence id for fenced flushes; the device signals the fence when the flush is
/// is complete, which its used-ring ack already gates our synchronous present on. /// complete, which its used-ring ack already gates our synchronous present on. Completion
/// feedback, not vblank — nothing here is paced to the display's refresh.
var fence_next: u64 = 1; var fence_next: u64 = 1;
/// Whether the device offered VIRTIO_GPU_F_EDID, so `get_edid` is worth issuing. /// Whether the device offered VIRTIO_GPU_F_EDID, so `get_edid` is worth issuing.
var edid_available = false; var edid_available = false;
/// The panel refresh rate parsed from the EDID preferred timing (0 = unknown). Carried to
/// the compositor in the announce so its frame clock paces to the panel, not a guess.
var edid_refresh_hz: u32 = 0;
/// The control virtqueue. We drive it synchronously — one command, notify, poll the used /// The control virtqueue. We drive it synchronously — one command, notify, poll the used
/// ring — so a depth of 16 is ample; we ask the device to shrink to it (virtio 1.0 lets the /// ring — so a depth of 16 is ample; we ask the device to shrink to it (virtio 1.0 lets the
/// driver reduce queue_size), keeping the whole ring inside one page. /// driver reduce queue_size), keeping the whole ring inside one page.
@@ -88,23 +93,16 @@ var bar_virtual: [6]usize = .{ 0, 0, 0, 0, 0, 0 };
var ring: dma.Region = undefined; var ring: dma.Region = undefined;
var command: dma.Region = undefined; var command: dma.Region = undefined;
// The scanout backing is a **shared** (shm) region, not DMA: cacheable so the compositor // The scanout backing is a **shared** (shared-memory) region, not DMA: cacheable so the compositor
// composites into it cheaply (x86 DMA is coherent, so the device still sees the writes), and // composites into it cheaply (x86 DMA is coherent, so the device still sees the writes), and
// shareable so the same physical pages the device scans out of are the ones the compositor // shareable so the same physical pages the device scans out of are the ones the compositor
// paints. The driver keeps the capability to hand to the compositor in the announce. // paints. The driver keeps the capability to hand to the compositor in the announce.
var surface: shm.Region = undefined; var surface: shared_memory.Region = undefined;
// Split-virtqueue producer/consumer shadows. // Split-virtqueue producer/consumer shadows.
var avail_shadow: u16 = 0; var avail_shadow: u16 = 0;
var used_shadow: u16 = 0; var used_shadow: u16 = 0;
/// Format one whole log line and emit it in a single `write`, so this driver's output can
/// never interleave mid-line with the other drivers the manager runs concurrently.
fn log(comptime fmt: []const u8, arguments: anytype) void {
var line: [160]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// --- common-config register access (little-endian MMIO at `common_base`) --------------- // --- common-config register access (little-endian MMIO at `common_base`) ---------------
fn cfgRead(comptime T: type, comptime field: []const u8) T { fn cfgRead(comptime T: type, comptime field: []const u8) T {
@@ -149,7 +147,7 @@ fn mapBar(config: usize, descriptor: *const device.DeviceDescriptor, bar: u8) ?u
return v; return v;
} }
} }
log("virtio-gpu: BAR {d} (physical 0x{x}) is not a mapped resource\n", .{ bar, base }); std.log.info("BAR {d} (physical 0x{x}) is not a mapped resource", .{ bar, base });
return null; return null;
} }
@@ -157,7 +155,7 @@ fn mapBar(config: usize, descriptor: *const device.DeviceDescriptor, bar: u8) ?u
/// notify structures (the only two V3 needs). Returns false if either is missing. /// notify structures (the only two V3 needs). Returns false if either is missing.
fn walkCapabilities(config: usize, descriptor: *const device.DeviceDescriptor) bool { fn walkCapabilities(config: usize, descriptor: *const device.DeviceDescriptor) bool {
if (mmio.read(u16, config + 0x06) & 0x10 == 0) { // Status bit 4: capabilities list present if (mmio.read(u16, config + 0x06) & 0x10 == 0) { // Status bit 4: capabilities list present
log("virtio-gpu: device has no PCI capability list\n", .{}); std.log.info("device has no PCI capability list", .{});
return false; return false;
} }
var cap: u8 = @as(u8, @truncate(mmio.read(u8, config + 0x34))) & 0xFC; var cap: u8 = @as(u8, @truncate(mmio.read(u8, config + 0x34))) & 0xFC;
@@ -188,7 +186,7 @@ fn walkCapabilities(config: usize, descriptor: *const device.DeviceDescriptor) b
cap = next; cap = next;
} }
if (common_base == 0 or notify_base == 0) { if (common_base == 0 or notify_base == 0) {
log("virtio-gpu: missing common-config or notify capability\n", .{}); std.log.info("missing common-config or notify capability", .{});
return false; return false;
} }
return true; return true;
@@ -270,7 +268,7 @@ fn testPixel(index: u32) u32 {
fn initialise(endpoint: ipc.Handle) bool { fn initialise(endpoint: ipc.Handle) bool {
_ = endpoint; _ = endpoint;
if (!device.claim(device_id)) { if (!device.claim(device_id)) {
log("virtio-gpu: unable to claim device {d}\n", .{device_id}); std.log.info("unable to claim device {d}", .{device_id});
return false; return false;
} }
@@ -279,7 +277,7 @@ fn initialise(endpoint: ipc.Handle) bool {
const descriptor = for (descriptors[0..@min(total, descriptors.len)]) |*d| { const descriptor = for (descriptors[0..@min(total, descriptors.len)]) |*d| {
if (d.id == device_id) break d; if (d.id == device_id) break d;
} else { } else {
log("virtio-gpu: device {d} not in the device tree\n", .{device_id}); std.log.info("device {d} not in the device tree", .{device_id});
return false; return false;
}; };
@@ -287,13 +285,13 @@ fn initialise(endpoint: ipc.Handle) bool {
// decode + bus mastering (the device DMAs the ring and backing out of RAM); pci-bus only // decode + bus mastering (the device DMAs the ring and backing out of RAM); pci-bus only
// preserves whatever the firmware left, and a secondary display is often left disabled. // preserves whatever the firmware left, and a secondary display is often left disabled.
const config = device.mmioMap(device_id, 0) orelse { const config = device.mmioMap(device_id, 0) orelse {
log("virtio-gpu: config-space map failed\n", .{}); std.log.info("config-space map failed", .{});
return false; return false;
}; };
const vendor = mmio.read(u16, config + 0x00); const vendor = mmio.read(u16, config + 0x00);
const dev = mmio.read(u16, config + 0x02); const dev = mmio.read(u16, config + 0x02);
if (vendor != virtio_vendor or dev != virtio_gpu_device) { if (vendor != virtio_vendor or dev != virtio_gpu_device) {
log("virtio-gpu: not a virtio-gpu (vendor 0x{x} device 0x{x})\n", .{ vendor, dev }); std.log.info("not a virtio-gpu (vendor 0x{x} device 0x{x})", .{ vendor, dev });
return false; return false;
} }
mmio.write(u16, config + 0x04, mmio.read(u16, config + 0x04) | 0x06); // MEM + bus master mmio.write(u16, config + 0x04, mmio.read(u16, config + 0x04) | 0x06); // MEM + bus master
@@ -312,7 +310,7 @@ fn initialise(endpoint: ipc.Handle) bool {
// High feature word: VERSION_1 (bit 32) is required for a modern device. // High feature word: VERSION_1 (bit 32) is required for a modern device.
cfgWrite(u32, "device_feature_select", vp.feature_version_1_word); cfgWrite(u32, "device_feature_select", vp.feature_version_1_word);
if (cfgRead(u32, "device_feature") & vp.feature_version_1_bit == 0) { if (cfgRead(u32, "device_feature") & vp.feature_version_1_bit == 0) {
log("virtio-gpu: device does not offer VERSION_1 (not a modern device)\n", .{}); std.log.info("device does not offer VERSION_1 (not a modern device)", .{});
return false; return false;
} }
// Accept exactly VERSION_1, plus EDID when the device offered it (never a feature it didn't). // Accept exactly VERSION_1, plus EDID when the device offered it (never a feature it didn't).
@@ -322,7 +320,7 @@ fn initialise(endpoint: ipc.Handle) bool {
cfgWrite(u32, "driver_feature", vp.feature_version_1_bit); cfgWrite(u32, "driver_feature", vp.feature_version_1_bit);
orStatus(vp.status_features_ok); orStatus(vp.status_features_ok);
if (cfgRead(u8, "device_status") & vp.status_features_ok == 0) { if (cfgRead(u8, "device_status") & vp.status_features_ok == 0) {
log("virtio-gpu: device rejected the negotiated features\n", .{}); std.log.info("device rejected the negotiated features", .{});
return false; return false;
} }
@@ -330,15 +328,15 @@ fn initialise(endpoint: ipc.Handle) bool {
cfgWrite(u16, "queue_select", 0); cfgWrite(u16, "queue_select", 0);
const device_qsize = cfgRead(u16, "queue_size"); const device_qsize = cfgRead(u16, "queue_size");
if (device_qsize < queue_size) { if (device_qsize < queue_size) {
log("virtio-gpu: control queue too small ({d})\n", .{device_qsize}); std.log.info("control queue too small ({d})", .{device_qsize});
return false; return false;
} }
ring = dma.alloc(4096, dma.coherent) orelse { ring = dma.alloc(4096, dma.coherent) orelse {
log("virtio-gpu: virtqueue allocation failed\n", .{}); std.log.info("virtqueue allocation failed", .{});
return false; return false;
}; };
command = dma.alloc(4096, dma.coherent) orelse { command = dma.alloc(4096, dma.coherent) orelse {
log("virtio-gpu: command-buffer allocation failed\n", .{}); std.log.info("command-buffer allocation failed", .{});
return false; return false;
}; };
mmio.write(u16, ring.virtual + avail_offset, 1); // VIRTQ_AVAIL_F_NO_INTERRUPT: we poll mmio.write(u16, ring.virtual + avail_offset, 1); // VIRTQ_AVAIL_F_NO_INTERRUPT: we poll
@@ -366,19 +364,19 @@ fn initialise(endpoint: ipc.Handle) bool {
.height = max_height, .height = max_height,
}; };
if (command_nodata(@sizeOf(vg.ResourceCreate2d)) != ok_nodata) { if (command_nodata(@sizeOf(vg.ResourceCreate2d)) != ok_nodata) {
log("virtio-gpu: resource_create_2d failed\n", .{}); std.log.info("resource_create_2d failed", .{});
return false; return false;
} }
} }
// Back the resource with a shared (shm) surface, so the compositor and the device work // Back the resource with a shared (shared-memory) surface, so the compositor and the device work
// the same physical pages. The device needs the guest-physical base for attach_backing. // the same physical pages. The device needs the guest-physical base for attach_backing.
surface = shm.create(scanout_bytes) orelse { surface = shared_memory.create(scanout_bytes) orelse {
log("virtio-gpu: scanout surface allocation failed\n", .{}); std.log.info("scanout surface allocation failed", .{});
return false; return false;
}; };
const surface_physical = shm.physical(surface.handle) orelse { const surface_physical = shared_memory.physical(surface.handle) orelse {
log("virtio-gpu: could not resolve the scanout surface physical address\n", .{}); std.log.info("could not resolve the scanout surface physical address", .{});
return false; return false;
}; };
{ {
@@ -391,15 +389,15 @@ fn initialise(endpoint: ipc.Handle) bool {
const entry: *vg.MemEntry = @ptrFromInt(command.virtual + request_offset + @sizeOf(vg.ResourceAttachBacking)); const entry: *vg.MemEntry = @ptrFromInt(command.virtual + request_offset + @sizeOf(vg.ResourceAttachBacking));
entry.* = .{ .addr = surface_physical, .length = @intCast(scanout_bytes) }; entry.* = .{ .addr = surface_physical, .length = @intCast(scanout_bytes) };
if (command_nodata(@sizeOf(vg.ResourceAttachBacking) + @sizeOf(vg.MemEntry)) != ok_nodata) { if (command_nodata(@sizeOf(vg.ResourceAttachBacking) + @sizeOf(vg.MemEntry)) != ok_nodata) {
log("virtio-gpu: resource_attach_backing failed\n", .{}); std.log.info("resource_attach_backing failed", .{});
return false; return false;
} }
} }
if (!setScanoutRect()) { if (!setScanoutRect()) {
log("virtio-gpu: set_scanout failed\n", .{}); std.log.info("set_scanout failed", .{});
return false; return false;
} }
log("virtio-gpu: scanout {d}x{d} online\n", .{ current_width, current_height }); std.log.info("scanout {d}x{d} online", .{ current_width, current_height });
// Hello the device manager so it counts us as up (and does not stop us at the hello // Hello the device manager so it counts us as up (and does not stop us at the hello
// deadline). A restarted instance re-hellos here and re-announces below — the compositor // deadline). A restarted instance re-hellos here and re-announces below — the compositor
@@ -417,17 +415,17 @@ fn initialise(endpoint: ipc.Handle) bool {
for (0..pixel_count) |i| pixels[i] = testPixel(@intCast(i)); for (0..pixel_count) |i| pixels[i] = testPixel(@intCast(i));
if (!presentFull()) { if (!presentFull()) {
log("virtio-gpu: initial present failed\n", .{}); std.log.info("initial present failed", .{});
return false; return false;
} }
// The scanout surface is CPU-visible RAM: read the pattern back to prove the mapping, // The scanout surface is CPU-visible RAM: read the pattern back to prove the mapping,
// which together with the flush ack above is the automated stand-in for "it's on screen". // which together with the flush ack above is the automated stand-in for "it's on screen".
mmio.rmb(); mmio.rmb();
if (pixels[0] != testPixel(0) or pixels[pixel_count / 2] != testPixel(@intCast(pixel_count / 2))) { if (pixels[0] != testPixel(0) or pixels[pixel_count / 2] != testPixel(@intCast(pixel_count / 2))) {
log("virtio-gpu: pixel read-back mismatch\n", .{}); std.log.info("pixel read-back mismatch", .{});
return false; return false;
} }
log("virtio-gpu: flush acked, pixel check ok\n", .{}); std.log.info("flush acked, pixel check ok", .{});
// Offer the shared surface to the compositor so it upgrades off the GOP floor (V4). // Offer the shared surface to the compositor so it upgrades off the GOP floor (V4).
announce(); announce();
@@ -451,26 +449,34 @@ fn setScanoutRect() bool {
/// a device that doesn't offer EDID, or a missing/short block, is logged and ignored. /// a device that doesn't offer EDID, or a missing/short block, is logged and ignored.
fn readEdid() void { fn readEdid() void {
if (!edid_available) { if (!edid_available) {
log("virtio-gpu: EDID not offered by device\n", .{}); std.log.info("EDID not offered by device", .{});
return; return;
} }
const request = requestAt(vg.GetEdid); const request = requestAt(vg.GetEdid);
request.* = .{ .hdr = .{ .type = @intFromEnum(vg.CmdType.get_edid) }, .scanout = 0 }; request.* = .{ .hdr = .{ .type = @intFromEnum(vg.CmdType.get_edid) }, .scanout = 0 };
if (!submit(@sizeOf(vg.GetEdid), @sizeOf(vg.RespEdid))) { if (!submit(@sizeOf(vg.GetEdid), @sizeOf(vg.RespEdid))) {
log("virtio-gpu: EDID request not acked\n", .{}); std.log.info("EDID request not acked", .{});
return; return;
} }
const response: *vg.RespEdid = @ptrFromInt(command.virtual + response_offset); const response: *vg.RespEdid = @ptrFromInt(command.virtual + response_offset);
if (response.hdr.type != @intFromEnum(vg.CmdType.resp_ok_edid) or response.size < 64) { if (response.hdr.type != @intFromEnum(vg.CmdType.resp_ok_edid) or response.size < 64) {
log("virtio-gpu: EDID unavailable\n", .{}); std.log.info("EDID unavailable", .{});
return; return;
} }
// The first detailed timing descriptor (EDID base-block offset 54) is the preferred mode: // The first detailed timing descriptor (EDID base-block offset 54) is the preferred mode:
// active pixels are 12-bit, low byte + high nibble (bytes 2/4 horizontal, 5/7 vertical). // active pixels are 12-bit, low byte + high nibble (bytes 2/4 horizontal, 5/7 vertical).
// The refresh rate is derived from the same descriptor: pixel clock (bytes 0-1, 10 kHz
// units) over total (active + blanking) pixels per frame — the loader does the identical
// computation for the boot framebuffer (boot/efi.zig edidNative).
const e = &response.edid; const e = &response.edid;
const h_active = @as(u32, e[56]) | (@as(u32, e[58] & 0xF0) << 4); const h_active = @as(u32, e[56]) | (@as(u32, e[58] & 0xF0) << 4);
const v_active = @as(u32, e[59]) | (@as(u32, e[61] & 0xF0) << 4); const v_active = @as(u32, e[59]) | (@as(u32, e[61] & 0xF0) << 4);
log("virtio-gpu: EDID preferred mode {d}x{d}\n", .{ h_active, v_active }); const clock_hz = (@as(u64, e[54]) | (@as(u64, e[55]) << 8)) * 10_000;
const h_blank = @as(u64, e[57]) | (@as(u64, e[58] & 0x0F) << 8);
const v_blank = @as(u64, e[60]) | (@as(u64, e[61] & 0x0F) << 8);
const total = (@as(u64, h_active) + h_blank) * (@as(u64, v_active) + v_blank);
if (total != 0) edid_refresh_hz = @intCast((clock_hz + total / 2) / total);
std.log.info("EDID preferred mode {d}x{d} @ {d} Hz", .{ h_active, v_active, edid_refresh_hz });
} }
/// Present the whole surface: copy the guest backing into the host resource, then flush it to /// Present the whole surface: copy the guest backing into the host resource, then flush it to
@@ -492,8 +498,9 @@ fn presentFull() bool {
if (command_nodata(@sizeOf(vg.TransferToHost2d)) != ok_nodata) return false; if (command_nodata(@sizeOf(vg.TransferToHost2d)) != ok_nodata) return false;
} }
{ {
// A fenced flush (vsync): the device signals the fence when the frame is actually on // A fenced flush: the device signals the fence once it has consumed the frame — which
// screen — which its used-ring ack, what our synchronous submit waits on, already gates. // its used-ring ack, what our synchronous submit waits on, already gates. Completion
// feedback and a tear-free snapshot, not vblank pacing.
const request = requestAt(vg.ResourceFlush); const request = requestAt(vg.ResourceFlush);
request.* = .{ request.* = .{
.hdr = .{ .type = @intFromEnum(vg.CmdType.resource_flush), .flags = vg.flag_fence, .fence_id = fence_next }, .hdr = .{ .type = @intFromEnum(vg.CmdType.resource_flush), .flags = vg.flag_fence, .fence_id = fence_next },
@@ -515,20 +522,20 @@ fn helloManager() void {
if (ipc.lookup(.device_manager)) |h| break h; if (ipc.lookup(.device_manager)) |h| break h;
system.sleep(20); system.sleep(20);
} else { } else {
log("virtio-gpu: no device manager to hello\n", .{}); std.log.info("no device manager to hello", .{});
return; return;
}; };
const hello = dm.Hello{ .role = @intFromEnum(dm.Role.bus), .device_id = device_id }; const hello = dm.Hello{ .role = @intFromEnum(dm.Role.bus), .device_id = device_id };
var reply: [dm.reply_size]u8 = undefined; var reply: [dm.reply_size]u8 = undefined;
const n = ipc.call(manager, std.mem.asBytes(&hello), &reply) catch { const n = ipc.call(manager, std.mem.asBytes(&hello), &reply) catch {
log("virtio-gpu: hello call failed\n", .{}); std.log.info("hello call failed", .{});
return; return;
}; };
if (n < dm.reply_size or std.mem.bytesToValue(dm.HelloReply, reply[0..dm.reply_size]).status != 0) { if (n < dm.reply_size or std.mem.bytesToValue(dm.HelloReply, reply[0..dm.reply_size]).status != 0) {
log("virtio-gpu: hello refused\n", .{}); std.log.info("hello refused", .{});
return; return;
} }
log("virtio-gpu: hello acknowledged\n", .{}); std.log.info("hello acknowledged", .{});
} }
/// Announce the scanout to the display service so it upgrades off the GOP framebuffer: hand it /// Announce the scanout to the display service so it upgrades off the GOP framebuffer: hand it
@@ -542,22 +549,23 @@ fn announce() void {
if (ipc.lookup(.display)) |h| break h; if (ipc.lookup(.display)) |h| break h;
system.sleep(20); system.sleep(20);
} else { } else {
log("virtio-gpu: no display service to announce to (scanout-only)\n", .{}); std.log.info("no display service to announce to (scanout-only)", .{});
return; return;
}; };
var request = dp.Request{ var request = dp.Request{
.operation = @intFromEnum(dp.Operation.attach_scanout), .operation = @intFromEnum(dp.Operation.attach_scanout),
.x = max_width, // the shared surface's row stride in pixels (it is sized to the max mode) .x = max_width, // the shared surface's row stride in pixels (it is sized to the max mode)
.y = edid_refresh_hz, // the panel refresh from EDID (0 = unknown) — the frame-clock seed
.width = current_width, .width = current_width,
.height = current_height, .height = current_height,
.colour = display_format_bgrx, .colour = display_format_bgrx,
}; };
var reply: [dp.reply_size]u8 = undefined; var reply: [dp.reply_size]u8 = undefined;
_ = ipc.callCap(display, std.mem.asBytes(&request), &reply, surface.handle) catch { _ = ipc.callCap(display, std.mem.asBytes(&request), &reply, surface.handle) catch {
log("virtio-gpu: announce to display failed\n", .{}); std.log.info("announce to display failed", .{});
return; return;
}; };
log("virtio-gpu: announced scanout to display\n", .{}); std.log.info("announced scanout to display", .{});
} }
/// A `sp.Reply{status}` written into `reply`. /// A `sp.Reply{status}` written into `reply`.
@@ -606,7 +614,7 @@ pub fn main(init: runtime.process.Init) void {
return; return;
}; };
device_id = std.fmt.parseInt(u64, argument, 10) catch { device_id = std.fmt.parseInt(u64, argument, 10) catch {
log("virtio-gpu: malformed device id '{s}'\n", .{argument}); std.log.info("malformed device id '{s}'", .{argument});
return; return;
}; };
runtime.service.run(256, .{ runtime.service.run(256, .{
+92 -8
View File
@@ -1,7 +1,13 @@
//! The initial_ramdisk (initial ramdisk) container format — shared by the build-time //! The initial_ramdisk (initial ramdisk) container format — built in RAM by the
//! packer (tools/make-initial-ramdisk.py) and the kernel that unpacks it. Deliberately //! bootloader (boot/efi.zig walks the boot volume's /system tree) and unpacked by
//! trivial: a header, a table of fixed-size entries, then the concatenated file //! the kernel. Deliberately trivial: a header, a table of fixed-size entries, then
//! blobs. We own both producer and consumer, so it need be no fancier. //! the concatenated file blobs. We own both producer and consumer, so it need be
//! no fancier.
//!
//! v2: entry names are full FHS paths ("/system/services/init"), 64 bytes — the
//! same limit as a task name (abi.maximum_process_name), so a path-named task is
//! never truncated. The boot volume's file tree is the single source of truth;
//! this image is only the loader→kernel handoff snapshot of it.
//! //!
//! Layout: //! Layout:
//! Header (magic, count) //! Header (magic, count)
@@ -10,8 +16,14 @@
const std = @import("std"); const std = @import("std");
/// "DNRD" — identifies a danos initial_ramdisk image. /// "DNR2" — identifies a danos initial_ramdisk image, format v2 (path names).
pub const magic: u32 = 0x444E5244; /// The v1 magic ("DNRD", basename entries) is rejected: a stale image should
/// fail loudly at Reader.init, not misparse names.
pub const magic: u32 = 0x32524E44;
/// Entry name capacity. Matches abi.maximum_process_name so a spawned task can
/// always carry its full binary path as its name.
pub const maximum_name = 64;
pub const Header = extern struct { pub const Header = extern struct {
magic: u32, magic: u32,
@@ -19,11 +31,17 @@ pub const Header = extern struct {
}; };
pub const Entry = extern struct { pub const Entry = extern struct {
name: [32]u8, // NUL-padded file name (basename) name: [maximum_name]u8, // NUL-padded FHS path, e.g. "/system/services/init"
offset: u64, // byte offset of the blob within the image offset: u64, // byte offset of the blob within the image
len: u64, // blob length in bytes len: u64, // blob length in bytes
}; };
/// The basename of a path: the final component after the last '/'.
pub fn basename(path: []const u8) []const u8 {
const i = std.mem.lastIndexOfScalar(u8, path, '/') orelse return path;
return path[i + 1 ..];
}
/// A validated view over an initial_ramdisk image. `init` checks the magic and that the /// A validated view over an initial_ramdisk image. `init` checks the magic and that the
/// entry table fits; `entry` bounds-checks each blob against the image. /// entry table fits; `entry` bounds-checks each blob against the image.
pub const Reader = struct { pub const Reader = struct {
@@ -48,11 +66,77 @@ pub const Reader = struct {
if (e.offset > self.image.len or e.len > self.image.len - e.offset) return null; if (e.offset > self.image.len or e.len > self.image.len - e.offset) return null;
// The name is stored in the entry's fixed field; return a stable slice // The name is stored in the entry's fixed field; return a stable slice
// into the image (not the value copy) up to the NUL terminator. // into the image (not the value copy) up to the NUL terminator.
const name_field = self.image[off .. off + 32]; const name_field = self.image[off .. off + maximum_name];
const nlen = std.mem.indexOfScalar(u8, name_field, 0) orelse name_field.len; const nlen = std.mem.indexOfScalar(u8, name_field, 0) orelse name_field.len;
return .{ return .{
.name = name_field[0..nlen], .name = name_field[0..nlen],
.blob = self.image[@intCast(e.offset)..][0..@intCast(e.len)], .blob = self.image[@intCast(e.offset)..][0..@intCast(e.len)],
}; };
} }
/// Look a binary up by name: an exact path match wins; otherwise a unique
/// basename match ("fat" finds "/system/services/fat") keeps pre-path callers
/// working. Comparisons are ASCII case-insensitive — the entries come from a
/// FAT volume, whose name lookups are case-insensitive by definition (and
/// whose short entries store uppercase). The returned Item's name is always
/// the stored full path.
pub fn find(self: Reader, name: []const u8) ?Item {
var i: u32 = 0;
while (i < self.count) : (i += 1) {
const item = self.entry(i) orelse continue;
if (std.ascii.eqlIgnoreCase(item.name, name)) return item;
}
i = 0;
while (i < self.count) : (i += 1) {
const item = self.entry(i) orelse continue;
if (std.ascii.eqlIgnoreCase(basename(item.name), name)) return item;
}
return null;
}
}; };
// --- tests (host) -----------------------------------------------------------
fn testImage(buffer: []u8, entries: []const struct { name: []const u8, blob: []const u8 }) []const u8 {
const table_end = @sizeOf(Header) + entries.len * @sizeOf(Entry);
var offset: usize = table_end;
std.mem.bytesAsValue(Header, buffer[0..@sizeOf(Header)]).* = .{ .magic = magic, .count = @intCast(entries.len) };
for (entries, 0..) |e, i| {
var record = Entry{ .name = @splat(0), .offset = offset, .len = e.blob.len };
@memcpy(record.name[0..e.name.len], e.name);
std.mem.bytesAsValue(Entry, buffer[@sizeOf(Header) + i * @sizeOf(Entry) ..][0..@sizeOf(Entry)]).* = record;
@memcpy(buffer[offset..][0..e.blob.len], e.blob);
offset += e.blob.len;
}
return buffer[0..offset];
}
test "find matches exact path, then unique basename; name is the stored path" {
var buffer: [1024]u8 = undefined;
const image = testImage(&buffer, &.{
.{ .name = "/system/services/init", .blob = "INIT" },
.{ .name = "/system/drivers/ps2-bus", .blob = "PS2" },
});
const rd = Reader.init(image).?;
const by_path = rd.find("/system/services/init").?;
try std.testing.expectEqualStrings("/system/services/init", by_path.name);
try std.testing.expectEqualStrings("INIT", by_path.blob);
const by_base = rd.find("ps2-bus").?;
try std.testing.expectEqualStrings("/system/drivers/ps2-bus", by_base.name);
try std.testing.expectEqualStrings("PS2", by_base.blob);
try std.testing.expect(rd.find("no-such-binary") == null);
}
test "v1 magic is rejected" {
var buffer: [64]u8 = @splat(0);
std.mem.bytesAsValue(Header, buffer[0..@sizeOf(Header)]).* = .{ .magic = 0x444E5244, .count = 0 };
try std.testing.expect(Reader.init(&buffer) == null);
}
test "basename" {
try std.testing.expectEqualStrings("fat", basename("/system/services/fat"));
try std.testing.expectEqualStrings("fat", basename("fat"));
}
+26 -11
View File
@@ -529,8 +529,22 @@ var warp_ap_ready: u32 = 0;
var warp_stop: u32 = 0; var warp_stop: u32 = 0;
var warp_checks: u32 = 0; // completed per-AP rendezvous count (for the tsc-sync test) var warp_checks: u32 = 0; // completed per-AP rendezvous count (for the tsc-sync test)
const warp_rounds: u32 = 1 << 20; // locked reads on the BSP: ~1 ms at GHz rates // The warp check is bounded by TIME, not iterations: a warp tick is a locked
const warp_spin_limit: u64 = 1 << 32; // bound every rendezvous wait so a lost core can't hang boot // read-modify-write on a cacheline two cores are fighting over — microseconds
// under real contention, not the nanosecond an uncontended count assumes (a
// 1<<20-round budget measured 2-21 SECONDS per core on a 16-core machine), and
// a PAUSE costs ~140 cycles on modern Intel, so an iteration-counted await
// mis-measures by two orders of magnitude too. ~5 ms of pairwise hammering per
// core is plenty to catch a lagging TSC (Linux's check_tsc_warp budget), and
// ~100 ms is a generous rendezvous window for a healthy core.
const warp_check_ns: u64 = 5_000_000; // per-AP pairwise check duration
const warp_await_ns: u64 = 100_000_000; // rendezvous wait before giving up
/// TSC ticks for `ns` nanoseconds (valid whenever the warp check runs: the TSC
/// is the clocksource, so tsc_hz is calibrated).
fn warpTicksFor(ns: u64) u64 {
return @intCast(@as(u128, ns) * tsc_hz / 1_000_000_000);
}
fn warpTick() void { fn warpTick() void {
while (@cmpxchgWeak(u32, &warp_lock, 0, 1, .acquire, .monotonic) != null) asm volatile ("pause"); while (@cmpxchgWeak(u32, &warp_lock, 0, 1, .acquire, .monotonic) != null) asm volatile ("pause");
@@ -545,11 +559,11 @@ fn warpTick() void {
@atomicStore(u32, &warp_lock, 0, .release); @atomicStore(u32, &warp_lock, 0, .release);
} }
/// Spin (bounded) until `flag` is nonzero; false on timeout. /// Spin (time-bounded) until `flag` is nonzero; false on timeout.
fn warpAwait(flag: *u32) bool { fn warpAwait(flag: *u32) bool {
var spins: u64 = 0; const deadline = rdtsc() +% warpTicksFor(warp_await_ns);
while (@atomicLoad(u32, flag, .acquire) == 0) : (spins += 1) { while (@atomicLoad(u32, flag, .acquire) == 0) {
if (spins >= warp_spin_limit) return false; if (rdtsc() -% deadline < (1 << 62)) return false; // past the deadline
asm volatile ("pause"); asm volatile ("pause");
} }
return true; return true;
@@ -568,8 +582,8 @@ pub fn checkWarpSource() void {
@atomicStore(u32, &warp_bsp_ready, 0, .release); @atomicStore(u32, &warp_bsp_ready, 0, .release);
return; return;
} }
var i: u32 = 0; const deadline = rdtsc() +% warpTicksFor(warp_check_ns);
while (i < warp_rounds) : (i += 1) warpTick(); while (rdtsc() -% deadline >= (1 << 62)) warpTick(); // until the time budget is spent
@atomicStore(u32, &warp_stop, 1, .release); @atomicStore(u32, &warp_stop, 1, .release);
@atomicStore(u32, &warp_bsp_ready, 0, .release); @atomicStore(u32, &warp_bsp_ready, 0, .release);
warp_checks += 1; warp_checks += 1;
@@ -583,9 +597,10 @@ pub fn checkWarpTarget() void {
if (clock_source != .tsc) return; if (clock_source != .tsc) return;
if (!warpAwait(&warp_bsp_ready)) return; if (!warpAwait(&warp_bsp_ready)) return;
@atomicStore(u32, &warp_ap_ready, 1, .release); @atomicStore(u32, &warp_ap_ready, 1, .release);
var spins: u64 = 0; // The BSP owns the budget; this bound only protects against a lost BSP.
while (@atomicLoad(u32, &warp_stop, .acquire) == 0) : (spins += 1) { const deadline = rdtsc() +% warpTicksFor(2 * warp_check_ns + warp_await_ns);
if (spins >= warp_spin_limit) return; while (@atomicLoad(u32, &warp_stop, .acquire) == 0) {
if (rdtsc() -% deadline < (1 << 62)) return; // past the deadline
warpTick(); warpTick();
} }
} }
+2 -2
View File
@@ -189,8 +189,8 @@ pub fn mapUserDmaInto(root: u64, virtual: u64, physical: u64, len: u64) void {
} }
/// Map shared cacheable RAM into address space `root`: write-back cacheable, RW+NX, and /// Map shared cacheable RAM into address space `root`: write-back cacheable, RW+NX, and
/// marked so teardown won't free the frames (they're owned by a refcounted shm object, /// marked so teardown won't free the frames (they're owned by a refcounted shared-memory object,
/// freed when its last capability drops). For shm_create/shm_map. /// freed when its last capability drops). For shared_memory_create/shared_memory_map.
pub fn mapUserSharedInto(root: u64, virtual: u64, physical: u64, len: u64) void { pub fn mapUserSharedInto(root: u64, virtual: u64, physical: u64, len: u64) void {
paging.mapUserSharedInto(root, virtual, physical, len); paging.mapUserSharedInto(root, virtual, physical, len);
} }
+6 -10
View File
@@ -138,14 +138,16 @@ pub const Console = struct {
} }
} }
/// Shift the visible text up one glyph row and clear the freed bottom row, /// The screen is full: start a fresh page at the top. NEVER scroll by
/// leaving the cursor on that now-blank last line. /// copying pixel rows — that READS the framebuffer, and VRAM reads are
/// uncached-slow on real hardware (measured: 16-core bring-up took ~90 s
/// purely from boot lines each paying a whole-screen scroll copy). A page
/// clear is writes only, and only once per screenful.
fn scroll(self: *Console) void { fn scroll(self: *Console) void {
const visible = self.rows * glyph_h; const visible = self.rows * glyph_h;
var y: u32 = 0; var y: u32 = 0;
while (y + glyph_h < visible) : (y += 1) self.copyRow(y, y + glyph_h);
while (y < visible) : (y += 1) self.fillRow(y, self.bg); while (y < visible) : (y += 1) self.fillRow(y, self.bg);
self.row = self.rows - 1; self.row = 0;
} }
inline fn rowPtr(self: *Console, y: u32) [*]volatile u32 { inline fn rowPtr(self: *Console, y: u32) [*]volatile u32 {
@@ -163,10 +165,4 @@ pub const Console = struct {
while (x < self.fb.width) : (x += 1) row[x] = color; while (x < self.fb.width) : (x += 1) row[x] = color;
} }
fn copyRow(self: *Console, destination_y: u32, source_y: u32) void {
const destination = self.rowPtr(destination_y);
const source = self.rowPtr(source_y);
var x: u32 = 0;
while (x < self.fb.width) : (x += 1) destination[x] = source[x];
}
}; };
+2 -2
View File
@@ -61,7 +61,7 @@ pub fn init(device_tree: *const platform.DeviceTree) void {
/// [[boot-handoff]], not the device tree), so it is seeded explicitly, after `init`. /// [[boot-handoff]], not the device tree), so it is seeded explicitly, after `init`.
/// Returns the new device id, or null when there is no framebuffer (headless) or the /// Returns the new device id, or null when there is no framebuffer (headless) or the
/// table is full. Idempotent-ish: only ever call once per boot. /// table is full. Idempotent-ish: only ever call once per boot.
pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32) ?u64 { pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32) ?u64 {
if (base == 0 or width == 0 or height == 0) return null; // headless if (base == 0 or width == 0 or height == 0) return null; // headless
if (count >= maximum_devices) { if (count >= maximum_devices) {
dropped += 1; dropped += 1;
@@ -79,7 +79,7 @@ pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32)
.len = @as(u64, height) * pitch, .len = @as(u64, height) * pitch,
.flags = device_abi.resource_flag_write_combining, .flags = device_abi.resource_flag_write_combining,
}; };
d.display = .{ .width = width, .height = height, .pitch = pitch, .format = format }; d.display = .{ .width = width, .height = height, .pitch = pitch, .format = format, .refresh_hz = refresh_hz };
devices[count] = d; devices[count] = d;
display_device = d.id; display_device = d.id;
count += 1; count += 1;
+35 -21
View File
@@ -153,35 +153,35 @@ pub fn dropRef(endpoint: *Endpoint) void {
/// The `kind` tag on a `scheduler.HandleObject` — which capability object a handle names. /// The `kind` tag on a `scheduler.HandleObject` — which capability object a handle names.
/// Defined here (not in scheduler) because the meaning is the IPC/capability layer's. /// Defined here (not in scheduler) because the meaning is the IPC/capability layer's.
pub const handle_kind_endpoint: u8 = 0; pub const handle_kind_endpoint: u8 = 0;
pub const handle_kind_shm: u8 = 1; pub const handle_kind_shared_memory: u8 = 1;
/// A page-aligned block of **shared cacheable RAM** (docs/display-v2.md), referenced by /// A page-aligned block of **shared cacheable RAM** (docs/display-v2.md), referenced by
/// capability handles across processes and freed when the last one drops. `phys` is its /// capability handles across processes and freed when the last one drops. `phys` is its
/// contiguous physical base, `pages` its length. A sharer's address-space teardown never /// contiguous physical base, `pages` its length. A sharer's address-space teardown never
/// reclaims these frames (the mapping carries `device_grant`); this object owns them. /// reclaims these frames (the mapping carries `device_grant`); this object owns them.
pub const ShmObject = struct { pub const SharedMemoryObject = struct {
refcount: u32 = 1, refcount: u32 = 1,
phys: u64, phys: u64,
pages: usize, pages: usize,
}; };
/// Wrap `pages` contiguous frames at `phys` (already allocated + zeroed by the caller) in a /// Wrap `pages` contiguous frames at `phys` (already allocated + zeroed by the caller) in a
/// refcounted shm object, or null if the heap is out of room. /// refcounted shared-memory object, or null if the heap is out of room.
pub fn createShm(phys: u64, pages: usize) ?*ShmObject { pub fn createSharedMemory(phys: u64, pages: usize) ?*SharedMemoryObject {
const shm = heap.allocator().create(ShmObject) catch return null; const shared_memory = heap.allocator().create(SharedMemoryObject) catch return null;
shm.* = .{ .phys = phys, .pages = pages }; shared_memory.* = .{ .phys = phys, .pages = pages };
return shm; return shared_memory;
} }
/// Drop a shared-memory reference; when the last one goes, return its frames to the /// Drop a shared-memory reference; when the last one goes, return its frames to the
/// allocator and free the object. (The mappings themselves are torn down with each /// allocator and free the object. (The mappings themselves are torn down with each
/// sharer's address space; `device_grant` keeps that from freeing the frames early.) /// sharer's address space; `device_grant` keeps that from freeing the frames early.)
pub fn dropShmRef(shm: *ShmObject) void { pub fn dropSharedMemoryReference(shared_memory: *SharedMemoryObject) void {
if (shm.refcount > 1) { if (shared_memory.refcount > 1) {
shm.refcount -= 1; shared_memory.refcount -= 1;
} else { } else {
for (0..shm.pages) |i| pmm.free(shm.phys + i * page_size); for (0..shared_memory.pages) |i| pmm.free(shared_memory.phys + i * page_size);
heap.allocator().destroy(shm); heap.allocator().destroy(shared_memory);
} }
} }
@@ -294,8 +294,8 @@ fn shareCapability(from: *Task, to: *Task, cap: u64) i64 {
const e: *Endpoint = @ptrCast(@alignCast(entry.ptr)); const e: *Endpoint = @ptrCast(@alignCast(entry.ptr));
e.refcount += 1; e.refcount += 1;
}, },
handle_kind_shm => { handle_kind_shared_memory => {
const s: *ShmObject = @ptrCast(@alignCast(entry.ptr)); const s: *SharedMemoryObject = @ptrCast(@alignCast(entry.ptr));
s.refcount += 1; s.refcount += 1;
}, },
else => return -EBADF, else => return -EBADF,
@@ -512,12 +512,26 @@ pub fn installHandle(t: *Task, endpoint: *Endpoint) i64 {
} }
/// Install a shared-memory handle. /// Install a shared-memory handle.
pub fn installShmHandle(t: *Task, shm: *ShmObject) i64 { pub fn installSharedMemoryHandle(t: *Task, shared_memory: *SharedMemoryObject) i64 {
return installEntry(t, .{ .kind = handle_kind_shm, .ptr = @ptrCast(shm) }); return installEntry(t, .{ .kind = handle_kind_shared_memory, .ptr = @ptrCast(shared_memory) });
}
/// Install an endpoint handle, reusing an existing slot that already names this
/// endpoint (no new reference taken in that case). For callers that install per
/// operation — fs_resolve — so a 16-slot table can't be exhausted by repeats.
/// Any subsystem installing handles per-call should come through here.
pub fn installHandleDeduped(t: *Task, endpoint: *Endpoint) i64 {
for (t.handles, 0..) |slot, i| {
const entry = slot orelse continue;
if (entry.kind == handle_kind_endpoint and entry.ptr == @as(*anyopaque, @ptrCast(endpoint))) return @intCast(i);
}
const h = installHandle(t, endpoint);
if (h >= 0) endpoint.refcount += 1; // the table entry owns a reference
return h;
} }
/// Resolve a handle to its endpoint, or null if out of range, unused, or a different kind /// Resolve a handle to its endpoint, or null if out of range, unused, or a different kind
/// (e.g. an shm handle used where an endpoint is expected). /// (e.g. a shared-memory handle used where an endpoint is expected).
pub fn resolveHandle(t: *Task, h: u64) ?*Endpoint { pub fn resolveHandle(t: *Task, h: u64) ?*Endpoint {
if (h >= t.handles.len) return null; if (h >= t.handles.len) return null;
const entry = t.handles[@intCast(h)] orelse return null; const entry = t.handles[@intCast(h)] orelse return null;
@@ -526,11 +540,11 @@ pub fn resolveHandle(t: *Task, h: u64) ?*Endpoint {
} }
/// Resolve a handle to its shared-memory object, or null if out of range, unused, or not /// Resolve a handle to its shared-memory object, or null if out of range, unused, or not
/// an shm handle. /// a shared-memory handle.
pub fn resolveShm(t: *Task, h: u64) ?*ShmObject { pub fn resolveSharedMemory(t: *Task, h: u64) ?*SharedMemoryObject {
if (h >= t.handles.len) return null; if (h >= t.handles.len) return null;
const entry = t.handles[@intCast(h)] orelse return null; const entry = t.handles[@intCast(h)] orelse return null;
if (entry.kind != handle_kind_shm) return null; if (entry.kind != handle_kind_shared_memory) return null;
return @ptrCast(@alignCast(entry.ptr)); return @ptrCast(@alignCast(entry.ptr));
} }
@@ -550,7 +564,7 @@ pub fn closeHandles(t: *Task) void {
fn dropEntry(entry: scheduler.HandleObject) void { fn dropEntry(entry: scheduler.HandleObject) void {
switch (entry.kind) { switch (entry.kind) {
handle_kind_endpoint => dropRef(@ptrCast(@alignCast(entry.ptr))), handle_kind_endpoint => dropRef(@ptrCast(@alignCast(entry.ptr))),
handle_kind_shm => dropShmRef(@ptrCast(@alignCast(entry.ptr))), handle_kind_shared_memory => dropSharedMemoryReference(@ptrCast(@alignCast(entry.ptr))),
else => {}, else => {},
} }
} }
+21 -18
View File
@@ -75,7 +75,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// Retain the whole stream in a RAM buffer too, so a user program can later // Retain the whole stream in a RAM buffer too, so a user program can later
// read it back (klog_read) and persist the boot log to disk — the only way to // read it back (klog_read) and persist the boot log to disk — the only way to
// see it on a headless/real machine with no host capturing serial. // see it on a headless/real machine with no host capturing serial.
log.addSink(log.ramSink); // (Retention is the tagged ring inside log.zig — not a sink.)
// The **framebuffer** is deliberately *not* a log sink. It's a separate output // The **framebuffer** is deliberately *not* a log sink. It's a separate output
// surface — a bootstrap text console today, a graphics device driver later — so // surface — a bootstrap text console today, a graphics device driver later — so
@@ -162,6 +162,13 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// uncached crawl). Routine boot output goes only to the log; this console now exists for // uncached crawl). Routine boot output goes only to the log; this console now exists for
// early-boot and fatal (`fatal`/panic) output, until the display service takes over. // early-boot and fatal (`fatal`/panic) output, until the display service takes over.
console.init(fb); console.init(fb);
// The console joins the log sinks: the boot transcript — kernel AND
// userspace lines, each timestamped by the renderer — shows on screen
// until the display service claims the framebuffer (which flips the
// console's `suppressed` and silences this sink). On a machine with no
// serial this is the only live view of the boot, and a slow boot becomes
// diagnosable by eye: the timeline is right there.
if (console.present()) log.addSink(console.write);
log.write(if (console.present()) log.write(if (console.present())
"/system/kernel: framebuffer ready (early-boot + fatal fallback; the display service drives it in normal operation)\n" "/system/kernel: framebuffer ready (early-boot + fatal fallback; the display service drives it in normal operation)\n"
else else
@@ -200,8 +207,8 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// Publish the loader's framebuffer as a claimable `display` device, so a // Publish the loader's framebuffer as a claimable `display` device, so a
// user-space display service can take it over the same claim + mmio_map path as // user-space display service can take it over the same claim + mmio_map path as
// any other hardware (it is not firmware-discovered; it rides the boot handoff). // any other hardware (it is not firmware-discovered; it rides the boot handoff).
if (devices_broker.seedDisplay(fb.base, fb.width, fb.height, fb.pitch, @intFromEnum(fb.format))) |display_id| { if (devices_broker.seedDisplay(fb.base, fb.width, fb.height, fb.pitch, @intFromEnum(fb.format), fb.refresh_hz)) |display_id| {
log.print("/system/kernel: framebuffer device {d} seeded ({d}x{d}, pitch {d}, write-combining)\n", .{ display_id, fb.width, fb.height, fb.pitch }); log.print("/system/kernel: framebuffer device {d} seeded ({d}x{d}, pitch {d}, {d} Hz, write-combining)\n", .{ display_id, fb.width, fb.height, fb.pitch, fb.refresh_hz });
} }
// Install the device-IRQ trampolines, so a driver's irq_bind has vectors to // Install the device-IRQ trampolines, so a driver's irq_bind has vectors to
@@ -329,20 +336,16 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// service supervisor and the device manager spawns the drivers it discovers. // service supervisor and the device manager spawns the drivers it discovers.
publishInitialRamdisk(boot_information); publishInitialRamdisk(boot_information);
// Hand over to user space: load /system/services/init (read off the boot volume by // Hand over to user space: spawn /system/services/init out of the ramdisk as a
// the loader) and spawn it as a real ring-3 process, PID 1. As the supervisor it // real ring-3 process, PID 1 — it rides the same table as every other binary.
// brings up the system services (the VFS server, the device manager); the device // As the supervisor it brings up the system services (the VFS server, the device
// manager then discovers the hardware and spawns each driver. init runs on its own // manager); the device manager then discovers the hardware and spawns each
// address space, preemptively — this boot context becomes the BSP's idle loop. // driver. init runs on its own address space, preemptively — this boot context
if (boot_information.init_len != 0) { // becomes the BSP's idle loop.
status("/system/kernel: starting /system/services/init...\n"); status("/system/kernel: starting /system/services/init...\n");
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; process.spawnBundled("/system/services/init") catch |err| {
process.spawnProcess(image, 4, &.{"/system/services/init"}) catch |err| { statusPrint("/system/kernel: /system/services/init failed to start: {s}\n", .{@errorName(err)});
statusPrint("/system/kernel: /system/services/init failed to load: {s}\n", .{@errorName(err)}); };
};
} else {
status("no /system/services/init on the boot volume.\n");
}
// Become the idle task: drop below every real task and halt until an // Become the idle task: drop below every real task and halt until an
// interrupt. The timer keeps preempting into init and any other work. // interrupt. The timer keeps preempting into init and any other work.
@@ -430,7 +433,7 @@ fn status(message: []const u8) void {
/// any display service holding the framebuffer. The console is otherwise silent in normal /// any display service holding the framebuffer. The console is otherwise silent in normal
/// operation (see `status`); it exists now only for early-boot and fatal output. /// operation (see `status`); it exists now only for early-boot and fatal output.
fn fatal(message: []const u8) void { fn fatal(message: []const u8) void {
log.write(message); log.appendPanic(message); // bounded lock wait: a panic never deadlocks on the log
console.setSuppressed(false); console.setSuppressed(false);
console.write(message); console.write(message);
} }
+234
View File
@@ -0,0 +1,234 @@
//! The tagged kernel log ring — a circular byte buffer of framed records, each
//! stamped by the writer (the kernel) with the sender's pid, task name, level,
//! per-boot sequence number, and monotonic timestamp. Pure code over an
//! embedded buffer — no architecture or lock imports — so it host-tests
//! alongside the other pure kernel pieces (`zig build test`).
//!
//! `head` and `tail` are free-running u64 positions in a logical byte stream;
//! the physical wrap is invisible to readers (all copies are modulo the
//! buffer), so a record never splits logically and no padding records exist.
//! Reclaim happens record by record: the writer parses the header at `tail`
//! (which it wrote itself) and advances until the new record fits — `tail`
//! always sits on a record boundary, and sequence-number gaps tell a reader
//! exactly how many records it lost.
//!
//! Locking is the caller's job (log.zig holds its log lock around every call);
//! the ring itself is single-writer, snapshot-reader.
const std = @import("std");
const abi = @import("abi");
pub fn Ring(comptime capacity: usize) type {
comptime std.debug.assert(std.math.isPowerOfTwo(capacity));
return struct {
const Self = @This();
buffer: [capacity]u8 = undefined,
head: u64 = 0,
tail: u64 = 0,
next_sequence: u64 = 0,
/// Append one record; returns its sequence number. `name` and `message`
/// are clamped to their ABI caps (the syscall clamps earlier too — the
/// clamp here makes the ring safe in isolation).
pub fn append(
self: *Self,
pid: u32,
name: []const u8,
level: abi.KlogLevel,
timestamp_ns: u64,
message: []const u8,
truncated: bool,
) u64 {
const name_len: usize = @min(name.len, abi.maximum_process_name);
const message_len: usize = @min(message.len, abi.klog_maximum_message);
const record_len = recordLength(name_len, message_len);
// Reclaim whole records until the new one fits.
while (self.head + record_len - self.tail > capacity) self.reclaimOne();
const sequence = self.next_sequence;
self.next_sequence += 1;
const header = abi.KlogRecordHeader{
.magic = abi.klog_record_magic,
.level = level,
.name_len = @intCast(name_len),
.pid = pid,
.sequence = sequence,
.timestamp_ns = timestamp_ns,
.message_len = @intCast(message_len),
.flags = if (truncated) abi.klog_flag_truncated else 0,
._reserved = @splat(0),
};
self.put(self.head, std.mem.asBytes(&header));
self.put(self.head + abi.klog_record_header_size, name[0..name_len]);
self.put(self.head + abi.klog_record_header_size + name_len, message[0..message_len]);
// The alignment pad is dead space; zero it so raw dumps stay tidy.
var pad = abi.klog_record_header_size + name_len + message_len;
while (pad < record_len) : (pad += 1)
self.buffer[@intCast((self.head + pad) % capacity)] = 0;
self.head += record_len;
return sequence;
}
/// Copy stream bytes beginning at `offset` into `out`. Returns null if
/// `offset` fell behind `tail` (overwritten) or lies past `head` — the
/// reader re-syncs from status(). 0 bytes means caught up.
pub fn read(self: *const Self, offset: u64, out: []u8) ?usize {
if (offset < self.tail or offset > self.head) return null;
const n: usize = @intCast(@min(out.len, self.head - offset));
self.get(offset, out[0..n]);
return n;
}
/// Cursors for klog_status. boot_unix_seconds is the kernel wrapper's
/// to fill — the ring knows nothing of wall clocks.
pub fn status(self: *const Self) abi.KlogStatus {
return .{
.tail = self.tail,
.head = self.head,
.next_sequence = self.next_sequence,
.boot_unix_seconds = 0,
};
}
fn reclaimOne(self: *Self) void {
var header_bytes: [abi.klog_record_header_size]u8 = undefined;
self.get(self.tail, &header_bytes);
const header = std.mem.bytesToValue(abi.KlogRecordHeader, &header_bytes);
// The writer wrote this header itself: the assert guards against
// memory corruption, not bad input.
std.debug.assert(header.magic == abi.klog_record_magic);
self.tail += recordLength(header.name_len, header.message_len);
}
fn recordLength(name_len: usize, message_len: usize) usize {
return std.mem.alignForward(usize, abi.klog_record_header_size + name_len + message_len, abi.klog_record_alignment);
}
// Byte-at-a-time modulo copies keep the wrap logic obviously correct;
// if they ever show in a profile, split into two @memcpy spans.
fn put(self: *Self, offset: u64, bytes: []const u8) void {
for (bytes, 0..) |b, i| self.buffer[@intCast((offset + i) % capacity)] = b;
}
fn get(self: *const Self, offset: u64, out: []u8) void {
for (out, 0..) |*b, i| b.* = self.buffer[@intCast((offset + i) % capacity)];
}
};
}
// --- tests (host) -----------------------------------------------------------
const TestRing = Ring(4096);
/// Parse the record at `offset` out of `ring`, returning the header plus name
/// and message copies — the same walk a userspace drainer performs.
const Parsed = struct {
header: abi.KlogRecordHeader,
name: [abi.maximum_process_name]u8 = undefined,
message: [abi.klog_maximum_message]u8 = undefined,
fn nameSlice(self: *const Parsed) []const u8 {
return self.name[0..self.header.name_len];
}
fn messageSlice(self: *const Parsed) []const u8 {
return self.message[0..self.header.message_len];
}
fn next(self: *const Parsed, offset: u64) u64 {
return offset + std.mem.alignForward(usize, abi.klog_record_header_size + self.header.name_len + self.header.message_len, abi.klog_record_alignment);
}
};
fn parseAt(ring: *const TestRing, offset: u64) Parsed {
var p: Parsed = undefined;
var header_bytes: [abi.klog_record_header_size]u8 = undefined;
std.debug.assert(ring.read(offset, &header_bytes).? == header_bytes.len);
p.header = std.mem.bytesToValue(abi.KlogRecordHeader, &header_bytes);
std.debug.assert(p.header.magic == abi.klog_record_magic);
_ = ring.read(offset + abi.klog_record_header_size, p.name[0..p.header.name_len]);
_ = ring.read(offset + abi.klog_record_header_size + p.header.name_len, p.message[0..p.header.message_len]);
return p;
}
test "header size is pinned" {
try std.testing.expectEqual(abi.klog_record_header_size, @sizeOf(abi.KlogRecordHeader));
}
test "append/read round trip" {
var ring = std.testing.allocator.create(TestRing) catch unreachable;
defer std.testing.allocator.destroy(ring);
ring.* = .{};
_ = ring.append(7, "/system/services/fat", .info, 123, "mounted /mnt/usb", false);
_ = ring.append(0, "kernel", .raw, 456, "wall clock online", false);
const first = parseAt(ring, ring.tail);
try std.testing.expectEqual(@as(u32, 7), first.header.pid);
try std.testing.expectEqual(abi.KlogLevel.info, first.header.level);
try std.testing.expectEqual(@as(u64, 123), first.header.timestamp_ns);
try std.testing.expectEqualStrings("/system/services/fat", first.nameSlice());
try std.testing.expectEqualStrings("mounted /mnt/usb", first.messageSlice());
const second = parseAt(ring, first.next(ring.tail));
try std.testing.expectEqual(@as(u32, 0), second.header.pid);
try std.testing.expectEqualStrings("kernel", second.nameSlice());
try std.testing.expectEqual(@as(u64, 1), second.header.sequence);
}
test "wrap reclaims whole records and keeps tail on a boundary" {
var ring = std.testing.allocator.create(TestRing) catch unreachable;
defer std.testing.allocator.destroy(ring);
ring.* = .{};
// Fill far past capacity so the ring wraps many times.
var i: u32 = 0;
while (i < 200) : (i += 1) {
var message: [64]u8 = undefined;
const m = std.fmt.bufPrint(&message, "line {d} padding padding padding", .{i}) catch unreachable;
_ = ring.append(1, "/system/tests/writer", .info, i, m, false);
}
try std.testing.expect(ring.head - ring.tail <= 4096);
// The record at tail parses cleanly (boundary held), and walking to head
// yields consecutive sequence numbers.
var offset = ring.tail;
var previous: ?u64 = null;
while (offset < ring.head) {
const p = parseAt(ring, offset);
if (previous) |q| try std.testing.expectEqual(q + 1, p.header.sequence);
previous = p.header.sequence;
offset = p.next(offset);
}
try std.testing.expectEqual(ring.head, offset);
// Records were lost (sequence at tail > 0), and the count is the gap.
try std.testing.expect(parseAt(ring, ring.tail).header.sequence > 0);
}
test "stale offset returns null; head offset reads zero bytes" {
var ring = std.testing.allocator.create(TestRing) catch unreachable;
defer std.testing.allocator.destroy(ring);
ring.* = .{};
var i: u32 = 0;
while (i < 300) : (i += 1)
_ = ring.append(1, "w", .info, i, "0123456789abcdef0123456789abcdef", false);
var out: [16]u8 = undefined;
try std.testing.expect(ring.read(0, &out) == null); // long overwritten
try std.testing.expect(ring.read(ring.head + 1, &out) == null); // past the end
try std.testing.expectEqual(@as(usize, 0), ring.read(ring.head, &out).?); // caught up
}
test "truncation flag and clamping" {
var ring = std.testing.allocator.create(TestRing) catch unreachable;
defer std.testing.allocator.destroy(ring);
ring.* = .{};
const long = "x" ** 300; // past klog_maximum_message
_ = ring.append(2, "w", .warn, 0, long, true);
const p = parseAt(ring, ring.tail);
try std.testing.expectEqual(@as(u16, abi.klog_maximum_message), p.header.message_len);
try std.testing.expect(p.header.flags & abi.klog_flag_truncated != 0);
}
+179 -39
View File
@@ -4,22 +4,37 @@
//! Output is a *diagnostic convenience, never a correctness dependency* — the //! Output is a *diagnostic convenience, never a correctness dependency* — the
//! kernel must boot and run correctly with zero output channels. So logging fans //! kernel must boot and run correctly with zero output channels. So logging fans
//! out to a set of registered **sinks**, each best-effort and self-guarding: the //! out to a set of registered **sinks**, each best-effort and self-guarding: the
//! serial UART, the 0xE9 debug console, and — later — a file on a ramdisk/USB/SSD. //! serial UART and the 0xE9 debug console. A message reaches whatever channels
//! A message reaches whatever channels exist; if none do, the kernel runs on, //! exist; if none do, the kernel runs on, silent but correct.
//! silent but correct. //!
//! Retention is the tagged RING (log-ring.zig): every emission becomes one
//! record per line, stamped with the sender's pid, task name (its binary path),
//! level, sequence number, and monotonic timestamp — attribution is structural,
//! stamped by the kernel, not a naming convention a process could forge. The
//! stamping is per LINE: an embedded '\n' ends the record, so a payload cannot
//! imitate another sender on the line that follows. Oldest records are
//! overwritten when the ring is full; sequence gaps make the loss countable.
//! `klog_read`/`klog_status` expose the stream to userspace (the logger service
//! drains it into per-process files once storage is up).
//!
//! Locking: a dedicated log spinlock, NOT the big kernel lock. `print` is
//! called both inside and outside BKL sections (and from ISRs), so the log
//! lock is taken with interrupts off and nothing inside it ever takes the BKL —
//! lock order is strictly BKL -> log lock, never the reverse. Panic paths use a
//! bounded try-acquire and fall back to sinks-only: a panic must never deadlock
//! on its own diagnostics.
//! //!
//! The **framebuffer is deliberately not a sink here.** It's a separate output //! The **framebuffer is deliberately not a sink here.** It's a separate output
//! surface (a bootstrap text console today, a graphics device driver later), so //! surface (a bootstrap text console today, a graphics device driver later), so
//! the log never assumes the machine is text-based. `main.zig` mirrors a few //! the log never assumes the machine is text-based. Two channels bypass the
//! user-facing status lines and panics to it explicitly; the verbose log does not. //! sink list because they must survive even a total-output failure:
//! //! `checkpoint` (a one-byte POST code) and `recordPanic` (a fixed breadcrumb).
//! No allocation: the sink table is fixed, so the log works before the heap is up
//! and inside a panic. Two channels don't go through the sink list because they
//! must survive even a total-output failure: `checkpoint` (a one-byte POST code)
//! and `recordPanic` (a breadcrumb in a fixed record).
const std = @import("std"); const std = @import("std");
const architecture = @import("architecture"); const architecture = @import("architecture");
const abi = @import("abi");
const log_ring = @import("log-ring.zig");
const wall_clock = @import("wall-clock.zig");
pub const SinkFn = *const fn ([]const u8) void; pub const SinkFn = *const fn ([]const u8) void;
@@ -36,42 +51,149 @@ pub fn addSink(sink: SinkFn) void {
} }
} }
/// Fan `bytes` out to every registered sink. // --- the log lock ------------------------------------------------------------
pub fn write(bytes: []const u8) void {
for (sinks[0..sink_count]) |sink| sink(bytes); var lock_held = std.atomic.Value(u32).init(0);
fn lockAcquire() u64 {
const flags = architecture.saveInterrupts();
while (lock_held.cmpxchgWeak(0, 1, .acquire, .monotonic) != null) std.atomic.spinLoopHint();
return flags;
} }
// --- the RAM sink: a retained copy of the whole diagnostic stream ------------ fn lockTryAcquire(spins: usize) ?u64 {
// const flags = architecture.saveInterrupts();
// A fixed in-image buffer that accumulates every logged byte, so a user program var i: usize = 0;
// (`log-flush`, and init at shutdown) can read it back through `klog_read` and while (i < spins) : (i += 1) {
// persist it to a file — the boot log survives on a headless/real machine that if (lock_held.cmpxchgWeak(0, 1, .acquire, .monotonic) == null) return flags;
// has no host capturing serial. It is a *sink like any other*: register it with std.atomic.spinLoopHint();
// `addSink(ramSink)` at boot. No allocation (works pre-heap and in a panic). }
// architecture.restoreInterrupts(flags);
// It fills linearly and stops when full: the earliest output — the most valuable return null;
// for diagnosing a boot — is kept, and the tail is still on the live serial sink. }
// 256 KiB comfortably holds a full boot plus a long run (a boot is ~15 KiB).
const ram_capacity = 256 * 1024; fn lockRelease(flags: u64) void {
var ram_buffer: [ram_capacity]u8 = undefined; lock_held.store(0, .release);
var ram_len: usize = 0; architecture.restoreInterrupts(flags);
}
/// The RAM sink. Best-effort and self-guarding like every sink: appends what fits // --- the ring + renderer -----------------------------------------------------
/// and silently drops the rest once full. (Concurrency matches the other sinks —
/// the dominant writer, debug_write, already holds the kernel lock; a rare torn /// 512 KiB: the tagged frames cost ~30% over the raw text, and the ring only
/// append on a kernel-internal line is an accepted diagnostic imperfection.) /// needs to cover the pre-mount backlog (a boot is ~15 KiB of text) — the
pub fn ramSink(bytes: []const u8) void { /// logger service tails it continuously once storage is up.
const n = @min(ram_buffer.len - ram_len, bytes.len); const ring_capacity = 512 * 1024;
if (n != 0) { var ring: log_ring.Ring(ring_capacity) = .{};
@memcpy(ram_buffer[ram_len..][0..n], bytes[0..n]);
ram_len += n; /// Renderer state: whether the sinks sit at a line start, and which pid's line
/// is currently open — when a different sender interleaves mid-line, the
/// renderer closes the line so serial output can't visually merge two senders.
var at_line_start: bool = true;
var open_line_pid: u32 = 0;
/// Append `bytes` as one tagged record per line and render them to the sinks.
/// The core emission path: `debug_write` calls this with the sender's identity;
/// kernel-internal `write`/`print` funnel here as pid 0 ("kernel", raw).
pub fn append(pid: u32, name: []const u8, level: abi.KlogLevel, bytes: []const u8) void {
if (bytes.len == 0) return;
const now = architecture.nanos();
const flags = lockAcquire();
defer lockRelease(flags);
appendLocked(pid, name, level, now, bytes);
}
/// The panic-safe variant: bounded lock wait; on failure, sinks only — the ring
/// entry is lost but the message still reaches serial, and the panic cannot
/// deadlock on a core that died holding the log lock.
pub fn appendPanic(bytes: []const u8) void {
if (lockTryAcquire(100_000)) |flags| {
defer lockRelease(flags);
appendLocked(0, "kernel", .raw, architecture.nanos(), bytes);
} else {
for (sinks[0..sink_count]) |sink| sink(bytes);
} }
} }
/// The accumulated log so far — what `klog_read` copies out. fn appendLocked(pid: u32, name: []const u8, level: abi.KlogLevel, now: u64, bytes: []const u8) void {
pub fn ramSnapshot() []const u8 { var rest = bytes;
return ram_buffer[0..ram_len]; while (rest.len != 0) {
const newline = std.mem.indexOfScalar(u8, rest, '\n');
// The record payload excludes the newline: a record IS a line. Raw
// emissions may leave a line open (kernel boot tables build lines from
// pieces); a LEVELED record is a complete line by contract — std.log
// payloads carry no trailing newline.
const line = if (newline) |i| rest[0..i] else rest;
const line_complete = newline != null or level != .raw;
if (line.len != 0 or line_complete)
_ = ring.append(pid, name, level, now, line, line.len > abi.klog_maximum_message);
render(pid, name, level, now, line, line_complete);
rest = if (newline) |i| rest[i + 1 ..] else rest[rest.len..];
}
}
/// Serial/debugcon rendering. Kernel output and legacy raw user output pass
/// through byte-identical to the historical stream (services still write their
/// own "name: " prefixes until the std.log migration). Leveled (std.log)
/// records get a kernel-rendered "<name>: " prefix at line start — err/warn/
/// debug also get their level spelled out.
fn render(pid: u32, name: []const u8, level: abi.KlogLevel, now: u64, line: []const u8, line_complete: bool) void {
if (sink_count == 0) return;
if (line.len == 0 and !line_complete) return;
// Compose the whole rendered piece first and emit it in ONE sink call per
// sink: fewer, larger UART writes, and no partial-line window should any
// path ever reach a sink without the log lock.
var buffer: [render_buffer_size]u8 = undefined;
var used: usize = 0;
if (!at_line_start and open_line_pid != pid) {
buffer[used] = '\n';
used += 1;
at_line_start = true;
}
if (at_line_start) {
// Every line starts with its boot-relative time: the live transcript
// (serial AND the on-screen boot console) is a readable timeline —
// which is how a slow real-hardware boot gets diagnosed by eye.
const seconds = now / 1_000_000_000;
const millis = (now / 1_000_000) % 1000;
used += (std.fmt.bufPrint(buffer[used..], "[{d:>4}.{d:0>3}] ", .{ seconds, millis }) catch buffer[used..used]).len;
}
if (at_line_start and level != .raw) {
used += place(buffer[used..], name);
used += place(buffer[used..], ": ");
used += place(buffer[used..], switch (level) {
.err => "error: ",
.warn => "warning: ",
.debug => "debug: ",
.info, .raw => "",
});
}
used += place(buffer[used..], line);
if (line_complete and used < buffer.len) {
buffer[used] = '\n';
used += 1;
}
fanOut(buffer[0..used]);
at_line_start = line_complete;
open_line_pid = pid;
}
/// newline + timestamp + name + ": warning: " + a full payload line + newline.
const render_buffer_size = 1 + 16 + abi.maximum_process_name + 11 + abi.klog_maximum_message + 1;
fn place(destination: []u8, bytes: []const u8) usize {
const n = @min(destination.len, bytes.len);
@memcpy(destination[0..n], bytes[0..n]);
return n;
}
fn fanOut(bytes: []const u8) void {
for (sinks[0..sink_count]) |sink| sink(bytes);
}
/// Kernel-internal write — a raw record from "kernel" (pid 0). The signature is
/// unchanged so every existing kernel call site stays as it is.
pub fn write(bytes: []const u8) void {
append(0, "kernel", .raw, bytes);
} }
/// A formatted log line. Truncates past 256 bytes; the buffer is on the stack, so /// A formatted log line. Truncates past 256 bytes; the buffer is on the stack, so
@@ -81,6 +203,24 @@ pub fn print(comptime fmt: []const u8, args: anytype) void {
write(std.fmt.bufPrint(&buffer, fmt, args) catch return); write(std.fmt.bufPrint(&buffer, fmt, args) catch return);
} }
/// klog_read: copy ring stream bytes from `offset` into `out`. Null when the
/// cursor was overwritten or lies past the end — the reader re-syncs via
/// status(). Zero bytes means caught up.
pub fn readAt(offset: u64, out: []u8) ?usize {
const flags = lockAcquire();
defer lockRelease(flags);
return ring.read(offset, out);
}
/// klog_status: the ring cursors plus the boot wall-clock anchor.
pub fn status() abi.KlogStatus {
const flags = lockAcquire();
defer lockRelease(flags);
var s = ring.status();
s.boot_unix_seconds = wall_clock.bootSeconds();
return s;
}
/// Emit a one-byte checkpoint/POST code (I/O port 0x80) — the always-available /// Emit a one-byte checkpoint/POST code (I/O port 0x80) — the always-available
/// progress channel for when there is no text output at all. Independent of the /// progress channel for when there is no text output at all. Independent of the
/// sink list, so it works even before any sink is registered. /// sink list, so it works even before any sink is registered.
+229 -69
View File
@@ -34,6 +34,7 @@ const ipc = @import("ipc-synchronous.zig");
const devices_broker = @import("devices-broker.zig"); const devices_broker = @import("devices-broker.zig");
const irq = @import("irq.zig"); const irq = @import("irq.zig");
const initial_ramdisk = @import("initial-ramdisk"); const initial_ramdisk = @import("initial-ramdisk");
const vfs = @import("vfs.zig");
const log = @import("log.zig"); const log = @import("log.zig");
const wall_clock = @import("wall-clock.zig"); const wall_clock = @import("wall-clock.zig");
@@ -82,17 +83,17 @@ pub const device_arena_end: u64 = device_arena_base + (4 << 30);
pub const dma_arena_base: u64 = 0x0000_7200_0000_0000; pub const dma_arena_base: u64 = 0x0000_7200_0000_0000;
pub const dma_arena_end: u64 = dma_arena_base + (256 << 20); // 256 MiB per process pub const dma_arena_end: u64 = dma_arena_base + (256 << 20); // 256 MiB per process
/// The shared-memory arena: where `shm_create`/`shm_map` place shared cacheable regions, in /// The shared-memory arena: where `shared_memory_create`/`shared_memory_map` place shared cacheable regions, in
/// PML4[230] — a user-exclusive region distinct from the DMA arena. The frames are owned by /// PML4[230] — a user-exclusive region distinct from the DMA arena. The frames are owned by
/// a refcounted shm object and freed when its last capability drops, not on teardown, so the /// a refcounted shared-memory object and freed when its last capability drops, not on teardown, so the
/// mapping carries `device_grant`. Per-process cursor in `Task.shm_map_next` (docs/display-v2.md). /// mapping carries `device_grant`. Per-process cursor in `Task.shared_memory_map_next` (docs/display-v2.md).
pub const shm_arena_base: u64 = 0x0000_7300_0000_0000; pub const shared_memory_arena_base: u64 = 0x0000_7300_0000_0000;
pub const shm_arena_end: u64 = shm_arena_base + (256 << 20); // 256 MiB per process pub const shared_memory_arena_end: u64 = shared_memory_arena_base + (256 << 20); // 256 MiB per process
/// Largest single `shm_create`, in pages (32 MiB) — enough for a 4K framebuffer surface; /// Largest single `shared_memory_create`, in pages (32 MiB) — enough for a 4K framebuffer surface;
/// also an overflow guard on the page count. shm frames are contiguous (like DMA), so this /// also an overflow guard on the page count. shared-memory frames are contiguous (like DMA), so this
/// bounds the contiguous allocation asked of the frame allocator. /// bounds the contiguous allocation asked of the frame allocator.
const maximum_shm_pages = 8192; const maximum_shared_memory_pages = 8192;
/// Largest single `mmap` grant, in pages (32 MiB). Big enough for a display service's /// Largest single `mmap` grant, in pages (32 MiB). Big enough for a display service's
/// back buffer at up to 4K (3840x2160x4 ≈ 8100 pages); the user heap otherwise grows in /// back buffer at up to 4K (3840x2160x4 ≈ 8100 pages); the user heap otherwise grows in
@@ -138,6 +139,17 @@ var ramdisk_image: ?[]const u8 = null;
/// handoff) so a user-space supervisor can `system_spawn` binaries out of it. /// handoff) so a user-space supervisor can `system_spawn` binaries out of it.
pub fn setInitialRamdisk(image: []const u8) void { pub fn setInitialRamdisk(image: []const u8) void {
ramdisk_image = image; ramdisk_image = image;
vfs.setInitialRamdisk(image); // the kernel VFS serves the same bytes at /system
}
/// Spawn a bundled binary from the kernel by path. Used exactly once, to start
/// /system/services/init (PID 1) — every other spawn goes through the
/// `system_spawn` syscall.
pub fn spawnBundled(name: []const u8) !void {
const image = ramdisk_image orelse return error.NoInitialRamdisk;
const rd = initial_ramdisk.Reader.init(image) orelse return error.BadInitialRamdisk;
const item = rd.find(name) orelse return error.NotBundled;
try spawnProcess(item.blob, 4, &.{item.name});
} }
/// The system_call surface, dispatched on the saved system_call number (`abi.SystemCall`). /// The system_call surface, dispatched on the saved system_call number (`abi.SystemCall`).
@@ -180,9 +192,12 @@ fn system_call(state: *architecture.CpuState) void {
.exit => { .exit => {
exit_code = architecture.systemCallArg(state, 0); exit_code = architecture.systemCallArg(state, 0);
// A scheduled process tears down fully (terminateCurrent); a borrowed // A scheduled process tears down fully (terminateCurrent); a borrowed
// test thread unwinds back to the kernel that entered it. // test thread unwinds back to the kernel that entered it. A NONZERO
// code is a deliberate failure exit (`.aborted`): "the work exists
// but I could not do it" — supervisors restart those, unlike a clean
// `.exited` ("nothing for me here"), which they let lie.
if (scheduler.currentIsUserProcess()) { if (scheduler.currentIsUserProcess()) {
scheduler.current().exit_reason = .exited; scheduler.current().exit_reason = if (exit_code == 0) .exited else .aborted;
terminateCurrent(); terminateCurrent();
} else architecture.userExit(); } else architecture.userExit();
}, },
@@ -224,10 +239,15 @@ fn system_call(state: *architecture.CpuState) void {
.process_signal => systemProcessSignal(state), .process_signal => systemProcessSignal(state),
.timer_bind => systemTimerBind(state), .timer_bind => systemTimerBind(state),
.klog_read => systemKlogRead(state), .klog_read => systemKlogRead(state),
.klog_status => systemKlogStatus(state),
.fs_resolve => systemFsResolve(state),
.fs_node => systemFsNode(state),
.fs_mount => systemFsMount(state),
.fs_unmount => systemFsUnmount(state),
.wall_clock => systemWallClock(state), .wall_clock => systemWallClock(state),
.shm_create => systemShmCreate(state), .shared_memory_create => systemSharedMemoryCreate(state),
.shm_map => systemShmMap(state), .shared_memory_map => systemSharedMemoryMap(state),
.shm_physical => systemShmPhysical(state), .shared_memory_physical => systemSharedMemoryPhysical(state),
.thread_spawn => systemThreadSpawn(state), .thread_spawn => systemThreadSpawn(state),
.current_core => systemCurrentCore(state), .current_core => systemCurrentCore(state),
.thread_self => systemThreadSelf(state), .thread_self => systemThreadSelf(state),
@@ -513,79 +533,79 @@ fn systemDmaFree(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, 0); architecture.setSystemCallResult(state, 0);
} }
/// shm_create(len) -> virtual_address (rax), handle (rdx): grant `len` bytes (rounded up to whole /// shared_memory_create(len) -> virtual_address (rax), handle (rdx): grant `len` bytes (rounded up to whole
/// pages) of **shareable, zeroed, cacheable** RAM — contiguous frames mapped into the /// pages) of **shareable, zeroed, cacheable** RAM — contiguous frames mapped into the
/// caller's shm arena — and hand back the virtual address plus a capability handle. Unlike /// caller's shared-memory arena — and hand back the virtual address plus a capability handle. Unlike
/// `dma_alloc` the memory is write-back cacheable (for CPU compositing, not device DMA) and /// `dma_alloc` the memory is write-back cacheable (for CPU compositing, not device DMA) and
/// its frames are owned by a refcounted object: the handle is passed to another process as /// its frames are owned by a refcounted object: the handle is passed to another process as
/// an `ipc_call` send_cap, that process `shm_map`s it, and the frames free only when the /// an `ipc_call` send_cap, that process `shared_memory_map`s it, and the frames free only when the
/// last capability drops (docs/display-v2.md — the compositor↔native-driver and /// last capability drops (docs/display-v2.md — the compositor↔native-driver and
/// app↔compositor surface path). /// app↔compositor surface path).
fn systemShmCreate(state: *architecture.CpuState) void { fn systemSharedMemoryCreate(state: *architecture.CpuState) void {
const len = architecture.systemCallArg(state, 0); const len = architecture.systemCallArg(state, 0);
const t = scheduler.current(); const t = scheduler.current();
if (t.address_space == 0 or len == 0) return fail(state); if (t.address_space == 0 or len == 0) return fail(state);
const pages: usize = @intCast((len + page_size - 1) / page_size); const pages: usize = @intCast((len + page_size - 1) / page_size);
if (pages == 0 or pages > maximum_shm_pages) return fail(state); if (pages == 0 or pages > maximum_shared_memory_pages) return fail(state);
// Reserve arena virtual space up front, so a mapping failure needs no rollback. // Reserve arena virtual space up front, so a mapping failure needs no rollback.
if (t.shm_map_next == 0) t.shm_map_next = shm_arena_base; if (t.shared_memory_map_next == 0) t.shared_memory_map_next = shared_memory_arena_base;
const base_v = t.shm_map_next; const base_v = t.shared_memory_map_next;
if (base_v + pages * page_size > shm_arena_end) return fail(state); // arena exhausted if (base_v + pages * page_size > shared_memory_arena_end) return fail(state); // arena exhausted
const phys = pmm.allocContiguous(pages, ~@as(u64, 0)) orelse return fail(state); const phys = pmm.allocContiguous(pages, ~@as(u64, 0)) orelse return fail(state);
// Zero through the physmap (the frames aren't mapped in the caller yet). // Zero through the physmap (the frames aren't mapped in the caller yet).
const kernel_view: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(phys)); const kernel_view: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(phys));
@memset(kernel_view[0 .. pages * page_size], 0); @memset(kernel_view[0 .. pages * page_size], 0);
const shm = ipc.createShm(phys, pages) orelse { const shared_memory = ipc.createSharedMemory(phys, pages) orelse {
for (0..pages) |i| pmm.free(phys + i * page_size); for (0..pages) |i| pmm.free(phys + i * page_size);
return fail(state); return fail(state);
}; };
const handle = ipc.installShmHandle(t, shm); const handle = ipc.installSharedMemoryHandle(t, shared_memory);
if (handle < 0) { if (handle < 0) {
ipc.dropShmRef(shm); // last ref: frees the object and its frames ipc.dropSharedMemoryReference(shared_memory); // last ref: frees the object and its frames
return fail(state); return fail(state);
} }
architecture.mapUserSharedInto(t.address_space, base_v, phys, pages * page_size); architecture.mapUserSharedInto(t.address_space, base_v, phys, pages * page_size);
t.shm_map_next = base_v + pages * page_size; t.shared_memory_map_next = base_v + pages * page_size;
architecture.setSystemCallResult(state, base_v); // virtual_address for the CPU architecture.setSystemCallResult(state, base_v); // virtual_address for the CPU
architecture.setSystemCallResult2(state, @intCast(handle)); // capability handle to pass on architecture.setSystemCallResult2(state, @intCast(handle)); // capability handle to pass on
} }
/// shm_map(cap) -> virtual_address: map the shared region named by a capability handle the caller /// shared_memory_map(cap) -> virtual_address: map the shared region named by a capability handle the caller
/// received (via an `ipc_call` send_cap) into its shm arena — the same physical frames the /// received (via an `ipc_call` send_cap) into its shared-memory arena — the same physical frames the
/// creator sees — returning the virtual address. The handle already holds a reference (taken /// creator sees — returning the virtual address. The handle already holds a reference (taken
/// when the capability was shared), so this only adds a mapping; it never bumps the refcount. /// when the capability was shared), so this only adds a mapping; it never bumps the refcount.
fn systemShmMap(state: *architecture.CpuState) void { fn systemSharedMemoryMap(state: *architecture.CpuState) void {
const cap = architecture.systemCallArg(state, 0); const cap = architecture.systemCallArg(state, 0);
const t = scheduler.current(); const t = scheduler.current();
if (t.address_space == 0) return fail(state); if (t.address_space == 0) return fail(state);
const shm = ipc.resolveShm(t, cap) orelse return fail(state); // not an shm handle we hold const shared_memory = ipc.resolveSharedMemory(t, cap) orelse return fail(state); // not a shared-memory handle we hold
if (t.shm_map_next == 0) t.shm_map_next = shm_arena_base; if (t.shared_memory_map_next == 0) t.shared_memory_map_next = shared_memory_arena_base;
const base_v = t.shm_map_next; const base_v = t.shared_memory_map_next;
const size = shm.pages * page_size; const size = shared_memory.pages * page_size;
if (base_v + size > shm_arena_end) return fail(state); if (base_v + size > shared_memory_arena_end) return fail(state);
architecture.mapUserSharedInto(t.address_space, base_v, shm.phys, size); architecture.mapUserSharedInto(t.address_space, base_v, shared_memory.phys, size);
t.shm_map_next = base_v + size; t.shared_memory_map_next = base_v + size;
architecture.setSystemCallResult(state, base_v); architecture.setSystemCallResult(state, base_v);
} }
/// shm_physical(cap) -> physical_address: the guest-physical base of a shared region the caller holds a /// shared_memory_physical(cap) -> physical_address: the guest-physical base of a shared region the caller holds a
/// capability for. The frames are contiguous (allocated by `allocContiguous`), so a single /// capability for. The frames are contiguous (allocated by `allocContiguous`), so a single
/// physical base + length describes the whole region — which is exactly what a driver needs /// physical base + length describes the whole region — which is exactly what a driver needs
/// to hand a shm surface to a device (virtio-gpu `attach_backing`). Only a holder of the /// to hand a shared-memory surface to a device (virtio-gpu `attach_backing`). Only a holder of the
/// capability can ask; there is no ambient way to turn a virtual address into a physical one. /// capability can ask; there is no ambient way to turn a virtual address into a physical one.
fn systemShmPhysical(state: *architecture.CpuState) void { fn systemSharedMemoryPhysical(state: *architecture.CpuState) void {
const cap = architecture.systemCallArg(state, 0); const cap = architecture.systemCallArg(state, 0);
const t = scheduler.current(); const t = scheduler.current();
if (t.address_space == 0) return fail(state); if (t.address_space == 0) return fail(state);
const shm = ipc.resolveShm(t, cap) orelse return fail(state); // not an shm handle we hold const shared_memory = ipc.resolveSharedMemory(t, cap) orelse return fail(state); // not a shared-memory handle we hold
architecture.setSystemCallResult(state, shm.phys); architecture.setSystemCallResult(state, shared_memory.phys);
} }
/// device_register(parent_id, descriptor_ptr) -> id: publish a child device below a device /// device_register(parent_id, descriptor_ptr) -> id: publish a child device below a device
@@ -656,8 +676,11 @@ fn systemSpawn(state: *architecture.CpuState) void {
const rd = initial_ramdisk.Reader.init(image) orelse return fail(state); const rd = initial_ramdisk.Reader.init(image) orelse return fail(state);
const name = @as([*]const u8, @ptrFromInt(ptr))[0..len]; const name = @as([*]const u8, @ptrFromInt(ptr))[0..len];
// Exact path first, basename fallback second; either way argv[0] (and hence
// the task name, and the log ring's attribution) is the stored full path.
const item = rd.find(name) orelse return fail(state); // no bundled binary by that name
var argv: [maximum_arguments][]const u8 = undefined; var argv: [maximum_arguments][]const u8 = undefined;
argv[0] = name; argv[0] = item.name;
var argc: usize = 1; var argc: usize = 1;
if (arguments_len != 0) { if (arguments_len != 0) {
const blob = @as([*]const u8, @ptrFromInt(arguments_ptr))[0..arguments_len]; const blob = @as([*]const u8, @ptrFromInt(arguments_ptr))[0..arguments_len];
@@ -669,15 +692,8 @@ fn systemSpawn(state: *architecture.CpuState) void {
} }
} }
var i: u32 = 0; const child = spawnProcessSupervised(item.blob, 4, argv[0..argc], t.id, exit_endpoint) catch return fail(state);
while (i < rd.count) : (i += 1) { architecture.setSystemCallResult(state, child);
const item = rd.entry(i) orelse continue;
if (!std.mem.eql(u8, item.name, name)) continue;
const child = spawnProcessSupervised(item.blob, 4, argv[0..argc], t.id, exit_endpoint) catch return fail(state);
architecture.setSystemCallResult(state, child);
return;
}
fail(state); // no bundled binary by that name
} }
/// thread_spawn(entry, stack_top, arg) -> tid: start a task that shares the **caller's** /// thread_spawn(entry, stack_top, arg) -> tid: start a task that shares the **caller's**
@@ -1204,7 +1220,6 @@ fn systemIrqAck(state: *architecture.CpuState) void {
/// Whether the debug_write stream sits at the start of a line — the last emitted /// Whether the debug_write stream sits at the start of a line — the last emitted
/// byte was a newline (true at boot: nothing emitted yet). Guarded by the kernel /// byte was a newline (true at boot: nothing emitted yet). Guarded by the kernel
/// lock in `systemDebugWrite`, like the stream it describes. /// lock in `systemDebugWrite`, like the stream it describes.
var write_at_line_start: bool = true;
/// debug_write(ptr, len): copy bytes from user memory into the kernel log. /// debug_write(ptr, len): copy bytes from user memory into the kernel log.
/// A bring-up diagnostic — real output goes through the VFS/console later. /// A bring-up diagnostic — real output goes through the VFS/console later.
@@ -1225,31 +1240,42 @@ var write_at_line_start: bool = true;
fn systemDebugWrite(state: *architecture.CpuState) void { fn systemDebugWrite(state: *architecture.CpuState) void {
const ptr = architecture.systemCallArg(state, 0); const ptr = architecture.systemCallArg(state, 0);
const len = architecture.systemCallArg(state, 1); const len = architecture.systemCallArg(state, 1);
const level_raw = architecture.systemCallArg(state, 2);
if (len <= write_buffer.len and ptr < user_half_end and ptr + len <= user_half_end) { if (len <= write_buffer.len and ptr < user_half_end and ptr + len <= user_half_end) {
const source: [*]const u8 = @ptrFromInt(ptr); const source: [*]const u8 = @ptrFromInt(ptr);
// Levels above the enum range clamp to raw — old two-arg callers land
// there naturally (garbage in arg 2 stays harmless).
const level: abi.KlogLevel = if (level_raw <= @intFromEnum(abi.KlogLevel.raw))
@enumFromInt(level_raw)
else
.raw;
const t = scheduler.current();
const flags = sync.enter(); const flags = sync.enter();
defer sync.leave(flags); defer sync.leave(flags);
@memcpy(write_buffer[0..len], source[0..len]); // keep the latest message @memcpy(write_buffer[0..len], source[0..len]); // keep the latest message
write_len = len; write_len = len;
write_from_user = architecture.fromUser(state); write_from_user = architecture.fromUser(state);
write_count += 1; write_count += 1;
log.write(source[0..len]); // The kernel stamps the sender's identity — attribution is structural,
if (len != 0) write_at_line_start = source[len - 1] == '\n'; // not a prefix convention the payload could forge (and it is stamped
// per line inside log.append).
log.append(t.id, t.name(), level, source[0..len]);
architecture.setSystemCallResult(state, len); architecture.setSystemCallResult(state, len);
} else { } else {
fail(state); fail(state);
} }
} }
/// klog_read(offset, ptr, len) -> bytes copied: copy the kernel's in-memory /// klog_read(offset, ptr, len) -> bytes copied: copy tagged log-ring stream
/// diagnostic log (the RAM sink in log.zig) out to the user buffer at `ptr`, /// bytes beginning at stream offset `offset` out to the user buffer at `ptr`.
/// starting at `offset`. Returns the count copied — 0 once `offset` reaches the /// Returns the count copied — 0 means caught up — and fails once `offset` has
/// end — so a program reads the whole log by looping from 0 until it gets 0. /// fallen behind the ring's tail (the records were overwritten) or lies past
/// its head; the reader re-syncs via klog_status. A reader parses
/// [KlogRecordHeader][name][message] frames out of the byte stream (abi.zig).
/// ///
/// The mirror of `debug_write`: the same overflow-safe user-half bounds check, /// The mirror of `debug_write`: the same overflow-safe user-half bounds check,
/// but the copy runs kernel -> user. Written under the kernel lock so the source /// but the copy runs kernel -> user, under the log lock (inside log.readAt) so
/// snapshot can't grow underneath the copy. A read-only diagnostic — it exposes /// the stream can't move underneath the copy. A read-only diagnostic.
/// only the log the kernel already broadcasts to serial, nothing else.
fn systemKlogRead(state: *architecture.CpuState) void { fn systemKlogRead(state: *architecture.CpuState) void {
const offset = architecture.systemCallArg(state, 0); const offset = architecture.systemCallArg(state, 0);
const ptr = architecture.systemCallArg(state, 1); const ptr = architecture.systemCallArg(state, 1);
@@ -1257,21 +1283,155 @@ fn systemKlogRead(state: *architecture.CpuState) void {
// Confine the whole destination span to the user (low) half. `len <= // Confine the whole destination span to the user (low) half. `len <=
// user_half_end - ptr` bounds the length without an overflowing add. // user_half_end - ptr` bounds the length without an overflowing add.
if (ptr < user_half_end and len <= user_half_end - ptr) { if (ptr < user_half_end and len <= user_half_end - ptr) {
const flags = sync.enter(); const dest: [*]u8 = @ptrFromInt(ptr);
defer sync.leave(flags); const n = log.readAt(offset, dest[0..len]) orelse return fail(state);
const snapshot = log.ramSnapshot();
var n: usize = 0;
if (offset < snapshot.len) {
n = @min(len, snapshot.len - offset);
const dest: [*]u8 = @ptrFromInt(ptr);
@memcpy(dest[0..n], snapshot[offset..][0..n]);
}
architecture.setSystemCallResult(state, n); architecture.setSystemCallResult(state, n);
} else { } else {
fail(state); fail(state);
} }
} }
/// klog_status(ptr) -> 0: copy a KlogStatus — the ring's live cursors plus the
/// boot wall-clock anchor — out to the user buffer at `ptr`. How a log reader
/// finds the oldest retained offset, detects lost records (sequence gaps), and
/// names a per-boot log directory (boot_unix_seconds).
fn systemKlogStatus(state: *architecture.CpuState) void {
const ptr = architecture.systemCallArg(state, 0);
const size = @sizeOf(abi.KlogStatus);
if (ptr < user_half_end and size <= user_half_end - ptr) {
var status = log.status();
const dest: [*]u8 = @ptrFromInt(ptr);
@memcpy(dest[0..size], std.mem.asBytes(&status)[0..size]);
architecture.setSystemCallResult(state, 0);
} else {
fail(state);
}
}
/// fs_resolve(path_ptr, path_len, flags, out_ptr, out_cap): route a path
/// through the kernel mount table (docs/vfs-protocol.md). Kernel-served ->
/// rax=fs_route_kernel, rdx=node token. Backend-served -> rax=fs_route_backend,
/// rdx=an endpoint handle in the caller's table (deduplicated), and the
/// rewritten mount-relative path copied into `out` behind a u16 length prefix. Fails for unknown paths, create-intent on /system, or an
/// undersized out buffer.
fn systemFsResolve(state: *architecture.CpuState) void {
const path_ptr = architecture.systemCallArg(state, 0);
const path_len = architecture.systemCallArg(state, 1);
const flags = architecture.systemCallArg(state, 2);
const out_ptr = architecture.systemCallArg(state, 3);
const out_cap = architecture.systemCallArg(state, 4);
if (path_len == 0 or path_len > 224 or path_ptr >= user_half_end or path_ptr + path_len > user_half_end) return fail(state);
if (out_cap != 0 and (out_ptr >= user_half_end or out_ptr + out_cap > user_half_end)) return fail(state);
const path = @as([*]const u8, @ptrFromInt(path_ptr))[0..path_len];
const t = scheduler.current();
const flags_lock = sync.enter();
defer sync.leave(flags_lock);
switch (vfs.resolvePath(path, flags & abi.fs_flag_create != 0)) {
.kernel_node => |node_token| {
architecture.setSystemCallResult(state, abi.fs_route_kernel);
architecture.setSystemCallResult2(state, node_token);
},
.backend => |*backend| {
// The rewritten path goes back in the out buffer behind a u16
// length prefix (a third result register would collide with r8's
// argument role in the userspace stub).
if (backend.path_len + 2 > out_cap) return fail(state);
const handle = ipc.installHandleDeduped(t, backend.endpoint);
if (handle < 0) return fail(state);
const destination: [*]u8 = @ptrFromInt(out_ptr);
destination[0] = @intCast(backend.path_len & 0xFF);
destination[1] = @intCast(backend.path_len >> 8);
@memcpy(destination[2..][0..backend.path_len], backend.path[0..backend.path_len]);
architecture.setSystemCallResult(state, abi.fs_route_backend);
architecture.setSystemCallResult2(state, @intCast(handle));
},
.not_found => fail(state),
}
}
/// fs_node(op, node_token, offset, buf_ptr, buf_len) -> bytes/0/-errno: serve a
/// kernel-backed node. read copies file bytes; status copies a FileAttributes;
/// readdir copies [DirectoryEntryHeader][name] for the `offset`th child. Reads
/// of the immutable initrd never take the kernel lock.
fn systemFsNode(state: *architecture.CpuState) void {
const operation = architecture.systemCallArg(state, 0);
const node_token = architecture.systemCallArg(state, 1);
const offset = architecture.systemCallArg(state, 2);
const buf_ptr = architecture.systemCallArg(state, 3);
const buf_len = architecture.systemCallArg(state, 4);
if (buf_ptr >= user_half_end or buf_len > user_half_end - buf_ptr) return fail(state);
const capped = @min(buf_len, 64 * 1024); // bound any single copy
const destination: [*]u8 = @ptrFromInt(buf_ptr);
switch (operation) {
abi.fs_node_read => {
const n = vfs.nodeRead(node_token, offset, destination[0..capped]) orelse return fail(state);
architecture.setSystemCallResult(state, n);
},
abi.fs_node_status => {
var attributes = vfs.nodeStatus(node_token) orelse return fail(state);
if (capped < @sizeOf(abi.FileAttributes)) return fail(state);
@memcpy(destination[0..@sizeOf(abi.FileAttributes)], std.mem.asBytes(&attributes));
architecture.setSystemCallResult(state, @sizeOf(abi.FileAttributes));
},
abi.fs_node_readdir => {
const header_size = @sizeOf(abi.DirectoryEntryHeader);
if (capped < header_size) return fail(state);
var name_buffer: [64]u8 = undefined;
const result = vfs.nodeReaddir(node_token, offset, &name_buffer) orelse {
architecture.setSystemCallResult(state, 0); // past the end
return;
};
var header = result.header;
const total = header_size + @min(result.name_len, capped - header_size);
@memcpy(destination[0..header_size], std.mem.asBytes(&header));
@memcpy(destination[header_size..total], name_buffer[0 .. total - header_size]);
architecture.setSystemCallResult(state, total);
},
else => fail(state),
}
}
/// fs_mount(prefix_ptr, prefix_len, backend_handle, rewrite_ptr, rewrite_len):
/// mount a userspace filesystem at an absolute prefix. Possession of the
/// backend endpoint handle is the capability — the same trust as the old
/// router's cap-passing mount. The mount takes its own endpoint reference.
fn systemFsMount(state: *architecture.CpuState) void {
const prefix_ptr = architecture.systemCallArg(state, 0);
const prefix_len = architecture.systemCallArg(state, 1);
const backend_handle = architecture.systemCallArg(state, 2);
const rewrite_ptr = architecture.systemCallArg(state, 3);
const rewrite_len = architecture.systemCallArg(state, 4);
if (prefix_len == 0 or prefix_len > 64 or prefix_ptr >= user_half_end or prefix_ptr + prefix_len > user_half_end) return fail(state);
if (rewrite_len > 32) return fail(state);
if (rewrite_len != 0 and (rewrite_ptr >= user_half_end or rewrite_ptr + rewrite_len > user_half_end)) return fail(state);
const t = scheduler.current();
const prefix = @as([*]const u8, @ptrFromInt(prefix_ptr))[0..prefix_len];
const rewrite = if (rewrite_len == 0) "" else @as([*]const u8, @ptrFromInt(rewrite_ptr))[0..rewrite_len];
const flags = sync.enter();
defer sync.leave(flags);
const endpoint = ipc.resolveHandle(t, backend_handle) orelse return failErr(state, ipc.EBADF);
endpoint.refcount += 1; // the mount table's reference
if (!vfs.mountBackend(prefix, endpoint, rewrite)) {
ipc.dropRef(endpoint);
return fail(state);
}
architecture.setSystemCallResult(state, 0);
}
/// fs_unmount(prefix_ptr, prefix_len): remove a backend mount.
fn systemFsUnmount(state: *architecture.CpuState) void {
const prefix_ptr = architecture.systemCallArg(state, 0);
const prefix_len = architecture.systemCallArg(state, 1);
if (prefix_len == 0 or prefix_len > 64 or prefix_ptr >= user_half_end or prefix_ptr + prefix_len > user_half_end) return fail(state);
const prefix = @as([*]const u8, @ptrFromInt(prefix_ptr))[0..prefix_len];
const flags = sync.enter();
defer sync.leave(flags);
if (!vfs.unmount(prefix)) return fail(state);
architecture.setSystemCallResult(state, 0);
}
/// mmap(len, prot) -> base: grant `len` bytes (rounded up to whole pages) of /// mmap(len, prot) -> base: grant `len` bytes (rounded up to whole pages) of
/// fresh, zeroed, writable+NX memory in the caller's mmap arena, and return the /// fresh, zeroed, writable+NX memory in the caller's mmap arena, and return the
/// base virtual address. `prot` is accepted but not yet honoured (grants are /// base virtual address. `prot` is accepted but not yet honoured (grants are
+1 -1
View File
@@ -113,7 +113,7 @@ pub const Task = struct {
ipc_reply_cap: u64 = 0, ipc_reply_cap: u64 = 0,
ipc_status: i64 = 0, // client: reply length / -errno, written by the replier ipc_status: i64 = 0, // client: reply length / -errno, written by the replier
dma_map_next: u64 = 0, // bump pointer into this task's DMA arena (0 = unseeded) dma_map_next: u64 = 0, // bump pointer into this task's DMA arena (0 = unseeded)
shm_map_next: u64 = 0, // bump pointer into this task's shared-memory arena (0 = unseeded) shared_memory_map_next: u64 = 0, // bump pointer into this task's shared-memory arena (0 = unseeded)
ipc_send_cap: u64 = ~@as(u64, 0), // handle to transfer with this message (abi.no_cap = none) ipc_send_cap: u64 = ~@as(u64, 0), // handle to transfer with this message (abi.no_cap = none)
ipc_received_cap: u64 = ~@as(u64, 0), // client: handle the reply's transferred cap landed at (abi.no_cap = none) ipc_received_cap: u64 = ~@as(u64, 0), // client: handle the reply's transferred cap landed at (abi.no_cap = none)
next: ?*Task = null, // ready-queue link (also the endpoint sender-FIFO link) next: ?*Task = null, // ready-queue link (also the endpoint sender-FIFO link)
+197 -93
View File
@@ -26,11 +26,14 @@ const irq = @import("irq.zig");
const sync = @import("sync.zig"); const sync = @import("sync.zig");
const process = @import("process.zig"); const process = @import("process.zig");
const initial_ramdisk = @import("initial-ramdisk"); const initial_ramdisk = @import("initial-ramdisk");
const kernel_log = @import("log.zig");
const kernel_vfs = @import("vfs.zig");
/// Formatted write straight to serial, independent of the framebuffer console. /// Formatted test-marker write. Goes through the kernel log (not straight to
/// serial): the log lock is what keeps marker lines from interleaving with
/// concurrent user-process records on other cores.
fn log(comptime fmt: []const u8, args: anytype) void { fn log(comptime fmt: []const u8, args: anytype) void {
var buffer: [128]u8 = undefined; kernel_log.print(fmt, args);
architecture.serialWrite(std.fmt.bufPrint(&buffer, fmt, args) catch return);
} }
var passed: u32 = 0; var passed: u32 = 0;
@@ -103,8 +106,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
displayDemoTest(boot_information); displayDemoTest(boot_information);
} else if (eql(case, "display-cursor")) { } else if (eql(case, "display-cursor")) {
displayCursorTest(boot_information); displayCursorTest(boot_information);
} else if (eql(case, "shm")) { } else if (eql(case, "shared-memory")) {
shmTest(boot_information); sharedMemoryTest(boot_information);
} else if (eql(case, "virtio-gpu")) { } else if (eql(case, "virtio-gpu")) {
virtioGpuTest(boot_information); virtioGpuTest(boot_information);
} else if (eql(case, "display-native")) { } else if (eql(case, "display-native")) {
@@ -207,6 +210,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
initialRamdiskTest(boot_information); initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) { } else if (eql(case, "vfs")) {
vfsTest(boot_information); vfsTest(boot_information);
} else if (eql(case, "kvfs")) {
kernelVfsTest(boot_information);
} else if (eql(case, "input")) { } else if (eql(case, "input")) {
inputTest(boot_information); inputTest(boot_information);
} else if (eql(case, "iopass")) { } else if (eql(case, "iopass")) {
@@ -260,6 +265,31 @@ fn bufferHas(needle: []const u8) bool {
return std.mem.indexOf(u8, process.write_buffer[0..process.write_len], needle) != null; return std.mem.indexOf(u8, process.write_buffer[0..process.write_len], needle) != null;
} }
/// Substring search across the whole retained log ring (record payloads are
/// contiguous in the stream, so a one-line needle always matches if present).
/// Unlike `bufferHas` (the single LAST write), this survives busy-tree chatter.
fn ringHas(needle: []const u8) bool {
var chunk: [1024]u8 = undefined;
var overlap: [128]u8 = undefined;
var overlap_len: usize = 0;
var offset = kernel_log.status().tail;
while (true) {
const n = kernel_log.readAt(offset, &chunk) orelse return false;
if (n == 0) return false;
offset += n;
// Search the previous tail glued to this chunk, then the chunk itself.
if (overlap_len != 0) {
var glued: [1152]u8 = undefined;
@memcpy(glued[0..overlap_len], overlap[0..overlap_len]);
const m = @min(n, glued.len - overlap_len);
@memcpy(glued[overlap_len..][0..m], chunk[0..m]);
if (std.mem.indexOf(u8, glued[0 .. overlap_len + m], needle) != null) return true;
} else if (std.mem.indexOf(u8, chunk[0..n], needle) != null) return true;
overlap_len = @min(n, @min(overlap.len, needle.len));
@memcpy(overlap[0..overlap_len], chunk[n - overlap_len ..][0..overlap_len]);
}
}
/// How many `pci_device` functions the devices broker currently holds. A durable /// How many `pci_device` functions the devices broker currently holds. A durable
/// snapshot, unlike a `bufferHas` poll of the single-latest write_buffer line, so a /// snapshot, unlike a `bufferHas` poll of the single-latest write_buffer line, so a
/// test can wait on it without racing transient log output. `scratch` is /// test can wait on it without racing transient log output. `scratch` is
@@ -1327,12 +1357,12 @@ fn procWorker() void {
/// strongest cheap proof of address-space isolation. /// strongest cheap proof of address-space isolation.
fn processTest(boot_information: *const BootInformation) void { fn processTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: process\n", .{}); log("DANOS-TEST-BEGIN: process\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0); const image = bundledInit(boot_information) orelse {
if (boot_information.init_len == 0) { check("initial_ramdisk carries /system/services/init", false);
result(); result();
return; return;
} };
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; check("initial_ramdisk carries /system/services/init", true);
process.write_count = 0; process.write_count = 0;
process.write_from_user = false; process.write_from_user = false;
@@ -1414,12 +1444,12 @@ fn spawnFaultingProcess() ?u32 {
/// time the harness out. /// time the harness out.
fn faultRecoveryTest(boot_information: *const BootInformation) void { fn faultRecoveryTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: fault-recovery\n", .{}); log("DANOS-TEST-BEGIN: fault-recovery\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0); const image = bundledInit(boot_information) orelse {
if (boot_information.init_len == 0) { check("initial_ramdisk carries /system/services/init", false);
result(); result();
return; return;
} };
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; check("initial_ramdisk carries /system/services/init", true);
process.write_count = 0; process.write_count = 0;
process.fault_kill_count = 0; process.fault_kill_count = 0;
@@ -1550,7 +1580,7 @@ fn threadJoinTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "thread-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "thread-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "join" })) true else |_| false; started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "join" })) true else |_| false;
break; break;
} }
@@ -1594,7 +1624,7 @@ fn threadFutexTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "thread-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "thread-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "futex" })) true else |_| false; started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "futex" })) true else |_| false;
break; break;
} }
@@ -1642,7 +1672,7 @@ fn threadMutexTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "thread-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "thread-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "mutex" })) true else |_| false; started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "mutex" })) true else |_| false;
break; break;
} }
@@ -1683,7 +1713,7 @@ fn threadIdTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "thread-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "thread-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "id" })) true else |_| false; started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "id" })) true else |_| false;
break; break;
} }
@@ -1726,7 +1756,7 @@ fn threadAllocTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "thread-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "thread-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "alloc" })) true else |_| false; started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "alloc" })) true else |_| false;
break; break;
} }
@@ -1769,7 +1799,7 @@ fn threadTlsTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "thread-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "thread-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "tls" })) true else |_| false; started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "tls" })) true else |_| false;
break; break;
} }
@@ -1811,7 +1841,7 @@ fn threadRwlockTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "thread-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "thread-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "rwlock" })) true else |_| false; started = if (process.spawnProcess(item.blob, 4, &.{ "thread-test", "rwlock" })) true else |_| false;
break; break;
} }
@@ -1876,12 +1906,12 @@ fn taskReapTest(boot_information: *const BootInformation) void {
/// (write + sleep), and stays alive rather than exiting. /// (write + sleep), and stays alive rather than exiting.
fn initTest(boot_information: *const BootInformation) void { fn initTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: init\n", .{}); log("DANOS-TEST-BEGIN: init\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0); const image = bundledInit(boot_information) orelse {
if (boot_information.init_len == 0) { check("initial_ramdisk carries /system/services/init", false);
result(); result();
return; return;
} };
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; check("initial_ramdisk carries /system/services/init", true);
process.write_count = 0; process.write_count = 0;
const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |err| blk: { const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |err| blk: {
log("DANOS-INIT-ERR: {s}\n", .{@errorName(err)}); log("DANOS-INIT-ERR: {s}\n", .{@errorName(err)});
@@ -1912,12 +1942,12 @@ fn initTest(boot_information: *const BootInformation) void {
/// learns how big a buffer to bring). /// learns how big a buffer to bring).
fn processListTest(boot_information: *const BootInformation) void { fn processListTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: process-list\n", .{}); log("DANOS-TEST-BEGIN: process-list\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0); const image = bundledInit(boot_information) orelse {
if (boot_information.init_len == 0) { check("initial_ramdisk carries /system/services/init", false);
result(); result();
return; return;
} };
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; check("initial_ramdisk carries /system/services/init", true);
var spawned: u32 = 0; var spawned: u32 = 0;
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {} if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
@@ -1963,13 +1993,12 @@ fn processListTest(boot_information: *const BootInformation) void {
/// harness out rather than passing vacuously. /// harness out rather than passing vacuously.
fn processKillTest(boot_information: *const BootInformation) void { fn processKillTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: process-kill\n", .{}); log("DANOS-TEST-BEGIN: process-kill\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0); const image = bundledInit(boot_information) orelse {
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) { check("initial_ramdisk carries /system/services/init", false);
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
result(); result();
return; return;
} };
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; check("initial_ramdisk carries /system/services/init", true);
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len]; const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(ramdisk) orelse { const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
check("initial_ramdisk image is valid", false); check("initial_ramdisk image is valid", false);
@@ -2024,7 +2053,7 @@ fn processKillTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "process-test")) continue;
spinner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "spinner" }, me, endpoint) catch 0; spinner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "spinner" }, me, endpoint) catch 0;
break; break;
} }
@@ -2041,7 +2070,7 @@ fn processKillTest(boot_information: *const BootInformation) void {
i = 0; i = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "args-echo")) continue; if (!eql(initial_ramdisk.basename(item.name), "args-echo")) continue;
clean = process.spawnProcessSupervised(item.blob, 4, &.{ "args-echo", "clean-exit" }, me, endpoint) catch 0; clean = process.spawnProcessSupervised(item.blob, 4, &.{ "args-echo", "clean-exit" }, me, endpoint) catch 0;
break; break;
} }
@@ -2089,12 +2118,12 @@ fn claimReleaseTest(boot_information: *const BootInformation) void {
check("cleanup released owner 222", devices_broker.ownerOf(1) == null); check("cleanup released owner 222", devices_broker.ownerOf(1) == null);
// The death-path wiring: a real process dies holding a claim. // The death-path wiring: a real process dies holding a claim.
check("bootloader handed over /system/services/init", boot_information.init_len != 0); const image = bundledInit(boot_information) orelse {
if (boot_information.init_len == 0) { check("initial_ramdisk carries /system/services/init", false);
result(); result();
return; return;
} };
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; check("initial_ramdisk carries /system/services/init", true);
const me = scheduler.currentId(); const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse { const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false); check("exit endpoint allocated", false);
@@ -2116,11 +2145,11 @@ fn claimReleaseTest(boot_information: *const BootInformation) void {
result(); result();
} }
/// M17.3: the published exit events, proven by their first subscriber. The VFS /// M17.3: the published exit events, proven by a stateful server. The fat
/// subscribes at startup; a client opens a file and parks holding the handle; /// server subscribes at startup; a client opens a file on the volume and parks
/// the kill posts the exit event to the VFS's endpoint; the VFS releases the /// holding the handle; the kill posts the exit event to fat's endpoint; fat
/// dead client's handle and says so — the service-side mirror of iron rule 1 /// releases the dead client's handle and says so — the service-side mirror of
/// (a service must never depend on clients cleaning up after themselves). /// iron rule 1 (a service must never depend on clients cleaning up).
fn vfsClientDeathTest(boot_information: *const BootInformation) void { fn vfsClientDeathTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: vfs-client-death\n", .{}); log("DANOS-TEST-BEGIN: vfs-client-death\n", .{});
if (boot_information.initial_ramdisk_len == 0) { if (boot_information.initial_ramdisk_len == 0) {
@@ -2136,7 +2165,11 @@ fn vfsClientDeathTest(boot_information: *const BootInformation) void {
}; };
process.write_count = 0; process.write_count = 0;
check("vfs spawned", spawnNamed(rd, "vfs")); // The full tree: the storage chain must come up for /mnt/usb to exist —
// the fat server (not a router) now owns client file state and its sweep.
process.setInitialRamdisk(image);
const init_ok = if (process.spawnBundled("/system/services/init")) true else |_| false;
check("init spawned (boots the storage chain)", init_ok);
const me = scheduler.currentId(); const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse { const endpoint = ipcsync.createIpcEndpoint() orelse {
@@ -2148,22 +2181,24 @@ fn vfsClientDeathTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "vfs-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "vfs-test")) continue;
client = process.spawnProcessSupervised(item.blob, 4, &.{ "vfs-test", "park" }, me, endpoint) catch 0; client = process.spawnProcessSupervised(item.blob, 4, &.{ "vfs-test", "park" }, me, endpoint) catch 0;
break; break;
} }
check("parked client spawned (supervised)", client != 0); check("parked client spawned (supervised)", client != 0);
// Its heartbeat is the fence: once it beats, the handle is open. // Its heartbeat is the fence: once it beats, the handle is open. The park
// waits out the whole USB->block->fat chain, so give it room; with the full
// tree chattering, the ring (not the last-write buffer) is the evidence.
const parked = "vfstest: parked"; const parked = "vfstest: parked";
scheduler.setPriority(1); scheduler.setPriority(1);
var deadline = architecture.millis() + 10000; var deadline = architecture.millis() + 30000;
while (architecture.millis() < deadline) { while (architecture.millis() < deadline) {
if (bufferHas(parked)) break; if (ringHas(parked)) break;
scheduler.yield(); scheduler.yield();
} }
scheduler.setPriority(4); scheduler.setPriority(4);
check("client parked holding an open handle", bufferHas(parked)); check("client parked holding an open handle", ringHas(parked));
check("the kill is accepted", process.killProcess(me, client) == 0); check("the kill is accepted", process.killProcess(me, client) == 0);
var badge: u64 = 0; var badge: u64 = 0;
@@ -2171,16 +2206,17 @@ fn vfsClientDeathTest(boot_information: *const BootInformation) void {
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap); _ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | client); check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | client);
// The VFS heard the same published event; its release line is the proof. // The fat server heard the same published event; its release line in the
const released = "vfs: released 1 handle(s) for dead client"; // ring is the proof.
const released = "released 1 handle(s) for dead client";
scheduler.setPriority(1); scheduler.setPriority(1);
deadline = architecture.millis() + 10000; deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) { while (architecture.millis() < deadline) {
if (bufferHas(released)) break; if (ringHas(released)) break;
scheduler.yield(); scheduler.yield();
} }
scheduler.setPriority(4); scheduler.setPriority(4);
check("the VFS released the dead client's handle", bufferHas(released)); check("the fat server released the dead client's handle", ringHas(released));
result(); result();
} }
@@ -2209,7 +2245,7 @@ fn signalsTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "process-test")) continue;
runner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "signal-run" }, scheduler.currentId(), null) catch 0; runner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "signal-run" }, scheduler.currentId(), null) catch 0;
break; break;
} }
@@ -2257,7 +2293,7 @@ fn driverRestartTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue; if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-restart" }, scheduler.currentId(), null) catch 0; manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-restart" }, scheduler.currentId(), null) catch 0;
break; break;
} }
@@ -2295,7 +2331,7 @@ fn usbReportTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue; if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0; manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0;
break; break;
} }
@@ -2327,7 +2363,7 @@ fn deviceListTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue; if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0; manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0;
break; break;
} }
@@ -2368,7 +2404,7 @@ fn pciScanTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue; if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-pci-restart" }, scheduler.currentId(), null) catch 0; manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-pci-restart" }, scheduler.currentId(), null) catch 0;
break; break;
} }
@@ -2452,8 +2488,8 @@ fn usbStorageTest(boot_information: *const BootInformation) void {
/// the fat mount and the client's success. /// the fat mount and the client's success.
fn fatMountTest(boot_information: *const BootInformation) void { fn fatMountTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: fat-mount\n", .{}); log("DANOS-TEST-BEGIN: fat-mount\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) { if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over init and the initial_ramdisk", false); check("bootloader handed over the initial_ramdisk", false);
result(); result();
return; return;
} }
@@ -2464,8 +2500,7 @@ fn fatMountTest(boot_information: *const BootInformation) void {
return; return;
}; };
process.setInitialRamdisk(ramdisk); process.setInitialRamdisk(ramdisk);
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; const init_ok = if (process.spawnBundled("/system/services/init")) true else |_| false;
const init_ok = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned (boots the tree, incl. the fat server)", init_ok); check("init spawned (boots the tree, incl. the fat server)", init_ok);
check("fat-test client spawned", spawnNamed(rd, "fat-test")); check("fat-test client spawned", spawnNamed(rd, "fat-test"));
result(); result();
@@ -2473,30 +2508,28 @@ fn fatMountTest(boot_information: *const BootInformation) void {
fn bootServiceTreeTest(boot_information: *const BootInformation, comptime label: []const u8) void { fn bootServiceTreeTest(boot_information: *const BootInformation, comptime label: []const u8) void {
log("DANOS-TEST-BEGIN: " ++ label ++ "\n", .{}); log("DANOS-TEST-BEGIN: " ++ label ++ "\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) { if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over init and the initial_ramdisk", false); check("bootloader handed over the initial_ramdisk", false);
result(); result();
return; return;
} }
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len]; const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
process.setInitialRamdisk(ramdisk); process.setInitialRamdisk(ramdisk);
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; const spawned = if (process.spawnBundled("/system/services/init")) true else |_| false;
const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned (boots vfs, input, device-manager, and the USB chain)", spawned); check("init spawned (boots vfs, input, device-manager, and the USB chain)", spawned);
result(); result();
} }
fn orderlyShutdownTest(boot_information: *const BootInformation) void { fn orderlyShutdownTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: orderly-shutdown\n", .{}); log("DANOS-TEST-BEGIN: orderly-shutdown\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) { if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over init and the initial_ramdisk", false); check("bootloader handed over the initial_ramdisk", false);
result(); result();
return; return;
} }
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len]; const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
process.setInitialRamdisk(ramdisk); process.setInitialRamdisk(ramdisk);
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; const spawned = if (process.spawnBundled("/system/services/init")) true else |_| false;
const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned as PID root of user space", spawned); check("init spawned as PID root of user space", spawned);
result(); result();
} }
@@ -2524,7 +2557,7 @@ fn acpiReportTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue; if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
_ = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0; _ = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0;
spawned = true; spawned = true;
break; break;
@@ -2562,7 +2595,7 @@ fn acpiParseTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "discovery")) continue; if (!eql(initial_ramdisk.basename(item.name), "discovery")) continue;
_ = process.spawnProcessSupervised(item.blob, 4, &.{ "discovery", "1" }, scheduler.currentId(), null) catch 0; _ = process.spawnProcessSupervised(item.blob, 4, &.{ "discovery", "1" }, scheduler.currentId(), null) catch 0;
spawned = true; spawned = true;
break; break;
@@ -2597,7 +2630,7 @@ fn supervisionTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue; if (!eql(initial_ramdisk.basename(item.name), "process-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "process-test", "run" })) true else |_| false; started = if (process.spawnProcess(item.blob, 4, &.{ "process-test", "run" })) true else |_| false;
break; break;
} }
@@ -2682,9 +2715,10 @@ fn vfsTest(boot_information: *const BootInformation) void {
process.write_count = 0; process.write_count = 0;
process.write_from_user = false; process.write_from_user = false;
// Spawn just the server and its client (other initial_ramdisk binaries would write to // Seed the kernel VFS (/system) — the router the client exercises.
// the shared evidence buffer and confuse the marker check). process.setInitialRamdisk(image);
_ = spawnNamed(rd, "vfs"); // Spawn just the client: the kernel itself is the VFS root it exercises
// (resolve + fs_node over /system through the plain runtime.fs API).
_ = spawnNamed(rd, "vfs-test"); _ = spawnNamed(rd, "vfs-test");
// Wait for the client's success heartbeat (it round-trips, then beats ~1/s). // Wait for the client's success heartbeat (it round-trips, then beats ~1/s).
@@ -2860,14 +2894,14 @@ fn displayDemoTest(boot_information: *const BootInformation) void {
while (true) scheduler.yield(); while (true) scheduler.yield();
} }
/// V2 — cross-process shared memory (docs/display-v2.md). Spawn shm-server and shm-client: /// V2 — cross-process shared memory (docs/display-v2.md). Spawn shared-memory-server and shared-memory-client:
/// the client shm_creates a region, writes a pattern, and passes the region's capability to /// the client shared_memory_creates a region, writes a pattern, and passes the region's capability to
/// the server as an ipc_call send_cap; the server shm_maps it and confirms the pattern is /// the server as an ipc_call send_cap; the server shared_memory_maps it and confirms the pattern is
/// visible — proving the two processes share the same physical pages, and that the extended /// visible — proving the two processes share the same physical pages, and that the extended
/// capability-passing (endpoints → memory objects) works. Its `shm: shared 4096 bytes ok` /// capability-passing (endpoints → memory objects) works. Its `shared-memory: shared 4096 bytes ok`
/// heartbeat is the marker. /// heartbeat is the marker.
fn shmTest(boot_information: *const BootInformation) void { fn sharedMemoryTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: shm\n", .{}); log("DANOS-TEST-BEGIN: shared-memory\n", .{});
if (boot_information.initial_ramdisk_len == 0) { if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false); check("bootloader handed over an initial_ramdisk", false);
result(); result();
@@ -2880,12 +2914,12 @@ fn shmTest(boot_information: *const BootInformation) void {
return; return;
}; };
if (!spawnNamed(rd, "shm-server")) { if (!spawnNamed(rd, "shared-memory-server")) {
log("shm: could not spawn shm-server\n", .{}); log("shared-memory: could not spawn shared-memory-server\n", .{});
result(); result();
return; return;
} }
_ = spawnNamed(rd, "shm-client"); _ = spawnNamed(rd, "shared-memory-client");
scheduler.setPriority(1); // below the two, so they run scheduler.setPriority(1); // below the two, so they run
while (true) scheduler.yield(); while (true) scheduler.yield();
} }
@@ -2921,7 +2955,7 @@ fn virtioGpuTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue; if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0; manager = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0;
break; break;
} }
@@ -2962,7 +2996,7 @@ fn displayNativeTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue; if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0; manager = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0;
break; break;
} }
@@ -3006,7 +3040,7 @@ fn displayReattachTest(boot_information: *const BootInformation) void {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue; if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-scanout-restart" }, scheduler.currentId(), null) catch 0; manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-scanout-restart" }, scheduler.currentId(), null) catch 0;
break; break;
} }
@@ -3055,7 +3089,7 @@ fn argsTest(boot_information: *const BootInformation) void {
while (process.write_count < 1 and architecture.millis() < deadline) scheduler.yield(); while (process.write_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4); scheduler.setPriority(4);
const expected = "args: args-echo alpha beta-42\n"; const expected = "args: /system/tests/args-echo alpha beta-42\n";
const echoed = process.write_len == expected.len and eql(process.write_buffer[0..process.write_len], expected); const echoed = process.write_len == expected.len and eql(process.write_buffer[0..process.write_len], expected);
if (!echoed and process.write_len > 0) log("DANOS-ARGS: got \"{s}\"\n", .{process.write_buffer[0..process.write_len]}); if (!echoed and process.write_len > 0) log("DANOS-ARGS: got \"{s}\"\n", .{process.write_buffer[0..process.write_len]});
check("argv arrived intact (argv[0] = name, argv[1..] = spawn arguments)", echoed); check("argv arrived intact (argv[0] = name, argv[1..] = spawn arguments)", echoed);
@@ -3065,11 +3099,79 @@ fn argsTest(boot_information: *const BootInformation) void {
/// Spawn the initial_ramdisk binary named `name` as a ring-3 process. Returns false if it /// Spawn the initial_ramdisk binary named `name` as a ring-3 process. Returns false if it
/// isn't in the image or fails to load. /// isn't in the image or fails to load.
/// The init ELF image out of the initial_ramdisk — init rides the table like
/// every other binary since the loader packs the whole /system tree.
fn bundledInit(boot_information: *const BootInformation) ?[]const u8 {
if (boot_information.initial_ramdisk_len == 0) return null;
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse return null;
const item = rd.find("/system/services/init") orelse return null;
return item.blob;
}
/// The kernel VFS root (M-F): resolve initrd paths to node tokens, read an ELF
/// header through nodeRead, enumerate /system's derived directory table, and
/// verify the read-only + unknown-path refusals. Pure kernel-side — the
/// syscall surface gets its end-to-end coverage when runtime.fs cuts over.
fn kernelVfsTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: kvfs\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over the initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
process.setInitialRamdisk(image); // also seeds the kernel VFS /system mount
// A file resolves to a kernel node token; its status and bytes are served.
const resolved = kernel_vfs.resolvePath("/system/services/init", false);
const is_file = resolved == .kernel_node;
check("/system/services/init resolves to a kernel node", is_file);
if (is_file) {
const status = kernel_vfs.nodeStatus(resolved.kernel_node);
check("its status is a non-empty regular file", status != null and status.?.kind == abi.file_kind_regular and status.?.size > 0);
var header: [4]u8 = undefined;
const n = kernel_vfs.nodeRead(resolved.kernel_node, 0, &header) orelse 0;
check("its first bytes are an ELF magic", n == 4 and header[0] == 0x7f and header[1] == 'E' and header[2] == 'L' and header[3] == 'F');
}
// Directories resolve and enumerate: /system lists services/drivers/tests.
const root_directory = kernel_vfs.resolvePath("/system", false);
check("/system resolves to a directory node", root_directory == .kernel_node);
var saw_services = false;
var saw_drivers = false;
var saw_files_in_services = false;
if (root_directory == .kernel_node) {
var cursor: u64 = 0;
var name: [64]u8 = undefined;
while (kernel_vfs.nodeReaddir(root_directory.kernel_node, cursor, &name)) |entry| : (cursor += 1) {
if (eql(name[0..entry.name_len], "services")) saw_services = true;
if (eql(name[0..entry.name_len], "drivers")) saw_drivers = true;
}
}
check("readdir /system yields services and drivers", saw_services and saw_drivers);
const services = kernel_vfs.resolvePath("/system/services", false);
if (services == .kernel_node) {
var cursor: u64 = 0;
var name: [64]u8 = undefined;
while (kernel_vfs.nodeReaddir(services.kernel_node, cursor, &name)) |entry| : (cursor += 1) {
if (eql(name[0..entry.name_len], "init")) saw_files_in_services = true;
}
}
check("readdir /system/services yields init", saw_files_in_services);
// Refusals: unknown paths, and create-intent on the immutable initrd.
check("an unknown path does not resolve", kernel_vfs.resolvePath("/system/services/no-such", false) == .not_found);
check("an unmounted absolute path does not resolve", kernel_vfs.resolvePath("/elsewhere", false) == .not_found);
check("create on /system is refused (read-only)", kernel_vfs.resolvePath("/system/services/new-file", true) == .not_found);
result();
}
fn spawnNamed(rd: initial_ramdisk.Reader, name: []const u8) bool { fn spawnNamed(rd: initial_ramdisk.Reader, name: []const u8) bool {
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (eql(item.name, name)) { if (eql(initial_ramdisk.basename(item.name), name)) {
return if (process.spawnProcess(item.blob, 4, &.{item.name})) true else |_| false; return if (process.spawnProcess(item.blob, 4, &.{item.name})) true else |_| false;
} }
} }
@@ -3082,7 +3184,7 @@ fn spawnNamedWithArg(rd: initial_ramdisk.Reader, name: []const u8, arg: []const
var i: u32 = 0; var i: u32 = 0;
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (eql(item.name, name)) { if (eql(initial_ramdisk.basename(item.name), name)) {
return if (process.spawnProcess(item.blob, 4, &.{ item.name, arg })) true else |_| false; return if (process.spawnProcess(item.blob, 4, &.{ item.name, arg })) true else |_| false;
} }
} }
@@ -3236,7 +3338,8 @@ fn processRunning(name: []const u8) bool {
var table: [64]abi.ProcessDescriptor = undefined; var table: [64]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table); const total = scheduler.enumerate(&table);
for (table[0..@min(total, table.len)]) |d| { for (table[0..@min(total, table.len)]) |d| {
if (std.mem.eql(u8, d.name[0..d.name_length], name)) return true; // Task names are full binary paths; callers pass either form.
if (std.mem.eql(u8, initial_ramdisk.basename(d.name[0..d.name_length]), name)) return true;
} }
return false; return false;
} }
@@ -3404,6 +3507,7 @@ fn displayTest(boot_information: *const BootInformation) void {
check("the node is class display", d.class == @intFromEnum(device_abi.DeviceClass.display)); check("the node is class display", d.class == @intFromEnum(device_abi.DeviceClass.display));
check("it carries the framebuffer geometry", d.display.width == fb.width and d.display.height == fb.height and d.display.pitch == fb.pitch); check("it carries the framebuffer geometry", d.display.width == fb.width and d.display.height == fb.height and d.display.pitch == fb.pitch);
check("it carries the panel refresh rate", d.display.refresh_hz == fb.refresh_hz);
check("it has exactly one resource", d.resource_count == 1); check("it has exactly one resource", d.resource_count == 1);
const r = d.resources[0]; const r = d.resources[0];
check("that resource is a memory window", r.kind == @intFromEnum(device_abi.ResourceKind.memory)); check("that resource is a memory window", r.kind == @intFromEnum(device_abi.ResourceKind.memory));
+368
View File
@@ -0,0 +1,368 @@
//! The kernel-resident VFS root: the mount table and the kernel-backed nodes.
//!
//! The kernel's job here is NAMING, never data plumbing to userspace backends —
//! the mechanism is **resolve + redirect**:
//!
//! - `fs_resolve(path)` walks the mount table. A path under a KERNEL-backed
//! mount (the initrd at /system, the scratch ram nodes) resolves to a
//! stateless node TOKEN served directly by `fs_node` (read/status/readdir
//! with copy-out). A path under a USERSPACE mount (the fat server at
//! /mnt/usb and /var) resolves to the backend's ENDPOINT: the kernel
//! installs a (deduplicated) handle in the caller's table, rewrites the
//! path mount-relative, and the caller speaks the unchanged vfs-protocol
//! to the backend over the ordinary ipc_call rendezvous. The kernel never
//! blocks on a userspace server.
//!
//! - Kernel node tokens are PERMANENT for a boot: the initrd is immutable and
//! ram nodes are never reclaimed — no open-handle state, no close, no sweep
//! on client death. Backend file state lives in the backend, which sweeps
//! dead clients itself via the published exit events.
//!
//! Mounting is `fs_mount(prefix, backend_handle, rewrite)`: possession of the
//! backend endpoint handle is the capability, exactly the trust of the old
//! userspace router's op-6 cap-pass. An optional REWRITE prefix maps the mount
//! into the backend's namespace ("/var" -> fat's "/var" subtree while the same
//! backend also serves "/mnt/usb" from its root), so FHS paths stay decoupled
//! from which volume happens to carry them.
const std = @import("std");
const abi = @import("abi");
const initial_ramdisk = @import("initial-ramdisk");
const ipc = @import("ipc-synchronous.zig");
// --- node tokens -------------------------------------------------------------
/// Kind lives in the top byte of a token; the index below. Tokens are permanent
/// for a boot, so userspace may cache them freely.
pub const token_kind_shift = 56;
pub const token_kind_initrd_file: u64 = 1;
pub const token_kind_initrd_directory: u64 = 2;
pub const token_kind_ram: u64 = 3;
fn token(kind: u64, index: u64) u64 {
return (kind << token_kind_shift) | index;
}
fn tokenKind(t: u64) u64 {
return t >> token_kind_shift;
}
fn tokenIndex(t: u64) u64 {
return t & ((@as(u64, 1) << token_kind_shift) - 1);
}
// --- the mount table ---------------------------------------------------------
pub const maximum_mounts = 8;
const maximum_prefix = 64;
const maximum_rewrite = 32;
const MountKind = enum(u8) { kernel_initrd, backend };
const Mount = struct {
used: bool = false,
prefix: [maximum_prefix]u8 = undefined,
prefix_len: usize = 0,
kind: MountKind = .backend,
backend: ?*ipc.Endpoint = null, // referenced while mounted
rewrite: [maximum_rewrite]u8 = undefined,
rewrite_len: usize = 0,
fn prefixSlice(self: *const Mount) []const u8 {
return self.prefix[0..self.prefix_len];
}
fn rewriteSlice(self: *const Mount) []const u8 {
return self.rewrite[0..self.rewrite_len];
}
};
var mounts: [maximum_mounts]Mount = @splat(.{});
/// The initrd image (set once at boot) and its derived directory table.
var ramdisk_image: ?[]const u8 = null;
const maximum_directories = 8;
const Directory = struct {
path: [maximum_prefix]u8 = undefined,
path_len: usize = 0,
parent: usize = 0, // index into `directories`; 0 is /system itself
fn slice(self: *const Directory) []const u8 {
return self.path[0..self.path_len];
}
};
var directories: [maximum_directories]Directory = @splat(.{});
var directory_count: usize = 0;
// --- pure path helpers (ported from the userspace router, with its tests) ----
/// If `path` lies under `mount_prefix` — equal to it, or the prefix followed by
/// a path separator — return the path relative to the mount ("/" for an exact
/// match, otherwise the tail beginning with '/'). Null when not under the
/// mount, so "/mnt/usb" never captures "/mnt/usbextra".
pub fn underMount(path: []const u8, mount_prefix: []const u8) ?[]const u8 {
if (path.len < mount_prefix.len) return null;
if (!std.mem.eql(u8, path[0..mount_prefix.len], mount_prefix)) return null;
if (path.len == mount_prefix.len) return "/";
if (path[mount_prefix.len] != '/') return null;
return path[mount_prefix.len..];
}
pub fn isAbsolute(path: []const u8) bool {
return path.len > 0 and path[0] == '/';
}
/// The parent directory portion of an initrd path ("/system/services/fat" ->
/// "/system/services").
fn parentOf(path: []const u8) []const u8 {
const slash = std.mem.lastIndexOfScalar(u8, path, '/') orelse return path[0..0];
if (slash == 0) return path[0..1];
return path[0..slash];
}
// --- boot wiring -------------------------------------------------------------
/// Publish the initrd as the kernel-backed /system mount and derive its bounded
/// directory table (the unique parents of the entry paths). Called once at boot.
pub fn setInitialRamdisk(image: []const u8) void {
ramdisk_image = image;
installMount("/system", .kernel_initrd, null, "");
// Directory 0 is /system itself.
directories[0] = .{ .parent = 0 };
@memcpy(directories[0].path[0..7], "/system");
directories[0].path_len = 7;
directory_count = 1;
const rd = initial_ramdisk.Reader.init(image) orelse return;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
// Register every ancestor directory strictly below /system.
var parent = parentOf(item.name);
while (parent.len > 7) : (parent = parentOf(parent)) {
if (directoryIndex(parent) == null and directory_count < maximum_directories) {
var d = &directories[directory_count];
@memcpy(d.path[0..parent.len], parent);
d.path_len = parent.len;
d.parent = 0; // fixed up below once all exist
directory_count += 1;
}
}
}
// Parent links (a second pass so out-of-order registration doesn't matter).
for (directories[1..directory_count]) |*d| {
d.parent = directoryIndex(parentOf(d.slice())) orelse 0;
}
}
fn directoryIndex(path: []const u8) ?usize {
for (directories[0..directory_count], 0..) |*d, i| {
if (std.mem.eql(u8, d.slice(), path)) return i;
}
return null;
}
fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, rewrite: []const u8) void {
// Remount replaces: a restarted backend re-mounts its prefix.
var slot: ?*Mount = null;
for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefixSlice(), prefix)) {
if (m.backend) |old| ipc.dropRef(old);
slot = m;
break;
}
if (slot == null and !m.used) slot = m;
}
const m = slot orelse return;
m.* = .{ .used = true, .kind = kind, .backend = backend };
@memcpy(m.prefix[0..prefix.len], prefix);
m.prefix_len = prefix.len;
@memcpy(m.rewrite[0..rewrite.len], rewrite);
m.rewrite_len = rewrite.len;
}
// --- resolve -----------------------------------------------------------------
pub const Resolved = union(enum) {
/// Kernel-served: a permanent node token.
kernel_node: u64,
/// Backend-served: the endpoint plus the rewritten mount-relative path.
backend: struct { endpoint: *ipc.Endpoint, path: [maximum_rewrite + maximum_prefix + 160]u8, path_len: usize },
not_found: void,
};
/// Longest-prefix match over the mount table, then per-kind resolution.
/// `create`-intent on the immutable /system fails here (EROFS-style).
pub fn resolvePath(path: []const u8, wants_create: bool) Resolved {
if (!isAbsolute(path)) {
return .{ .not_found = {} }; // bare names have no kernel namespace (ramfs retired)
}
var best: ?*Mount = null;
var best_relative: []const u8 = undefined;
for (&mounts) |*m| {
if (!m.used) continue;
const relative = underMount(path, m.prefixSlice()) orelse continue;
if (best == null or m.prefix_len > best.?.prefix_len) {
best = m;
best_relative = relative;
}
}
const m = best orelse return .{ .not_found = {} };
switch (m.kind) {
.kernel_initrd => {
if (wants_create) return .{ .not_found = {} }; // read-only
return resolveInitrd(path);
},
.backend => {
const endpoint = m.backend orelse return .{ .not_found = {} };
if (endpoint.dead) {
// The backend died: treat the mount as gone (it re-mounts on
// restart) and release our reference lazily.
ipc.dropRef(endpoint);
m.backend = null;
m.used = false;
return .{ .not_found = {} };
}
var out: Resolved = .{ .backend = .{ .endpoint = endpoint, .path = undefined, .path_len = 0 } };
const rewrite = m.rewriteSlice();
const tail = if (std.mem.eql(u8, best_relative, "/") and rewrite.len != 0) "" else best_relative;
const total = rewrite.len + tail.len;
if (total > out.backend.path.len or total == 0) {
if (rewrite.len == 0 and tail.len == 0) return .{ .not_found = {} };
if (total > out.backend.path.len) return .{ .not_found = {} };
}
@memcpy(out.backend.path[0..rewrite.len], rewrite);
@memcpy(out.backend.path[rewrite.len..][0..tail.len], tail);
out.backend.path_len = total;
return out;
},
}
}
fn resolveInitrd(path: []const u8) Resolved {
if (directoryIndex(path)) |index| return .{ .kernel_node = token(token_kind_initrd_directory, index) };
const image = ramdisk_image orelse return .{ .not_found = {} };
const rd = initial_ramdisk.Reader.init(image) orelse return .{ .not_found = {} };
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (std.mem.eql(u8, item.name, path)) return .{ .kernel_node = token(token_kind_initrd_file, i) };
}
return .{ .not_found = {} };
}
// --- fs_node: serving kernel-backed nodes ------------------------------------
/// Read `out.len` bytes of an initrd file at `offset`. Returns bytes copied
/// (0 at EOF) or null for a bad token. Lock-free: the initrd is immutable.
pub fn nodeRead(node_token: u64, offset: u64, out: []u8) ?usize {
if (tokenKind(node_token) != token_kind_initrd_file) return null;
const image = ramdisk_image orelse return null;
const rd = initial_ramdisk.Reader.init(image) orelse return null;
const item = rd.entry(@intCast(tokenIndex(node_token))) orelse return null;
if (offset >= item.blob.len) return 0;
const n = @min(out.len, item.blob.len - @as(usize, @intCast(offset)));
@memcpy(out[0..n], item.blob[@intCast(offset)..][0..n]);
return n;
}
/// A node's metadata in vfs-protocol FileStatus shape (size, kind, mtime).
pub fn nodeStatus(node_token: u64) ?abi.FileAttributes {
switch (tokenKind(node_token)) {
token_kind_initrd_file => {
const image = ramdisk_image orelse return null;
const rd = initial_ramdisk.Reader.init(image) orelse return null;
const item = rd.entry(@intCast(tokenIndex(node_token))) orelse return null;
return .{ .size = item.blob.len, .kind = abi.file_kind_regular };
},
token_kind_initrd_directory => {
if (tokenIndex(node_token) >= directory_count) return null;
return .{ .size = 0, .kind = abi.file_kind_directory };
},
else => return null,
}
}
/// The `cursor`th child of an initrd directory: fills `name_out`, returns the
/// entry header, or null past the end / bad token. Cursor enumerates
/// subdirectories first, then files whose parent is this directory — stable,
/// because the initrd is immutable.
pub fn nodeReaddir(node_token: u64, cursor: u64, name_out: []u8) ?struct { header: abi.DirectoryEntryHeader, name_len: usize } {
if (tokenKind(node_token) != token_kind_initrd_directory) return null;
const directory_index = tokenIndex(node_token);
if (directory_index >= directory_count) return null;
const self_path = directories[@intCast(directory_index)].slice();
var index: u64 = 0;
// Subdirectories whose parent is this directory.
for (directories[0..directory_count], 0..) |*d, i| {
if (i == directory_index) continue;
if (d.parent != directory_index) continue;
if (i == 0) continue;
if (index == cursor) {
const name = d.slice()[self_path.len + 1 ..];
const n = @min(name.len, name_out.len);
@memcpy(name_out[0..n], name[0..n]);
return .{ .header = .{ .kind = abi.file_kind_directory, .name_len = @intCast(n), .size = 0 }, .name_len = n };
}
index += 1;
}
// Files directly inside this directory.
const image = ramdisk_image orelse return null;
const rd = initial_ramdisk.Reader.init(image) orelse return null;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!std.mem.eql(u8, parentOf(item.name), self_path)) continue;
if (index == cursor) {
const name = item.name[self_path.len + 1 ..];
const n = @min(name.len, name_out.len);
@memcpy(name_out[0..n], name[0..n]);
return .{ .header = .{ .kind = abi.file_kind_regular, .name_len = @intCast(n), .size = item.blob.len }, .name_len = n };
}
index += 1;
}
return null;
}
// --- mount/unmount (syscall bodies; caller resolved the handle) --------------
/// Mount `backend` at `prefix` with an optional backend-side `rewrite` prefix.
/// The endpoint reference is taken by the caller (process.zig bumps it); refuses
/// shadowing or replacing /system.
pub fn mountBackend(prefix: []const u8, backend: *ipc.Endpoint, rewrite: []const u8) bool {
if (!isAbsolute(prefix) or prefix.len < 2 or prefix.len > maximum_prefix) return false;
if (rewrite.len > maximum_rewrite) return false;
if (underMount(prefix, "/system") != null) return false; // the initrd is not shadowable
installMount(prefix, .backend, backend, rewrite);
return true;
}
pub fn unmount(prefix: []const u8) bool {
for (&mounts) |*m| {
if (m.used and m.kind == .backend and std.mem.eql(u8, m.prefixSlice(), prefix)) {
if (m.backend) |endpoint| ipc.dropRef(endpoint);
m.* = .{};
return true;
}
}
return false;
}
// --- tests (host) ------------------------------------------------------------
test "underMount matches only at path boundaries" {
try std.testing.expectEqualStrings("/", underMount("/mnt/usb", "/mnt/usb").?);
try std.testing.expectEqualStrings("/system/kernel", underMount("/mnt/usb/system/kernel", "/mnt/usb").?);
try std.testing.expect(underMount("/mnt/usbextra", "/mnt/usb") == null);
try std.testing.expect(underMount("/mnt", "/mnt/usb") == null);
try std.testing.expect(underMount("/other", "/mnt/usb") == null);
try std.testing.expect(underMount("greeting", "/mnt/usb") == null);
}
test "parentOf walks toward the root" {
try std.testing.expectEqualStrings("/system/services", parentOf("/system/services/fat"));
try std.testing.expectEqualStrings("/system", parentOf("/system/services"));
try std.testing.expectEqualStrings("/", parentOf("/system"));
}
+6
View File
@@ -24,3 +24,9 @@ pub fn init() void {
pub fn nowSeconds() u64 { pub fn nowSeconds() u64 {
return boot_unix_seconds + (architecture.nanos() -% boot_nanos) / 1_000_000_000; return boot_unix_seconds + (architecture.nanos() -% boot_nanos) / 1_000_000_000;
} }
/// The wall-clock time of boot itself (the RTC anchor) — what klog_status hands
/// the logger service to name a per-boot log directory. Zero until `init` runs.
pub fn bootSeconds() u64 {
return boot_unix_seconds;
}
+7 -12
View File
@@ -21,11 +21,6 @@ const power = runtime.power_protocol;
/// integer decode names the opcodes instead of bare 0x0A/0x0B/… (docs/coding-standards.md). /// integer decode names the opcodes instead of bare 0x0A/0x0B/… (docs/coding-standards.md).
const opcodes = aml.opcodes; const opcodes = aml.opcodes;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// The claimed acpi-tables node and the resource index of its broad io_port // The claimed acpi-tables node and the resource index of its broad io_port
// window — the Hal routes every port access through this one claim. // window — the Hal routes every port access through this one claim.
var node_id: u64 = 0; var node_id: u64 = 0;
@@ -163,12 +158,12 @@ pub fn main(init: runtime.process.Init) void {
}; };
var namespace = result.namespace; var namespace = result.namespace;
const devices = aml.deviceCount(&namespace); const devices = aml.deviceCount(&namespace);
writeLine("/system/services/acpi: parsed {d} AML blob(s), {d} namespace devices\n", .{ block_count, devices }); std.log.info("parsed {d} AML blob(s), {d} namespace devices", .{ block_count, devices });
if (floor) |minimum| { if (floor) |minimum| {
if (devices >= minimum) { if (devices >= minimum) {
_ = runtime.system.write("acpi-parse: ok\n"); _ = runtime.system.write("acpi-parse: ok\n");
} else { } else {
writeLine("acpi-parse: too few (ring-3 {d} < floor {d})\n", .{ devices, minimum }); std.log.info("acpi-parse: too few (ring-3 {d} < floor {d})", .{ devices, minimum });
} }
// Self-verify mode is standalone (no manager); stop before reporting. // Self-verify mode is standalone (no manager); stop before reporting.
while (true) runtime.system.sleep(1000); while (true) runtime.system.sleep(1000);
@@ -216,9 +211,9 @@ fn onInit(endpoint: runtime.ipc.Handle) bool {
const hid = entry.hid[0..entry.hid_len]; const hid = entry.hid[0..entry.hid_len];
const desc = acpi_ids.description(hid); const desc = acpi_ids.description(hid);
if (desc.len != 0) if (desc.len != 0)
writeLine("/system/services/acpi: reported {s} (device {d}, {d} resources) — {s}\n", .{ hid, entry.device_id, entry.resource_count, desc }) std.log.info("reported {s} (device {d}, {d} resources) — {s}", .{ hid, entry.device_id, entry.resource_count, desc })
else else
writeLine("/system/services/acpi: reported {s} (device {d}, {d} resources)\n", .{ hid, entry.device_id, entry.resource_count }); std.log.info("reported {s} (device {d}, {d} resources)", .{ hid, entry.device_id, entry.resource_count });
if (manager) |h| { if (manager) |h| {
var report = protocol.ChildAdded{ .parent = node_id, .bus_address = entry.device_id, .identity = 0, .device_id = entry.device_id }; var report = protocol.ChildAdded{ .parent = node_id, .bus_address = entry.device_id, .identity = 0, .device_id = entry.device_id };
@memcpy(report.hid[0..entry.hid_len], entry.hid[0..entry.hid_len]); @memcpy(report.hid[0..entry.hid_len], entry.hid[0..entry.hid_len]);
@@ -226,7 +221,7 @@ fn onInit(endpoint: runtime.ipc.Handle) bool {
_ = runtime.ipc.call(h, std.mem.asBytes(&report), &reply) catch {}; _ = runtime.ipc.call(h, std.mem.asBytes(&report), &reply) catch {};
} }
} }
writeLine("/system/services/acpi: reported {d} device(s) to the manager\n", .{registered_count}); std.log.info("reported {d} device(s) to the manager", .{registered_count});
armPowerButton(endpoint); armPowerButton(endpoint);
return true; return true;
@@ -373,7 +368,7 @@ fn publishNotify(node: *aml.Node, code: u64) void {
const which: power.Event = if (std.mem.eql(u8, hid[0..7], "PNP0C0A")) .battery else if (std.mem.eql(u8, hid[0..7], "ACPI0003")) .ac else if (std.mem.eql(u8, hid[0..7], "PNP0C0D")) .lid else .notify; const which: power.Event = if (std.mem.eql(u8, hid[0..7], "PNP0C0A")) .battery else if (std.mem.eql(u8, hid[0..7], "ACPI0003")) .ac else if (std.mem.eql(u8, hid[0..7], "PNP0C0D")) .lid else .notify;
var event = power.EventMessage{ .event = @intFromEnum(which), .code = @truncate(code) }; var event = power.EventMessage{ .event = @intFromEnum(which), .code = @truncate(code) };
event.hid = hid; event.hid = hid;
writeLine("power: notify {s} code {d}\n", .{ hid[0..7], code }); std.log.info("power: notify {s} code {d}", .{ hid[0..7], code });
publishEvent(std.mem.asBytes(&event)); publishEvent(std.mem.asBytes(&event));
} }
@@ -500,7 +495,7 @@ fn registerDevice(node: *aml.Node, hid: [8]u8, interpreter: *aml.Interpreter) vo
applyCrs(&descriptor, node, interpreter); applyCrs(&descriptor, node, interpreter);
const id = device.register(node_id, &descriptor) orelse { const id = device.register(node_id, &descriptor) orelse {
writeLine("/system/services/acpi: register refused for {s}\n", .{hid[0..@intCast(hid_len)]}); std.log.info("register refused for {s}", .{hid[0..@intCast(hid_len)]});
return; return;
}; };
registered[registered_count] = .{ .hid = hid, .hid_len = @intCast(hid_len), .device_id = id, .resource_count = descriptor.resource_count }; registered[registered_count] = .{ .hid = hid, .hid_len = @intCast(hid_len), .device_id = id, .resource_count = descriptor.resource_count };
@@ -24,14 +24,6 @@ const protocol = runtime.device_manager_protocol;
const device = runtime.device; const device = runtime.device;
const system = runtime.system; const system = runtime.system;
/// Format one whole log line and emit it in a single `debug_write`, so output
/// from the drivers this manager starts (which run concurrently) can never land
/// in the middle of it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// The PCI class/subclass/prog-IF triple of an xHCI (USB 3) host controller — /// The PCI class/subclass/prog-IF triple of an xHCI (USB 3) host controller —
/// Serial Bus Controller / USB Controller / XHCI — named from pci-class.zig rather /// Serial Bus Controller / USB Controller / XHCI — named from pci-class.zig rather
/// than written as the bare 0x0C0330 (docs/coding-standards.md, "Named values"). /// than written as the bare 0x0C0330 (docs/coding-standards.md, "Named values").
@@ -56,8 +48,8 @@ const virtio_gpu_pci_class: u64 = pci_class.ClassCode.pack(.{
/// its registered id as argv[1]. /// its registered id as argv[1].
fn pciDriverForIdentity(identity: u64) ?[]const u8 { fn pciDriverForIdentity(identity: u64) ?[]const u8 {
return switch (identity) { return switch (identity) {
xhci_pci_class => "usb-xhci-bus", xhci_pci_class => "/system/drivers/usb-xhci-bus",
virtio_gpu_pci_class => "virtio-gpu", virtio_gpu_pci_class => "/system/drivers/virtio-gpu",
else => null, else => null,
}; };
} }
@@ -67,8 +59,8 @@ fn pciDriverForIdentity(identity: u64) ?[]const u8 {
/// nodes the kernel used to build). ps2-bus is a singleton that finds both its /// nodes the kernel used to build). ps2-bus is a singleton that finds both its
/// devices by hid once spawned, so keyboard and mouse map to the same name. /// devices by hid once spawned, so keyboard and mouse map to the same name.
fn hidDriverFor(hid: []const u8) ?[]const u8 { fn hidDriverFor(hid: []const u8) ?[]const u8 {
if (std.mem.eql(u8, hid, "PNP0303")) return "ps2-bus"; // PS/2 keyboard if (std.mem.eql(u8, hid, "PNP0303")) return "/system/drivers/ps2-bus"; // PS/2 keyboard
if (std.mem.eql(u8, hid, "PNP0F13")) return "ps2-bus"; // PS/2 mouse if (std.mem.eql(u8, hid, "PNP0F13")) return "/system/drivers/ps2-bus"; // PS/2 mouse
return null; return null;
} }
@@ -95,9 +87,9 @@ fn usbDriverForIdentity(identity: u64) ?[]const u8 {
@intFromEnum(usb_ids.mass_storage.Protocol.bulk_only), @intFromEnum(usb_ids.mass_storage.Protocol.bulk_only),
); );
return switch (identity) { return switch (identity) {
keyboard => "usb-hid-keyboard", keyboard => "/system/drivers/usb-hid-keyboard",
mouse => "usb-hid-mouse", mouse => "/system/drivers/usb-hid-mouse",
storage => "usb-storage", storage => "/system/drivers/usb-storage",
else => null, else => null,
}; };
} }
@@ -132,7 +124,7 @@ const DriverState = enum {
const Driver = struct { const Driver = struct {
used: bool = false, used: bool = false,
name_buffer: [24]u8 = undefined, name_buffer: [64]u8 = undefined, // fits a full binary path (abi.maximum_process_name)
name_len: usize = 0, name_len: usize = 0,
// The assigned device id (becomes argv[1]), or protocol.no_device. // The assigned device id (becomes argv[1]), or protocol.no_device.
device_id: u64 = protocol.no_device, device_id: u64 = protocol.no_device,
@@ -223,7 +215,7 @@ fn addChild(parent: u64, bus_address: u64, identity: u64, device_id: u64, report
fn pruneChildrenOf(reporter: u32) void { fn pruneChildrenOf(reporter: u32) void {
for (&children) |*child| { for (&children) |*child| {
if (child.used and child.reporter == reporter) { if (child.used and child.reporter == reporter) {
writeLine("/system/services/device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address }); std.log.info("child removed (device {d} port {d})", .{ child.parent, child.bus_address });
child.used = false; child.used = false;
const event = protocol.ChildRemoved{ .parent = child.parent, .bus_address = child.bus_address }; const event = protocol.ChildRemoved{ .parent = child.parent, .bus_address = child.bus_address };
publishEvent(std.mem.asBytes(&event)); publishEvent(std.mem.asBytes(&event));
@@ -270,7 +262,7 @@ fn addDriver(name: []const u8, device_id: u64, speaks_protocol: bool) void {
spawnDriver(driver); spawnDriver(driver);
return; return;
} }
writeLine("/system/services/device-manager: driver table full; cannot supervise {s}\n", .{name}); std.log.info("driver table full; cannot supervise {s}", .{name});
} }
/// (Re)spawn a driver instance: supervised on the manager's own endpoint, the /// (Re)spawn a driver instance: supervised on the manager's own endpoint, the
@@ -285,7 +277,7 @@ fn spawnDriver(driver: *Driver) void {
argument_count = 1; argument_count = 1;
} }
const child = system.spawnSupervised(driver.name(), arguments[0..argument_count], manager_endpoint) orelse { const child = system.spawnSupervised(driver.name(), arguments[0..argument_count], manager_endpoint) orelse {
writeLine("/system/services/device-manager: failed to spawn {s}\n", .{driver.name()}); std.log.info("failed to spawn {s}", .{driver.name()});
driver.state = .failed; driver.state = .failed;
return; return;
}; };
@@ -299,9 +291,9 @@ fn spawnDriver(driver: *Driver) void {
driver.state = .running; driver.state = .running;
} }
if (driver.device_id != protocol.no_device) { if (driver.device_id != protocol.no_device) {
writeLine("/system/services/device-manager: spawned {s} for device {d}\n", .{ driver.name(), driver.device_id }); std.log.info("spawned {s} for device {d}", .{ driver.name(), driver.device_id });
} else { } else {
writeLine("/system/services/device-manager: spawned {s}\n", .{driver.name()}); std.log.info("spawned {s}", .{driver.name()});
} }
} }
@@ -313,7 +305,7 @@ fn onDriverExit(driver: *Driver) void {
const reason = runtime.process.exitReason(driver.process_id) orelse .fault; const reason = runtime.process.exitReason(driver.process_id) orelse .fault;
if (reason == .exited) { if (reason == .exited) {
driver.state = .stopped; driver.state = .stopped;
writeLine("/system/services/device-manager: {s} exited cleanly; not restarting\n", .{driver.name()}); std.log.info("{s} exited cleanly; not restarting", .{driver.name()});
return; return;
} }
const now = system.clock(); const now = system.clock();
@@ -321,13 +313,13 @@ fn onDriverExit(driver: *Driver) void {
driver.restarts = if (alive_ns < fast_death_ns) driver.restarts + 1 else 1; driver.restarts = if (alive_ns < fast_death_ns) driver.restarts + 1 else 1;
if (driver.restarts >= crash_loop_cap) { if (driver.restarts >= crash_loop_cap) {
driver.state = .failed; driver.state = .failed;
writeLine("/system/services/device-manager: {s} is failing repeatedly (crash loop); giving up\n", .{driver.name()}); std.log.info("{s} is failing repeatedly (crash loop); giving up", .{driver.name()});
return; return;
} }
const delay_ms = backoff_base_ms << @intCast(driver.restarts - 1); const delay_ms = backoff_base_ms << @intCast(driver.restarts - 1);
driver.state = .restarting; driver.state = .restarting;
driver.restart_due_ns = now + delay_ms * 1_000_000; driver.restart_due_ns = now + delay_ms * 1_000_000;
writeLine("/system/services/device-manager: restarting {s} in {d} ms (died: {s})\n", .{ driver.name(), delay_ms, @tagName(reason) }); std.log.info("restarting {s} in {d} ms (died: {s})", .{ driver.name(), delay_ms, @tagName(reason) });
_ = system.timerOnce(manager_endpoint, delay_ms + 50); _ = system.timerOnce(manager_endpoint, delay_ms + 50);
} }
@@ -338,7 +330,7 @@ fn onDriverExit(driver: *Driver) void {
fn sweepDeadlines() void { fn sweepDeadlines() void {
const now = system.clock(); const now = system.clock();
if (test_kill_pid != 0 and now >= test_kill_due_ns) { if (test_kill_pid != 0 and now >= test_kill_due_ns) {
writeLine("/system/services/device-manager: test mode: killing the reporter\n", .{}); std.log.info("test mode: killing the reporter", .{});
_ = system.kill(test_kill_pid); _ = system.kill(test_kill_pid);
test_kill_pid = 0; test_kill_pid = 0;
} }
@@ -346,7 +338,7 @@ fn sweepDeadlines() void {
if (!driver.used) continue; if (!driver.used) continue;
switch (driver.state) { switch (driver.state) {
.awaiting_hello => if (now >= driver.hello_deadline_ns) { .awaiting_hello => if (now >= driver.hello_deadline_ns) {
writeLine("/system/services/device-manager: {s} missed its hello deadline\n", .{driver.name()}); std.log.info("{s} missed its hello deadline", .{driver.name()});
_ = system.kill(driver.process_id); _ = system.kill(driver.process_id);
// The exit notification finishes the job via onDriverExit. // The exit notification finishes the job via onDriverExit.
}, },
@@ -423,13 +415,13 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime
var status: i32 = 0; var status: i32 = 0;
if (hello.version != protocol.version) { if (hello.version != protocol.version) {
status = -1; status = -1;
writeLine("/system/services/device-manager: refused hello (version {d}) from process {d}\n", .{ hello.version, sender }); std.log.info("refused hello (version {d}) from process {d}", .{ hello.version, sender });
} else if (driverByProcess(sender)) |driver| { } else if (driverByProcess(sender)) |driver| {
driver.state = .running; driver.state = .running;
writeLine("/system/services/device-manager: hello from {s} (device {d})\n", .{ driver.name(), hello.device_id }); std.log.info("hello from {s} (device {d})", .{ driver.name(), hello.device_id });
// Resilience drill (V6): once, kill the virtio-gpu driver a moment after it hellos, so // Resilience drill (V6): once, kill the virtio-gpu driver a moment after it hellos, so
// the normal restart policy respawns it — the compositor must survive and re-attach. // the normal restart policy respawns it — the compositor must survive and re-attach.
if (test_scanout_restart_mode and !test_scanout_killed and std.mem.eql(u8, driver.name(), "virtio-gpu")) { if (test_scanout_restart_mode and !test_scanout_killed and std.mem.eql(u8, driver.name(), "/system/drivers/virtio-gpu")) {
test_scanout_killed = true; test_scanout_killed = true;
test_kill_pid = sender; test_kill_pid = sender;
test_kill_due_ns = system.clock() + 1_500_000_000; test_kill_due_ns = system.clock() + 1_500_000_000;
@@ -437,7 +429,7 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime
} }
} else { } else {
status = -1; status = -1;
writeLine("/system/services/device-manager: hello from unknown process {d}\n", .{sender}); std.log.info("hello from unknown process {d}", .{sender});
} }
const hello_reply = protocol.HelloReply{ .status = status }; const hello_reply = protocol.HelloReply{ .status = status };
@memcpy(reply[0..protocol.reply_size], std.mem.asBytes(&hello_reply)); @memcpy(reply[0..protocol.reply_size], std.mem.asBytes(&hello_reply));
@@ -453,7 +445,7 @@ fn onChildAdded(message: []const u8, reply: []u8, sender: u32) usize {
var status: i32 = 0; var status: i32 = 0;
if (driverByProcess(sender)) |driver| { if (driverByProcess(sender)) |driver| {
if (!addChild(report.parent, report.bus_address, report.identity, report.device_id, sender)) status = -1; if (!addChild(report.parent, report.bus_address, report.identity, report.device_id, sender)) status = -1;
writeLine("/system/services/device-manager: child added (device {d} port {d}, identity {d}) by {s}\n", .{ report.parent, report.bus_address, report.identity, driver.name() }); std.log.info("child added (device {d} port {d}, identity {d}) by {s}", .{ report.parent, report.bus_address, report.identity, driver.name() });
if (status == 0) publishEvent(message[0..protocol.child_added_size]); if (status == 0) publishEvent(message[0..protocol.child_added_size]);
// Matching from reports (M19.3): a registered child whose identity // Matching from reports (M19.3): a registered child whose identity
// names a driver gets one, once — re-reports after a bus restart // names a driver gets one, once — re-reports after a bus restart
@@ -498,7 +490,7 @@ fn onChildAdded(message: []const u8, reply: []u8, sender: u32) usize {
// Only the xHCI reporter is the drill's victim — pci-bus also reports // Only the xHCI reporter is the drill's victim — pci-bus also reports
// now, and whichever finishes second must not trigger the kill. // now, and whichever finishes second must not trigger the kill.
if (driverByProcess(sender)) |driver| { if (driverByProcess(sender)) |driver| {
if (std.mem.eql(u8, driver.name(), "usb-xhci-bus")) { if (std.mem.eql(u8, driver.name(), "/system/drivers/usb-xhci-bus")) {
// Delayed, not immediate: the device-list scenario's subscriber // Delayed, not immediate: the device-list scenario's subscriber
// needs a window to enumerate and subscribe before the events. // needs a window to enumerate and subscribe before the events.
test_usb_killed = true; test_usb_killed = true;
@@ -519,7 +511,7 @@ fn onChildRemoved(message: []const u8, reply: []u8, sender: u32) usize {
var status: i32 = -1; var status: i32 = -1;
for (&children) |*child| { for (&children) |*child| {
if (child.used and child.parent == report.parent and child.bus_address == report.bus_address and child.reporter == sender) { if (child.used and child.parent == report.parent and child.bus_address == report.bus_address and child.reporter == sender) {
writeLine("/system/services/device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address }); std.log.info("child removed (device {d} port {d})", .{ child.parent, child.bus_address });
child.used = false; child.used = false;
status = 0; status = 0;
} }
+62 -26
View File
@@ -16,8 +16,10 @@ const scanout_protocol = runtime.scanout_protocol;
const Rect = compositor.Rect; const Rect = compositor.Rect;
const Surface = compositor.Surface; const Surface = compositor.Surface;
/// The current display mode, as a backend reports it. /// The current display mode, as a backend reports it. `refresh_hz` is the panel's
pub const Info = struct { width: u32, height: u32, pitch: u32, format: u32 }; /// refresh rate from EDID (0 = unknown) — the frame clock's pacing seed; without vblank
/// it fixes the rate, never the phase (docs/display-v2.md, "Fenced is not vsync").
pub const Info = struct { width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32 };
/// Enumeration scratch — a `DeviceDescriptor` is large, and only one scan is ever needed. /// Enumeration scratch — a `DeviceDescriptor` is large, and only one scan is ever needed.
var device_table: [64]device.DeviceDescriptor = undefined; var device_table: [64]device.DeviceDescriptor = undefined;
@@ -26,7 +28,7 @@ var device_table: [64]device.DeviceDescriptor = undefined;
/// framebuffer write-combining as the front buffer, and keeps a cacheable back buffer of /// framebuffer write-combining as the front buffer, and keeps a cacheable back buffer of
/// the same geometry as the compose target. `present` streams the damaged rectangle from /// the same geometry as the compose target. `present` streams the damaged rectangle from
/// the back buffer to the LFB (sequential WC writes; the LFB is never read). No mode-set, /// the back buffer to the LFB (sequential WC writes; the LFB is never read). No mode-set,
/// no vsync — the portable floor (docs/display-v2.md). /// no present fence — the portable floor (docs/display-v2.md).
pub const Gop = struct { pub const Gop = struct {
device_id: u64, device_id: u64,
front: [*]volatile u8, // the LFB (write-combining) front: [*]volatile u8, // the LFB (write-combining)
@@ -35,13 +37,14 @@ pub const Gop = struct {
height: u32, height: u32,
pitch: u32, pitch: u32,
format: u32, format: u32,
refresh_hz: u32, // from the boot EDID via the display0 node (0 = unknown)
/// The framebuffer's id and geometry, captured together. `findDisplay` reads these out of /// The framebuffer's id and geometry, captured together. `findDisplay` reads these out of
/// the enumeration table and returns them by value, so the caller never re-reads the table /// the enumeration table and returns them by value, so the caller never re-reads the table
/// across later syscalls (`device_enumerate` writes the whole table straight into this /// across later syscalls (`device_enumerate` writes the whole table straight into this
/// process's memory; reading a descriptor's tail again after other syscalls have run is a /// process's memory; reading a descriptor's tail again after other syscalls have run is a
/// window we simply avoid by copying the few fields we need up front). /// window we simply avoid by copying the few fields we need up front).
const Found = struct { id: u64, width: u32, height: u32, pitch: u32, format: u32 }; const Found = struct { id: u64, width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32 };
/// The first `display`-class device with a *valid* (non-zero) geometry, or null. A zero /// The first `display`-class device with a *valid* (non-zero) geometry, or null. A zero
/// geometry is treated as "not ready yet" so the caller retries — a real framebuffer always /// geometry is treated as "not ready yet" so the caller retries — a real framebuffer always
@@ -52,7 +55,7 @@ pub const Gop = struct {
for (device_table[0..n]) |*d| { for (device_table[0..n]) |*d| {
if (d.class != @intFromEnum(device.DeviceClass.display)) continue; if (d.class != @intFromEnum(device.DeviceClass.display)) continue;
if (d.display.width == 0 or d.display.height == 0 or d.display.pitch == 0) continue; if (d.display.width == 0 or d.display.height == 0 or d.display.pitch == 0) continue;
return .{ .id = d.id, .width = d.display.width, .height = d.display.height, .pitch = d.display.pitch, .format = d.display.format }; return .{ .id = d.id, .width = d.display.width, .height = d.display.height, .pitch = d.display.pitch, .format = d.display.format, .refresh_hz = d.display.refresh_hz };
} }
return null; return null;
} }
@@ -93,11 +96,12 @@ pub const Gop = struct {
.height = found.height, .height = found.height,
.pitch = found.pitch, .pitch = found.pitch,
.format = found.format, .format = found.format,
.refresh_hz = found.refresh_hz,
}; };
} }
pub fn info(self: *const Gop) Info { pub fn info(self: *const Gop) Info {
return .{ .width = self.width, .height = self.height, .pitch = self.pitch, .format = self.format }; return .{ .width = self.width, .height = self.height, .pitch = self.pitch, .format = self.format, .refresh_hz = self.refresh_hz };
} }
/// The cacheable compose target (the back buffer). /// The cacheable compose target (the back buffer).
@@ -110,27 +114,53 @@ pub const Gop = struct {
}; };
} }
/// Stream the damaged rectangle from the back buffer to the write-combining LFB, row by /// Stream each damaged rectangle from the back buffer to the write-combining LFB, row
/// row (sequential writes — what WC memory wants; the LFB is never read). /// by row (sequential writes — what WC memory wants; the LFB is never read). The rows
pub fn present(self: *const Gop, damage: Rect) void { /// are copied by `presentSpan` below, which widens the stores by hand: `volatile`
const c = damage.intersect(.{ .x = 0, .y = 0, .w = @intCast(self.width), .h = @intCast(self.height) }); /// keeps the compiler from eliding or reordering framebuffer writes, but it also
if (c.isEmpty()) return; /// forbids it from merging them, so a naive per-pixel loop is stuck at one 4-byte
var y: i32 = c.y; /// store per iteration. Keeping each copy small (the damage list) and each store wide
while (y < c.bottom()) : (y += 1) { /// shrinks the window in which scanout can sample a half-written frame.
const off = @as(usize, @intCast(y)) * self.pitch; pub fn present(self: *const Gop, damage: []const Rect) void {
const src: [*]const u32 = @ptrCast(@alignCast(self.back + off)); const bounds = Rect{ .x = 0, .y = 0, .w = @intCast(self.width), .h = @intCast(self.height) };
const dst: [*]volatile u32 = @ptrCast(@alignCast(self.front + off)); for (damage) |rect| {
var x: i32 = c.x; const c = rect.intersect(bounds);
while (x < c.right()) : (x += 1) dst[@intCast(x)] = src[@intCast(x)]; if (c.isEmpty()) continue;
const span: usize = @intCast(c.w);
var y: i32 = c.y;
while (y < c.bottom()) : (y += 1) {
const offset = @as(usize, @intCast(y)) * self.pitch + @as(usize, @intCast(c.x)) * 4;
const source: [*]const u32 = @ptrCast(@alignCast(self.back + offset));
const front_row: [*]volatile u32 = @ptrCast(@alignCast(self.front + offset));
presentSpan(front_row, source, span);
}
} }
} }
}; };
/// Copy `count` pixels into the write-combining front buffer with 8-byte volatile stores
/// (plus a 4-byte head/tail where the span isn't 8-aligned — pixel spans are always
/// 4-aligned). The loads come from the cacheable back buffer and are assembled into a
/// `u64` in registers, so nothing here reads the front buffer.
fn presentSpan(destination: [*]volatile u32, source: [*]const u32, count: usize) void {
var i: usize = 0;
if (i < count and (@intFromPtr(destination) & 7) != 0) {
destination[0] = source[0];
i = 1;
}
while (i + 2 <= count) : (i += 2) {
const pair = @as(u64, source[i]) | (@as(u64, source[i + 1]) << 32);
const wide: *volatile u64 = @ptrCast(@alignCast(destination + i));
wide.* = pair;
}
if (i < count) destination[i] = source[i];
}
/// A display mode the native backend can switch to. /// A display mode the native backend can switch to.
pub const Mode = scanout_protocol.Mode; pub const Mode = scanout_protocol.Mode;
/// The native virtio-gpu backend: the compositor composes into a **shared** scanout surface /// The native virtio-gpu backend: the compositor composes into a **shared** scanout surface
/// (an `shm` region the driver created and handed over) and `present` asks the driver to put /// (a shared-memory region the driver created and handed over) and `present` asks the driver to put
/// a frame on the panel over its `.scanout` endpoint. Unlike GOP there is no local copy — the /// a frame on the panel over its `.scanout` endpoint. Unlike GOP there is no local copy — the
/// surface *is* the device's resource backing, so compositing writes land straight where the /// surface *is* the device's resource backing, so compositing writes land straight where the
/// driver transfers-and-flushes from (x86 DMA is cache-coherent, so the cacheable shared pages /// driver transfers-and-flushes from (x86 DMA is cache-coherent, so the cacheable shared pages
@@ -143,17 +173,19 @@ pub const VirtioGpu = struct {
width: u32, // the active mode width: u32, // the active mode
height: u32, height: u32,
format: u32, format: u32,
refresh_hz: u32, // from the driver's EDID read, carried in the announce (0 = unknown)
scanout: ipc.Handle, // the driver's present + mode channel (looked up on `.scanout`) scanout: ipc.Handle, // the driver's present + mode channel (looked up on `.scanout`)
pub fn info(self: *const VirtioGpu) Info { pub fn info(self: *const VirtioGpu) Info {
return .{ .width = self.width, .height = self.height, .pitch = self.stride * 4, .format = self.format }; return .{ .width = self.width, .height = self.height, .pitch = self.stride * 4, .format = self.format, .refresh_hz = self.refresh_hz };
} }
pub fn surface(self: *const VirtioGpu) Surface { pub fn surface(self: *const VirtioGpu) Surface {
return .{ .pixels = self.pixels, .stride = self.stride, .width = self.width, .height = self.height }; return .{ .pixels = self.pixels, .stride = self.stride, .width = self.width, .height = self.height };
} }
/// Ask the driver to present. The composited pixels are already in the shared surface, so /// Ask the driver to present. The composited pixels are already in the shared surface, so
/// this is a single request over `.scanout`; the driver transfers + fenced-flushes. /// this is a single request over `.scanout` regardless of how many damage rectangles
pub fn present(self: *const VirtioGpu, damage: Rect) void { /// accumulated; the driver transfers + fenced-flushes the whole frame.
pub fn present(self: *const VirtioGpu, damage: []const Rect) void {
_ = damage; _ = damage;
var request = scanout_protocol.Request{ var request = scanout_protocol.Request{
.operation = @intFromEnum(scanout_protocol.Operation.present), .operation = @intFromEnum(scanout_protocol.Operation.present),
@@ -210,7 +242,7 @@ pub const Backend = union(enum) {
inline else => |*b| b.surface(), inline else => |*b| b.surface(),
}; };
} }
pub fn present(self: *const Backend, damage: Rect) void { pub fn present(self: *const Backend, damage: []const Rect) void {
switch (self.*) { switch (self.*) {
inline else => |*b| b.present(damage), inline else => |*b| b.present(damage),
} }
@@ -236,9 +268,13 @@ pub const Backend = union(enum) {
.virtio => true, .virtio => true,
}; };
} }
/// Whether this backend has a vblank/fence for tear-free present (virtio-gpu: yes, V5 — every /// Whether this backend's present is **fenced** — it completes only once the device has
/// flush is fenced, so the device signals completion when the frame is actually on screen). /// consumed the frame (virtio-gpu: every flush carries a fence the used-ring ack waits on).
pub fn hasVsync(self: *const Backend) bool { /// A fence gives completion feedback and tear-free snapshot presents; it is *not* vblank —
/// nothing paces presents to the display's refresh (base virtio-gpu 2D has no vblank event
/// at all). True vsync needs a native driver's vblank interrupt. See docs/display-v2.md,
/// "Fenced is not vsync".
pub fn hasFencedPresent(self: *const Backend) bool {
return switch (self.*) { return switch (self.*) {
.gop => false, .gop => false,
.virtio => true, .virtio => true,
+285 -25
View File
@@ -57,6 +57,177 @@ pub const Rect = struct {
} }
}; };
/// The dirty screen regions accumulated between presents. Kept as a *list* of rectangles,
/// not one bounding box: when two small things move far apart — the cursor on one side of
/// the screen, an animating layer on the other — a single bounding box unites them into a
/// huge region, and presenting it streams megabytes to the framebuffer for a few thousand
/// changed pixels. The long copy widens the window in which scanout (or QEMU's display
/// refresh) samples a half-written frame — visible as tearing and cursor trails. Small
/// separate rectangles keep each copy, and that window, tight.
///
/// A new rectangle that overlaps an existing entry is united into it (repainting a modest
/// superset is harmless — compositing is idempotent); the grown entry is *not* re-merged
/// against the rest, so entries may overlap, which costs only a duplicate repaint. When
/// the table is full the newcomer folds into the last entry — degrading toward the old
/// bounding-box behaviour instead of dropping damage.
pub const DamageList = struct {
pub const capacity = 16;
rects: [capacity]Rect = [_]Rect{Rect.empty} ** capacity,
count: usize = 0,
pub fn add(self: *DamageList, r: Rect) void {
if (r.isEmpty()) return;
for (self.rects[0..self.count]) |*existing| {
if (!existing.intersect(r).isEmpty()) {
existing.* = existing.unite(r);
return;
}
}
if (self.count < capacity) {
self.rects[self.count] = r;
self.count += 1;
return;
}
self.rects[capacity - 1] = self.rects[capacity - 1].unite(r);
}
pub fn isEmpty(self: *const DamageList) bool {
return self.count == 0;
}
pub fn slice(self: *const DamageList) []const Rect {
return self.rects[0..self.count];
}
pub fn clear(self: *DamageList) void {
self.count = 0;
}
};
/// The alternative damage tracker: a **fixed tile grid**, the scheme browser compositors
/// and tile-based GPUs use. The screen is divided into `tile_size`-pixel tiles up front;
/// `add` marks the tiles a rectangle touches (a bit per tile — merging is free and exact,
/// no heuristics), and `collect` walks the grid turning runs of adjacent dirty tiles into
/// repaint rectangles (horizontal runs, then equal-span rows merged vertically, so
/// full-screen damage collapses back to a single rectangle).
///
/// Trade-off against `DamageList`: tracking is O(1) with a strictly bounded worst case
/// (never more than the dirty tiles), but repaints are quantized — a 1-pixel change
/// repaints a whole tile. Which wins depends on the workload; the display service has a
/// compile-time switch (`damage_mode`) to compare them.
pub const TileGrid = struct {
pub const tile_size = 64;
pub const maximum_columns = 128; // supports screens up to 8192 px wide…
pub const maximum_rows = 128; // …and 8192 px tall (beyond that, edge tiles stretch)
pub const maximum_tiles = maximum_columns * maximum_rows;
/// The most rectangles `collect` produces; extras fold into the last (never dropped).
pub const maximum_rects = 64;
width: u32 = 0,
height: u32 = 0,
columns: u32 = 0,
rows: u32 = 0,
dirty_count: u32 = 0,
dirty: [maximum_tiles]bool = [_]bool{false} ** maximum_tiles,
/// Size the grid for a screen. Also clears it — callers reset on a geometry change,
/// where the mode-set paths damage the whole new screen anyway.
pub fn reset(self: *TileGrid, width: u32, height: u32) void {
self.width = width;
self.height = height;
self.columns = @min((width + tile_size - 1) / tile_size, maximum_columns);
self.rows = @min((height + tile_size - 1) / tile_size, maximum_rows);
self.clear();
}
pub fn matches(self: *const TileGrid, width: u32, height: u32) bool {
return self.width == width and self.height == height;
}
pub fn isEmpty(self: *const TileGrid) bool {
return self.dirty_count == 0;
}
pub fn clear(self: *TileGrid) void {
@memset(&self.dirty, false);
self.dirty_count = 0;
}
/// Mark every tile `r` touches. Clips to the screen first, so out-of-range
/// rectangles are harmless.
pub fn add(self: *TileGrid, r: Rect) void {
const screen = Rect{ .x = 0, .y = 0, .w = @intCast(self.width), .h = @intCast(self.height) };
const c = r.intersect(screen);
if (c.isEmpty()) return;
const column_first: u32 = @intCast(@divTrunc(c.x, tile_size));
const row_first: u32 = @intCast(@divTrunc(c.y, tile_size));
const column_last: u32 = @min(@as(u32, @intCast(@divTrunc(c.right() - 1, tile_size))), self.columns - 1);
const row_last: u32 = @min(@as(u32, @intCast(@divTrunc(c.bottom() - 1, tile_size))), self.rows - 1);
var row = row_first;
while (row <= row_last) : (row += 1) {
var column = column_first;
while (column <= column_last) : (column += 1) {
const index = row * self.columns + column;
if (!self.dirty[index]) {
self.dirty[index] = true;
self.dirty_count += 1;
}
}
}
}
/// The screen rectangle covered by tiles [column_first, column_end) of `row`. Edge
/// tiles clamp to the true screen size (the last column/row may be partial — or, on a
/// screen wider than the grid supports, stretched to cover the remainder).
fn tileSpanRect(self: *const TileGrid, column_first: u32, column_end: u32, row: u32) Rect {
const x: i32 = @intCast(column_first * tile_size);
const y: i32 = @intCast(row * tile_size);
const right: i32 = if (column_end >= self.columns) @intCast(self.width) else @intCast(column_end * tile_size);
const bottom: i32 = if (row + 1 >= self.rows) @intCast(self.height) else @intCast((row + 1) * tile_size);
return .{ .x = x, .y = y, .w = right - x, .h = bottom - y };
}
/// Turn the dirty tiles into repaint rectangles in `out`: coalesce each row's runs of
/// adjacent dirty tiles, then merge a run into the rectangle directly above it when
/// the spans match — so a dirty block of tiles becomes one rectangle. Returns the
/// filled prefix of `out`.
pub fn collect(self: *const TileGrid, out: []Rect) []Rect {
var count: usize = 0;
var row: u32 = 0;
while (row < self.rows) : (row += 1) {
var column: u32 = 0;
while (column < self.columns) {
if (!self.dirty[row * self.columns + column]) {
column += 1;
continue;
}
var run_end = column + 1;
while (run_end < self.columns and self.dirty[row * self.columns + run_end]) run_end += 1;
const rect = self.tileSpanRect(column, run_end, row);
column = run_end;
var merged = false;
for (out[0..count]) |*existing| {
if (existing.x == rect.x and existing.w == rect.w and existing.bottom() == rect.y) {
existing.h += rect.h;
merged = true;
break;
}
}
if (merged) continue;
if (count < out.len) {
out[count] = rect;
count += 1;
} else {
out[count - 1] = out[count - 1].unite(rect);
}
}
}
return out[0..count];
}
};
/// A block of 32-bit pixels: `pixels` addressed row-major with `stride` pixels between /// A block of 32-bit pixels: `pixels` addressed row-major with `stride` pixels between
/// row starts (≥ width — the framebuffer's stride is pitch/4, a layer's is its width). /// row starts (≥ width — the framebuffer's stride is pitch/4, a layer's is its width).
pub const Surface = struct { pub const Surface = struct {
@@ -74,15 +245,17 @@ pub const Surface = struct {
} }
}; };
/// Fill `rect` of `s` with the native pixel `colour`, clipped to `s`'s bounds. /// Fill `rect` of `s` with the native pixel `colour`, clipped to `s`'s bounds. Each row is
/// one `@memset` over the clipped span, so the compiler vectorizes it and the bounds check
/// runs once per row, not once per pixel.
pub fn fillRect(s: Surface, rect: Rect, colour: u32) void { pub fn fillRect(s: Surface, rect: Rect, colour: u32) void {
const c = rect.intersect(s.bounds()); const c = rect.intersect(s.bounds());
if (c.isEmpty()) return; if (c.isEmpty()) return;
const x0: usize = @intCast(c.x);
const span: usize = @intCast(c.w);
var y: i32 = c.y; var y: i32 = c.y;
while (y < c.bottom()) : (y += 1) { while (y < c.bottom()) : (y += 1) {
const r = s.row(@intCast(y)); @memset((s.row(@intCast(y)) + x0)[0..span], colour);
var x: i32 = c.x;
while (x < c.right()) : (x += 1) r[@intCast(x)] = colour;
} }
} }
@@ -94,35 +267,36 @@ pub fn composite(dst: Surface, dx: i32, dy: i32, layer: Surface, clip: Rect) voi
const on_screen = Rect{ .x = dx, .y = dy, .w = @intCast(layer.width), .h = @intCast(layer.height) }; const on_screen = Rect{ .x = dx, .y = dy, .w = @intCast(layer.width), .h = @intCast(layer.height) };
const region = on_screen.intersect(clip).intersect(dst.bounds()); const region = on_screen.intersect(clip).intersect(dst.bounds());
if (region.isEmpty()) return; if (region.isEmpty()) return;
const span: usize = @intCast(region.w);
const dst_x: usize = @intCast(region.x);
const src_x: usize = @intCast(region.x - dx);
var y: i32 = region.y; var y: i32 = region.y;
while (y < region.bottom()) : (y += 1) { while (y < region.bottom()) : (y += 1) {
const src = layer.row(@intCast(y - dy)); const source_row = layer.row(@intCast(y - dy)) + src_x;
const d = dst.row(@intCast(y)); const destination_row = dst.row(@intCast(y)) + dst_x;
var x: i32 = region.x; @memcpy(destination_row[0..span], source_row[0..span]);
while (x < region.right()) : (x += 1) {
d[@intCast(x)] = src[@intCast(x - dx)];
}
} }
} }
/// Copy a `w`×`h` tile of native pixels from `src` (raw little-endian bytes, row-major, /// Copy a `w`×`h` tile of native pixels from `src` (raw little-endian bytes, row-major,
/// tightly packed) into `dst` at (`dx`, `dy`), clipped to `dst`'s bounds. `src` is read /// tightly packed) into `dst` at (`dx`, `dy`), clipped to `dst`'s bounds. `src` comes
/// with `readInt` because it comes straight out of an IPC message buffer and carries no /// straight out of an IPC message buffer and carries no alignment guarantee, so each
/// alignment guarantee. Returns without touching anything if `src` is short. /// clipped row is a byte-wise `@memcpy` — which equals the old per-pixel little-endian
/// `readInt` on every danos target (all little-endian) without the alignment concern.
/// Returns without touching anything if `src` is short.
pub fn blitTile(dst: Surface, dx: i32, dy: i32, src: []const u8, w: u32, h: u32) void { pub fn blitTile(dst: Surface, dx: i32, dy: i32, src: []const u8, w: u32, h: u32) void {
if (src.len < @as(usize, w) * h * 4) return; if (src.len < @as(usize, w) * h * 4) return;
var ty: u32 = 0; const region = Rect.init(dx, dy, @intCast(w), @intCast(h)).intersect(dst.bounds());
while (ty < h) : (ty += 1) { if (region.isEmpty()) return;
const yy = dy + @as(i32, @intCast(ty)); const span: usize = @intCast(region.w);
if (yy < 0 or yy >= dst.height) continue; const tile_x: usize = @intCast(region.x - dx);
const drow = dst.row(@intCast(yy)); const dst_x: usize = @intCast(region.x);
var tx: u32 = 0; var y: i32 = region.y;
while (tx < w) : (tx += 1) { while (y < region.bottom()) : (y += 1) {
const xx = dx + @as(i32, @intCast(tx)); const tile_y: usize = @intCast(y - dy);
if (xx < 0 or xx >= dst.width) continue; const offset = (tile_y * w + tile_x) * 4;
const off = (@as(usize, ty) * w + tx) * 4; const destination_row = dst.row(@intCast(y)) + dst_x;
drow[@intCast(xx)] = std.mem.readInt(u32, src[off..][0..4], .little); @memcpy(std.mem.sliceAsBytes(destination_row[0..span]), src[offset..][0 .. span * 4]);
}
} }
} }
@@ -179,6 +353,92 @@ test "composite honours the damage rectangle" {
try std.testing.expectEqual(@as(u32, 0), back[4 * 8 + 4]); // outside damage try std.testing.expectEqual(@as(u32, 0), back[4 * 8 + 4]); // outside damage
} }
test "damage list keeps disjoint rectangles separate and merges overlap" {
var list = DamageList{};
list.add(Rect.init(0, 0, 10, 10));
list.add(Rect.init(100, 100, 10, 10)); // far away: its own entry
try std.testing.expectEqual(@as(usize, 2), list.slice().len);
list.add(Rect.init(5, 5, 10, 10)); // overlaps the first: united into it
try std.testing.expectEqual(@as(usize, 2), list.slice().len);
try std.testing.expectEqual(Rect.init(0, 0, 15, 15), list.slice()[0]);
try std.testing.expect(!list.isEmpty());
list.clear();
try std.testing.expect(list.isEmpty());
}
test "damage list folds overflow into the last entry instead of dropping it" {
var list = DamageList{};
var i: i32 = 0;
while (i < DamageList.capacity) : (i += 1) {
list.add(Rect.init(i * 100, 0, 10, 10)); // disjoint: fills every slot
}
try std.testing.expectEqual(@as(usize, DamageList.capacity), list.slice().len);
const overflow = Rect.init(0, 5000, 10, 10);
list.add(overflow);
try std.testing.expectEqual(@as(usize, DamageList.capacity), list.slice().len);
const last = list.slice()[DamageList.capacity - 1];
try std.testing.expect(!last.intersect(overflow).isEmpty()); // still covered
}
test "damage list ignores empty rectangles" {
var list = DamageList{};
list.add(Rect.empty);
try std.testing.expect(list.isEmpty());
}
test "tile grid coalesces a run of adjacent tiles into one rectangle" {
var grid = TileGrid{};
grid.reset(256, 128); // 4×2 tiles of 64 px
grid.add(Rect.init(10, 10, 100, 10)); // spans tiles (0,0) and (1,0)
var scratch: [TileGrid.maximum_rects]Rect = undefined;
const rects = grid.collect(&scratch);
try std.testing.expectEqual(@as(usize, 1), rects.len);
try std.testing.expectEqual(Rect.init(0, 0, 128, 64), rects[0]);
}
test "tile grid: full-screen damage collapses back to a single rectangle" {
var grid = TileGrid{};
grid.reset(1280, 720); // 20×12 tiles; the bottom row is partial (720 = 11*64 + 16)
grid.add(Rect.init(0, 0, 1280, 720));
var scratch: [TileGrid.maximum_rects]Rect = undefined;
const rects = grid.collect(&scratch);
try std.testing.expectEqual(@as(usize, 1), rects.len);
try std.testing.expectEqual(Rect.init(0, 0, 1280, 720), rects[0]);
}
test "tile grid keeps far-apart damage as separate rectangles" {
var grid = TileGrid{};
grid.reset(1280, 720);
grid.add(Rect.init(0, 0, 10, 10)); // top-left tile
grid.add(Rect.init(1000, 600, 10, 10)); // a far-away tile
var scratch: [TileGrid.maximum_rects]Rect = undefined;
const rects = grid.collect(&scratch);
try std.testing.expectEqual(@as(usize, 2), rects.len);
}
test "tile grid clamps edge tiles to the true screen size" {
var grid = TileGrid{};
grid.reset(100, 100); // 2×2 tiles, both partial in each axis
grid.add(Rect.init(0, 0, 100, 100));
var scratch: [TileGrid.maximum_rects]Rect = undefined;
const rects = grid.collect(&scratch);
try std.testing.expectEqual(@as(usize, 1), rects.len);
try std.testing.expectEqual(Rect.init(0, 0, 100, 100), rects[0]);
}
test "tile grid clear empties it and reset resizes it" {
var grid = TileGrid{};
grid.reset(256, 256);
grid.add(Rect.init(0, 0, 256, 256));
try std.testing.expect(!grid.isEmpty());
grid.clear();
try std.testing.expect(grid.isEmpty());
try std.testing.expect(grid.matches(256, 256));
grid.reset(512, 512);
try std.testing.expect(!grid.matches(256, 256));
try std.testing.expect(grid.isEmpty());
}
test "blitTile copies a packed tile, clipping and reading unaligned bytes" { test "blitTile copies a packed tile, clipping and reading unaligned bytes" {
var back = [_]u32{0} ** (4 * 4); var back = [_]u32{0} ** (4 * 4);
const dst = Surface{ .pixels = &back, .stride = 4, .width = 4, .height = 4 }; const dst = Surface{ .pixels = &back, .stride = 4, .width = 4, .height = 4 };
+123 -40
View File
@@ -10,7 +10,9 @@
//! z-order, and visibility. Clients create layers, draw into them by command (`fill_rect`, //! z-order, and visibility. Clients create layers, draw into them by command (`fill_rect`,
//! `blit_tile`), mark `damage`, and ask for a `present`; the compositor repaints only the //! `blit_tile`), mark `damage`, and ask for a `present`; the compositor repaints only the
//! damaged region — clear it, paint the visible layers bottom-to-top into the backend's //! damaged region — clear it, paint the visible layers bottom-to-top into the backend's
//! surface, then `backend.present(damage)`. Shared-memory client surfaces are later //! surface, then `backend.present(damage)`. Presents are paced by a ~60 Hz **frame clock**
//! (see `schedulePresent`), so any number of client presents and cursor moves inside one
//! interval coalesce into a single frame. Shared-memory client surfaces are later
//! (docs/display-v2.md). //! (docs/display-v2.md).
const std = @import("std"); const std = @import("std");
@@ -50,8 +52,10 @@ var pending_modeset_check: bool = false;
var background: u32 = 0; var background: u32 = 0;
/// The layer stack. A fixed table (a compositor has few top-level surfaces during /// The layer stack. A fixed table (a compositor has few top-level surfaces during
/// bring-up); each used slot owns an mmap'd surface. `damage` accumulates the dirty /// bring-up); each used slot owns an mmap'd surface. `damage_list` accumulates the dirty
/// screen region since the last `present`, so a present touches only what changed. /// screen rectangles since the last `present`, so a present touches only what changed —
/// and keeps far-apart changes (the cursor here, an animating layer there) as *separate*
/// small copies rather than one huge bounding box (see compositor.DamageList).
const maximum_layers = 16; const maximum_layers = 16;
const Layer = struct { const Layer = struct {
@@ -65,7 +69,67 @@ const Layer = struct {
}; };
var layers: [maximum_layers]Layer = [_]Layer{.{}} ** maximum_layers; var layers: [maximum_layers]Layer = [_]Layer{.{}} ** maximum_layers;
var damage: Rect = Rect.empty;
/// Which damage tracker drives `present` — a compile-time A/B switch (both are in
/// compositor.zig with the trade-off discussion):
/// .list — free-form dirty rectangles (tight bounds, heuristic merging)
/// .grid — a fixed 64-px tile grid (exact O(1) merging, tile-quantized repaints)
const DamageMode = enum { list, grid };
const damage_mode: DamageMode = .grid;
var damage_list: compositor.DamageList = .{};
var damage_grid: compositor.TileGrid = .{};
/// The **frame clock**: client `present` requests and cursor motion don't repaint
/// immediately — they accumulate damage and arm a one-shot timer, and the tick composites
/// everything pending as one frame. That paces presents to ~60 Hz no matter how fast
/// clients draw or the mouse moves (previously every mouse event became a full present).
/// No backend has a real vblank to pace by (docs/display-v2.md, "Fenced is not vsync");
/// this is the software stand-in, the same strategy Linux uses atop virtio-gpu. Bring-up
/// paths that need pixels on screen *now* (initialise, the self-checks) still call
/// `present()` directly.
///
/// The interval comes from the *active backend's* panel refresh rate (EDID: the loader
/// captures it for the GOP floor while firmware still runs; the native driver reads its
/// own and carries it in the announce). `updateFrameClock` re-derives it whenever the
/// backend changes — the boot framebuffer's clock dies with the GOP floor at upgrade.
/// Without a rate the clock defaults to 60 Hz, and it is clamped to [30, 120] Hz so a
/// mis-parsed EDID can neither starve nor flood the compositor.
var frame_interval_milliseconds: u64 = 16;
var frame_timer_armed = false;
/// Derive the frame-clock interval from the active backend's refresh rate and log what
/// the clock is now pacing to. Called at bring-up and again on every backend change.
fn updateFrameClock() void {
const reported = backend.info().refresh_hz;
const rate: u64 = if (reported == 0) 60 else @min(@max(reported, 30), 120);
frame_interval_milliseconds = @max(1000 / rate, 1);
var line: [96]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "display: frame clock {d} Hz ({s})\n", .{
1000 / frame_interval_milliseconds,
if (reported == 0) "default" else "panel EDID",
}) catch return);
}
/// Arm the frame clock unless a tick is already pending: any number of requests inside
/// one interval coalesce into that single tick's present.
fn schedulePresent() void {
if (frame_timer_armed) return;
frame_timer_armed = true;
_ = system.timerOnce(service_endpoint, frame_interval_milliseconds);
}
/// A timer landing — the frame clock, or the deferred first native present armed by
/// `attach_scanout`: present the accumulated damage, then run the one-shot mode-set
/// self-check if the native upgrade queued it.
fn frameTick() void {
frame_timer_armed = false;
present();
if (pending_modeset_check) {
pending_modeset_check = false;
modesetSelfCheck();
}
}
// --- geometry helpers ------------------------------------------------------- // --- geometry helpers -------------------------------------------------------
@@ -78,9 +142,20 @@ fn layerScreenRect(l: *const Layer) Rect {
return .{ .x = l.x, .y = l.y, .w = @intCast(l.surface.width), .h = @intCast(l.surface.height) }; return .{ .x = l.x, .y = l.y, .w = @intCast(l.surface.width), .h = @intCast(l.surface.height) };
} }
/// Add `r` (screen coordinates) to the pending damage, clipped to the screen. /// Add `r` (screen coordinates) to the pending damage, clipped to the screen. In grid
/// mode the grid re-sizes itself lazily when the screen geometry changes — every
/// geometry-changing path (`attach_scanout`, `set_mode`) damages the whole new screen
/// right after, so damage pending from the old geometry is safely superseded.
fn addDamage(r: Rect) void { fn addDamage(r: Rect) void {
damage = damage.unite(r.intersect(screenRect())); const clipped = r.intersect(screenRect());
switch (damage_mode) {
.list => damage_list.add(clipped),
.grid => {
const mode = backend.info();
if (!damage_grid.matches(mode.width, mode.height)) damage_grid.reset(mode.width, mode.height);
damage_grid.add(clipped);
},
}
} }
// --- layer operations (called from onMessage and the self-check) ------------ // --- layer operations (called from onMessage and the self-check) ------------
@@ -183,21 +258,29 @@ fn compositeInto(clip: Rect) void {
} }
} }
/// Composite the accumulated damage into the backend's surface, hand it to the backend to /// Composite each accumulated damage rectangle into the backend's surface, hand the list
/// put on screen, then clear the damage. A no-op when nothing is dirty. The frame counter /// to the backend to put on screen, then clear the damage. A no-op when nothing is dirty.
/// advances regardless, so callers can name frames. /// The frame counter advances regardless, so callers can name frames.
fn present() void { fn present() void {
const dirty = damage.intersect(screenRect()); var scratch: [compositor.TileGrid.maximum_rects]Rect = undefined;
if (!dirty.isEmpty()) { const dirty: []const Rect = switch (damage_mode) {
compositeInto(dirty); .list => damage_list.slice(),
.grid => damage_grid.collect(&scratch),
};
const had_damage = dirty.len != 0;
if (had_damage) {
for (dirty) |region| compositeInto(region);
backend.present(dirty); backend.present(dirty);
} }
damage = Rect.empty; switch (damage_mode) {
.list => damage_list.clear(),
.grid => damage_grid.clear(),
}
frames += 1; frames += 1;
// The first present after a native upgrade confirms the composited frame actually reached // The first present after a native upgrade confirms the composited frame actually reached
// the shared scanout surface (the automated stand-in for "it's on screen"). // the shared scanout surface (the automated stand-in for "it's on screen").
if (pending_native_verify and !dirty.isEmpty()) { if (pending_native_verify and had_damage) {
pending_native_verify = false; pending_native_verify = false;
verifyNativePresent(); verifyNativePresent();
} }
@@ -221,13 +304,13 @@ fn verifyNativePresent() void {
/// present channel, switch the backend to virtio-gpu, and queue a full-screen repaint. The /// present channel, switch the backend to virtio-gpu, and queue a full-screen repaint. The
/// present is deferred to a timer (see `service_endpoint`) so it happens after this reply /// present is deferred to a timer (see `service_endpoint`) so it happens after this reply
/// unblocks the driver and it starts serving `.scanout`. /// unblocks the driver and it starts serving `.scanout`.
fn attachScanout(stride: u32, width: u32, height: u32, format: u32, capability: ?ipc.Handle, reply: []u8) usize { fn attachScanout(stride: u32, width: u32, height: u32, format: u32, refresh_hz: u32, capability: ?ipc.Handle, reply: []u8) usize {
const cap = capability orelse return fail(reply); const cap = capability orelse return fail(reply);
if (width == 0 or height == 0 or stride < width) return fail(reply); if (width == 0 or height == 0 or stride < width) return fail(reply);
const mapped = runtime.shm.map(cap) orelse return fail(reply); const mapped = runtime.shared_memory.map(cap) orelse return fail(reply);
const scanout = ipc.lookup(.scanout) orelse return fail(reply); const scanout = ipc.lookup(.scanout) orelse return fail(reply);
// A second announce means the driver died and was restarted (V6): re-attach to its fresh // A second announce means the driver died and was restarted (V6): re-attach to its fresh
// scanout. (The previous shared mapping leaks — there is no shm_unmap syscall yet — but the // scanout. (The previous shared mapping leaks — there is no shared_memory_unmap syscall yet — but the
// frames are the dead driver's, reclaimed on its exit; a handful across a crash is benign.) // frames are the dead driver's, reclaimed on its exit; a handful across a crash is benign.)
const reattach = switch (backend) { const reattach = switch (backend) {
.virtio => true, .virtio => true,
@@ -240,9 +323,11 @@ fn attachScanout(stride: u32, width: u32, height: u32, format: u32, capability:
.width = width, .width = width,
.height = height, .height = height,
.format = format, .format = format,
.refresh_hz = refresh_hz,
.scanout = scanout, .scanout = scanout,
} }; } };
background = protocol.pack(format, 0x20, 0x30, 0x48); // re-pack the wallpaper for the mode background = protocol.pack(format, 0x20, 0x30, 0x48); // re-pack the wallpaper for the mode
updateFrameClock(); // the GOP floor's clock dies here — pace by the GPU's EDID now
addDamage(screenRect()); // the whole new surface must be painted addDamage(screenRect()); // the whole new surface must be painted
pending_native_verify = true; pending_native_verify = true;
if (!reattach) pending_modeset_check = true; // the mode-set self-check runs once, on first upgrade if (!reattach) pending_modeset_check = true; // the mode-set self-check runs once, on first upgrade
@@ -257,7 +342,8 @@ fn attachScanout(stride: u32, width: u32, height: u32, format: u32, capability:
/// After the native upgrade is verified, prove the runtime-resolution-change and fenced-present /// After the native upgrade is verified, prove the runtime-resolution-change and fenced-present
/// paths: query the driver's modes, switch to one that differs from the current, re-composite /// paths: query the driver's modes, switch to one that differs from the current, re-composite
/// the whole screen at the new size, and confirm the backend now reports that geometry. The /// the whole screen at the new size, and confirm the backend now reports that geometry. The
/// present goes through the driver's fenced flush, so a clean present is a vsync present. /// present goes through the driver's fenced flush, so a clean present is a *fenced* present —
/// completion-acknowledged and tear-free, not vblank-paced (docs/display-v2.md).
fn modesetSelfCheck() void { fn modesetSelfCheck() void {
if (!backend.canModeSet()) return; if (!backend.canModeSet()) return;
var mode_list: [4]backend_mod.Mode = undefined; var mode_list: [4]backend_mod.Mode = undefined;
@@ -289,7 +375,7 @@ fn modesetSelfCheck() void {
if (now.width == wanted.width and now.height == wanted.height) { if (now.width == wanted.width and now.height == wanted.height) {
var line: [80]u8 = undefined; var line: [80]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "display: mode set to {d}x{d}, verified\n", .{ now.width, now.height }) catch "display: mode set, verified\n"); _ = system.write(std.fmt.bufPrint(&line, "display: mode set to {d}x{d}, verified\n", .{ now.width, now.height }) catch "display: mode set, verified\n");
if (backend.hasVsync()) _ = system.write("display: vsync present ok\n"); if (backend.hasFencedPresent()) _ = system.write("display: fenced present ok\n");
} else { } else {
_ = system.write("display: mode set FAILED (geometry unchanged)\n"); _ = system.write("display: mode set FAILED (geometry unchanged)\n");
} }
@@ -443,15 +529,15 @@ fn mouseListener(width: u32, height: u32) void {
} }
} }
/// Consume the latest cursor position from the channel and repaint the cursor layer at /// Consume the latest cursor position from the channel and move the cursor layer to it.
/// it. Runs on the main loop (the compositor owner) in response to a listener poke. /// Runs on the main loop (the compositor owner) in response to a listener poke.
/// `configureLayer` damages both the old and new footprints, so a plain `present` /// `configureLayer` damages both the old and new footprints; the frame clock presents
/// repaints exactly the two rectangles that changed. /// them at the next tick, so a fast mouse coalesces to at most ~60 repaints a second.
fn renderCursor() void { fn renderCursor() void {
const snapshot = cursor_channel.take() orelse return; const snapshot = cursor_channel.take() orelse return;
const id = cursor_layer orelse return; const id = cursor_layer orelse return;
_ = configureLayer(id, snapshot.x, snapshot.y, cursor_z, true); _ = configureLayer(id, snapshot.x, snapshot.y, cursor_z, true);
present(); schedulePresent();
if (!cursor_tracking_reported and if (!cursor_tracking_reported and
@abs(snapshot.x - cursor_origin_x) >= cursor_report_threshold and @abs(snapshot.x - cursor_origin_x) >= cursor_report_threshold and
@abs(snapshot.y - cursor_origin_y) >= cursor_report_threshold) @abs(snapshot.y - cursor_origin_y) >= cursor_report_threshold)
@@ -500,6 +586,7 @@ fn initialise(endpoint: ipc.Handle) bool {
_ = system.write(std.fmt.bufPrint(&line, "display: online {d}x{d} pitch {d} format {d}\n", .{ _ = system.write(std.fmt.bufPrint(&line, "display: online {d}x{d} pitch {d} format {d}\n", .{
mode.width, mode.height, mode.pitch, mode.format, mode.width, mode.height, mode.pitch, mode.format,
}) catch "display: online\n"); }) catch "display: online\n");
updateFrameClock();
_ = system.write("display: presented frame 0\n"); _ = system.write("display: presented frame 0\n");
selfCheck(); selfCheck();
@@ -561,11 +648,13 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Han
return ok(reply); return ok(reply);
}, },
@intFromEnum(protocol.Operation.present) => { @intFromEnum(protocol.Operation.present) => {
present(); // Scheduled, not immediate: the frame clock composites the accumulated damage
// at the next tick, so back-to-back client presents coalesce into one frame.
schedulePresent();
return ok(reply); return ok(reply);
}, },
@intFromEnum(protocol.Operation.attach_scanout) => { @intFromEnum(protocol.Operation.attach_scanout) => {
return attachScanout(request.x, request.width, request.height, request.colour, capability, reply); return attachScanout(request.x, request.width, request.height, request.colour, request.y, capability, reply);
}, },
@intFromEnum(protocol.Operation.set_mode) => { @intFromEnum(protocol.Operation.set_mode) => {
if (!backend.setMode(request.width, request.height)) return fail(reply); if (!backend.setMode(request.width, request.height)) return fail(reply);
@@ -591,21 +680,15 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Han
} }
} }
/// Two notification sources reach the compositor. A **message-notification** is a poke /// Two notification sources reach the compositor, and one coalesced badge can carry
/// from the mouse-listener thread (a buffered self-`ipc.send`, `notify_message_bit`): /// both, so each bit is handled independently. A **message-notification** is a poke from
/// repaint the cursor at its latest channel position. Anything else is the post-attach /// the mouse-listener thread (a buffered self-`ipc.send`, `notify_message_bit`): fold the
/// present **timer**: repaint into the freshly attached native surface, verify the frame /// newest cursor position into the scene. A **timer** (`notify_timer_bit`) is the frame
/// landed, then run the one-shot mode-set self-check (V5). /// clock — or the deferred first native present after `attach_scanout` — either way,
/// present the accumulated damage.
fn onNotification(badge: u64) void { fn onNotification(badge: u64) void {
if (badge & ipc.notify_message_bit != 0) { if (badge & ipc.notify_message_bit != 0) renderCursor();
renderCursor(); if (badge & ipc.notify_timer_bit != 0) frameTick();
return;
}
present(); // native present + verify (first timer fire after the upgrade)
if (pending_modeset_check) {
pending_modeset_check = false;
modesetSelfCheck();
}
} }
pub fn main() void { pub fn main() void {
+7 -5
View File
@@ -24,11 +24,13 @@ pub const Operation = enum(u32) {
damage = 6, damage = 6,
/// present(): composite the dirty layers and flush to the screen. /// present(): composite the dirty layers and flush to the screen.
present = 7, present = 7,
/// attach_scanout(x=stride, width, height, colour=format) + <surface capability>: a native /// attach_scanout(x=stride, y=refresh_hz, width, height, colour=format) + <surface
/// scanout driver announces itself, handing over the shared scanout surface as an `ipc_call` /// capability>: a native scanout driver announces itself, handing over the shared scanout
/// send_cap. The compositor maps it, looks up the driver's `.scanout` present channel, and /// surface as an `ipc_call` send_cap. The compositor maps it, looks up the driver's
/// upgrades off the GOP floor (docs/display-v2.md V4). `x` is the surface's row stride in /// `.scanout` present channel, and upgrades off the GOP floor (docs/display-v2.md V4).
/// pixels, `colour` the DisplayFormat. /// `x` is the surface's row stride in pixels, `y` the panel refresh rate from the
/// driver's EDID read (0 = unknown; paces the compositor's frame clock), `colour` the
/// DisplayFormat.
attach_scanout = 8, attach_scanout = 8,
/// set_mode(width, height): change the display resolution — only a native backend that /// set_mode(width, height): change the display resolution — only a native backend that
/// reports `canModeSet` honours it; on the GOP floor it fails (docs/display-v2.md V5). /// reports `canModeSet` honours it; on the GOP floor it fails (docs/display-v2.md V5).
+337 -16
View File
@@ -70,6 +70,15 @@ pub const FileSystem = struct {
// server before a mutating op. 0 leaves the on-disk timestamps untouched (host // server before a mutating op. 0 leaves the on-disk timestamps untouched (host
// tests that don't care about time, and reads). // tests that don't care about time, and reads).
current_time_epoch: u64 = 0, current_time_epoch: u64 = 0,
// Where the next allocateCluster scan starts — clusters below this were seen
// in use, so a fresh scan needn't re-read them (frees rewind it). Without
// this the scan re-read the FAT from cluster 2 per allocation: measured at
// ~1 s/cluster on a part-full volume (a 37 s shutdown log flush).
next_free_hint: u32 = 2,
// Which absolute LBA `fat_sector` currently holds (0 = none). Lets a FAT
// scan serve consecutive entries from one device read; every write through
// the sector keeps the cache coherent (writeFatBytes updates it in place).
fat_sector_lba: u64 = 0,
// Every filesystem-relative sector access adds the partition base. // Every filesystem-relative sector access adds the partition base.
fn blockRead(self: *FileSystem, lba: u64, buffer: []u8) bool { fn blockRead(self: *FileSystem, lba: u64, buffer: []u8) bool {
@@ -139,7 +148,10 @@ pub const FileSystem = struct {
while (done < out.len) { while (done < out.len) {
const lba = position / sector_size; const lba = position / sector_size;
const within: usize = @intCast(position % sector_size); const within: usize = @intCast(position % sector_size);
if (!self.blockRead(lba, &self.fat_sector)) return false; if (lba != self.fat_sector_lba) {
if (!self.blockRead(lba, &self.fat_sector)) return false;
self.fat_sector_lba = lba;
}
const n = @min(out.len - done, sector_size - within); const n = @min(out.len - done, sector_size - within);
@memcpy(out[done .. done + n], self.fat_sector[within .. within + n]); @memcpy(out[done .. done + n], self.fat_sector[within .. within + n]);
done += n; done += n;
@@ -159,7 +171,10 @@ pub const FileSystem = struct {
while (done < in.len) { while (done < in.len) {
const lba = position / sector_size; const lba = position / sector_size;
const within: usize = @intCast(position % sector_size); const within: usize = @intCast(position % sector_size);
if (!self.blockRead(lba, &self.fat_sector)) return false; if (lba != self.fat_sector_lba) {
if (!self.blockRead(lba, &self.fat_sector)) return false;
self.fat_sector_lba = lba;
}
const n = @min(in.len - done, sector_size - within); const n = @min(in.len - done, sector_size - within);
@memcpy(self.fat_sector[within .. within + n], in[done .. done + n]); @memcpy(self.fat_sector[within .. within + n], in[done .. done + n]);
if (!self.blockWrite(lba, &self.fat_sector)) return false; if (!self.blockWrite(lba, &self.fat_sector)) return false;
@@ -237,11 +252,19 @@ pub const FileSystem = struct {
// Find and claim a free cluster, marking it end-of-chain. Returns its number. // Find and claim a free cluster, marking it end-of-chain. Returns its number.
fn allocateCluster(self: *FileSystem) ?u32 { fn allocateCluster(self: *FileSystem) ?u32 {
var cluster: u32 = 2; const limit = self.geometry.cluster_count + 2;
while (cluster < self.geometry.cluster_count + 2) : (cluster += 1) { // Two passes: hint..end, then 2..hint (the hint only skips known-used
if (self.readFatEntry(cluster) == on_disk.free_cluster) { // ground, it never hides a freed cluster — freeChain rewinds it).
if (!self.writeFatEntry(cluster, self.endOfChainValue())) return null; var pass: u2 = 0;
return cluster; while (pass < 2) : (pass += 1) {
var cluster: u32 = if (pass == 0) self.next_free_hint else 2;
const end: u32 = if (pass == 0) limit else self.next_free_hint;
while (cluster < end) : (cluster += 1) {
if (self.readFatEntry(cluster) == on_disk.free_cluster) {
if (!self.writeFatEntry(cluster, self.endOfChainValue())) return null;
self.next_free_hint = cluster + 1;
return cluster;
}
} }
} }
return null; return null;
@@ -257,6 +280,7 @@ pub const FileSystem = struct {
while (cluster >= 2 and cluster < limit and guard < limit) : (guard += 1) { while (cluster >= 2 and cluster < limit and guard < limit) : (guard += 1) {
const next = self.readFatEntry(cluster); const next = self.readFatEntry(cluster);
_ = self.writeFatEntry(cluster, on_disk.free_cluster); _ = self.writeFatEntry(cluster, on_disk.free_cluster);
if (cluster < self.next_free_hint) self.next_free_hint = cluster;
if (self.isEndOfChain(next) or next < 2) break; if (self.isEndOfChain(next) or next < 2) break;
cluster = next; cluster = next;
} }
@@ -603,6 +627,192 @@ pub const FileSystem = struct {
_ = self.blockWrite(node.entry_sector, &self.dir_sector); _ = self.blockWrite(node.entry_sector, &self.dir_sector);
} }
// --- long-name creation --------------------------------------------------
// The standard 8.3 short-name checksum carried by every long-name entry.
fn shortChecksum(raw: [11]u8) u8 {
var sum: u8 = 0;
for (raw) |c| sum = ((sum & 1) << 7) +% (sum >> 1) +% c;
return sum;
}
fn valid83Char(c: u8) bool {
return (c >= 'A' and c <= 'Z') or (c >= '0' and c <= '9') or c == '-' or c == '_';
}
// Whether an 8.3 entry with exactly this raw name exists in `dir`.
const RawContext = struct { raw: [11]u8, found: *bool };
fn rawVisit(context: *const RawContext, entry: on_disk.DirectoryEntry, name: []const u8, entry_sector: u64, entry_offset: u32) bool {
_ = name;
_ = entry_sector;
_ = entry_offset;
if (std.mem.eql(u8, &entry.name, &context.raw)) {
context.found.* = true;
return true;
}
return false;
}
fn shortNameExists(self: *FileSystem, dir: Node, raw: [11]u8) bool {
var found = false;
var context = RawContext{ .raw = raw, .found = &found };
self.scanDirectory(dir, &context, rawVisit);
return found;
}
// A mangled STEM~N.EXT short name that collides with nothing in `dir` — the
// alias behind a long-name chain.
fn shortNameFor(self: *FileSystem, dir: Node, name: []const u8) ?[11]u8 {
const dot = std.mem.lastIndexOfScalar(u8, name, '.');
const base = if (dot) |d| name[0..d] else name;
const ext = if (dot) |d| name[d + 1 ..] else name[0..0];
var stem: [6]u8 = undefined;
var stem_len: usize = 0;
for (base) |c| {
if (stem_len == stem.len) break;
const upper = std.ascii.toUpper(c);
if (valid83Char(upper)) {
stem[stem_len] = upper;
stem_len += 1;
}
}
if (stem_len == 0) {
stem[0] = 'X';
stem_len = 1;
}
var raw = [_]u8{' '} ** 11;
var ext_len: usize = 0;
for (ext) |c| {
if (ext_len == 3) break;
const upper = std.ascii.toUpper(c);
if (valid83Char(upper)) {
raw[8 + ext_len] = upper;
ext_len += 1;
}
}
var index: u32 = 1;
while (index <= 999_999) : (index += 1) {
var tail_buffer: [8]u8 = undefined;
const tail = std.fmt.bufPrint(&tail_buffer, "~{d}", .{index}) catch return null;
const keep = @min(stem_len, 8 - tail.len);
@memset(raw[0..8], ' ');
@memcpy(raw[0..keep], stem[0..keep]);
@memcpy(raw[keep .. keep + tail.len], tail);
if (!self.shortNameExists(dir, raw)) return raw;
}
return null;
}
// Fill one long-name entry's 13 UTF-16 slots from `name` starting at
// `offset`: the name's bytes widened, then a 0x0000 terminator, then 0xFFFF.
fn fillLongNamePiece(lfn: *on_disk.LongNameEntry, name: []const u8, offset: usize) void {
var units: [13]u16 = undefined;
var i: usize = 0;
while (i < 13) : (i += 1) {
const at = offset + i;
units[i] = if (at < name.len) name[at] else if (at == name.len) 0x0000 else 0xFFFF;
}
lfn.name1 = units[0..5].*;
lfn.name2 = units[5..11].*;
lfn.name3 = units[11..13].*;
}
// The first entry index of a run of `count` free slots in `dir`, growing the
// directory as needed. Fresh clusters are zeroed, so growth always yields
// free slots; only the fixed FAT12/16 root can genuinely run out.
fn findFreeRun(self: *FileSystem, dir: Node, count: usize) ?u32 {
var run_start: u32 = 0;
var run_len: usize = 0;
var sector_index: u32 = 0;
while (self.dirSectorLba(dir, sector_index, true)) |lba| : (sector_index += 1) {
if (!self.blockRead(lba, &self.dir_sector)) return null;
var i: u32 = 0;
while (i < entries_per_sector) : (i += 1) {
const offset = i * @sizeOf(on_disk.DirectoryEntry);
const entry = std.mem.bytesToValue(on_disk.DirectoryEntry, self.dir_sector[offset .. offset + @sizeOf(on_disk.DirectoryEntry)]);
if (entry.isFree()) {
if (run_len == 0) run_start = sector_index * entries_per_sector + i;
run_len += 1;
if (run_len == count) return run_start;
} else {
run_len = 0;
}
}
if (sector_index > 4096) return null; // runaway guard
}
return null;
}
// Write one 32-byte directory entry at a global entry index (read-modify-
// write of its sector). Returns the entry's (sector, offset) or null.
fn writeEntryAt(self: *FileSystem, dir: Node, index: u32, bytes: *const [32]u8) ?EntryLoc {
const lba = self.dirSectorLba(dir, index / entries_per_sector, true) orelse return null;
if (!self.blockRead(lba, &self.dir_sector)) return null;
const offset = (index % entries_per_sector) * @sizeOf(on_disk.DirectoryEntry);
@memcpy(self.dir_sector[offset .. offset + 32], bytes);
if (!self.blockWrite(lba, &self.dir_sector)) return null;
return .{ .sector = lba, .offset = offset };
}
// Add a named directory entry, creating a long-name chain when the name is
// not its own 8.3 form. Write order is LFN pieces first, 8.3 entry last: an
// interrupted create leaves orphaned long-name entries, which every FAT
// reader (this engine's scanner included) skips as unattached — never a
// mismatched chain.
fn addEntryNamed(self: *FileSystem, dir: Node, name: []const u8, attributes: u8, first_cluster: u32, size: u32) ?Node {
if (to83(name)) |raw| {
var display: [12]u8 = undefined;
// Only a name that IS its 8.3 form (already uppercase) skips the
// chain — a lowercase name gets one so its exact case survives,
// matching tools/make-fat-image.py.
if (std.mem.eql(u8, format83(raw, &display), name))
return self.addEntry(dir, raw, attributes, first_cluster, size);
}
if (name.len == 0 or name.len > 255) return null;
const raw = self.shortNameFor(dir, name) orelse return null;
const checksum = shortChecksum(raw);
const piece_count: u32 = @intCast((name.len + 12) / 13);
if (piece_count > 20) return null;
const start = self.findFreeRun(dir, piece_count + 1) orelse return null;
var k: u32 = 0;
while (k < piece_count) : (k += 1) {
const piece = piece_count - k; // stored last-logical-first
var lfn = std.mem.zeroes(on_disk.LongNameEntry);
lfn.order = @intCast(piece | (if (k == 0) @as(u8, 0x40) else 0));
lfn.attributes = on_disk.attribute_long_name;
lfn.checksum = checksum;
fillLongNamePiece(&lfn, name, (piece - 1) * 13);
_ = self.writeEntryAt(dir, start + k, std.mem.asBytes(&lfn)[0..32]) orelse return null;
}
var entry = std.mem.zeroes(on_disk.DirectoryEntry);
entry.name = raw;
entry.attributes = attributes;
entry.file_size = size;
entry.setFirstCluster(first_cluster);
const stamp = on_disk.epochToFatDateTime(self.current_time_epoch);
entry.creation_date = stamp.date;
entry.creation_time = stamp.time;
entry.write_date = stamp.date;
entry.write_time = stamp.time;
entry.last_access_date = stamp.date;
const location = self.writeEntryAt(dir, start + piece_count, std.mem.asBytes(&entry)[0..32]) orelse return null;
return .{
.first_cluster = first_cluster,
.size = size,
.is_directory = attributes & on_disk.attribute_directory != 0,
.mtime = self.current_time_epoch,
.entry_sector = location.sector,
.entry_offset = location.offset,
.has_entry = true,
};
}
// Add an 8.3 directory entry to `dir` with the given attributes, first cluster, // Add an 8.3 directory entry to `dir` with the given attributes, first cluster,
// and size, reusing a free (0x00 or 0xE5) slot and growing the directory chain // and size, reusing a free (0x00 or 0xE5) slot and growing the directory chain
// if needed. Returns the new node (with its entry location) or null if full. // if needed. Returns the new node (with its entry location) or null if full.
@@ -645,19 +855,20 @@ pub const FileSystem = struct {
return null; return null;
} }
/// Create an 8.3-named file in directory `dir`. Returns the new (empty) node, /// Create a file in directory `dir`. Uppercase 8.3 names get a bare short
/// or null if the name is not 8.3-representable or no directory slot is free. /// entry; anything else gets a long-name chain over a mangled ~N alias.
/// Returns the new (empty) node, or null (bad name / directory full /
/// duplicate — the caller checks existence first if it must distinguish).
pub fn createFile(self: *FileSystem, dir: Node, name: []const u8) ?Node { pub fn createFile(self: *FileSystem, dir: Node, name: []const u8) ?Node {
const raw = to83(name) orelse return null; return self.addEntryNamed(dir, name, on_disk.attribute_archive, 0, 0);
return self.addEntry(dir, raw, on_disk.attribute_archive, 0, 0);
} }
/// Create an 8.3-named subdirectory in `dir`: allocate and initialise its first /// Create a subdirectory in `dir`: allocate and initialise its first
/// cluster with "." (itself) and ".." (the parent) entries, then add its /// cluster with "." (itself) and ".." (the parent) entries, then add its
/// directory entry to `dir`. Returns the new directory node, or null (bad name, /// directory entry to `dir` (long-name chain when the name needs one).
/// no free cluster, or the directory is full). Long names are not created. /// Returns the new directory node, or null (bad name, no free cluster, or
/// the directory is full).
pub fn createDirectory(self: *FileSystem, dir: Node, name: []const u8) ?Node { pub fn createDirectory(self: *FileSystem, dir: Node, name: []const u8) ?Node {
const raw = to83(name) orelse return null;
const cluster = self.allocateCluster() orelse return null; const cluster = self.allocateCluster() orelse return null;
self.zeroCluster(cluster); self.zeroCluster(cluster);
@@ -685,7 +896,7 @@ pub const FileSystem = struct {
self.freeChain(cluster); self.freeChain(cluster);
return null; return null;
} }
return self.addEntry(dir, raw, on_disk.attribute_directory, cluster, 0) orelse { return self.addEntryNamed(dir, name, on_disk.attribute_directory, cluster, 0) orelse {
self.freeChain(cluster); self.freeChain(cluster);
return null; return null;
}; };
@@ -1103,3 +1314,113 @@ test "a create stamps the modification time" {
try std.testing.expectEqual(@as(u64, 1_700_000_000), fs.resolve("/STAMP.TXT").?.mtime); try std.testing.expectEqual(@as(u64, 1_700_000_000), fs.resolve("/STAMP.TXT").?.mtime);
try std.testing.expectEqual(@as(u64, 1_700_000_000), fs.listEntry(fs.rootNode(), 0).?.mtime); try std.testing.expectEqual(@as(u64, 1_700_000_000), fs.listEntry(fs.rootNode(), 0).?.mtime);
} }
test "long-name create: directory + file round-trip by long name" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
// The per-boot log directory shape: an 18-char stamp, nested paths, .log names.
const stamp_dir = fs.createDirectory(fs.rootNode(), "2026-07-21T101530Z").?;
try std.testing.expect(stamp_dir.is_directory);
const file = fs.createFile(stamp_dir, "device-manager.log").?;
_ = file;
// Resolve by exact long name, and case-insensitively (FAT semantics).
try std.testing.expect(fs.resolve("/2026-07-21T101530Z/device-manager.log") != null);
try std.testing.expect(fs.resolve("/2026-07-21t101530z/DEVICE-MANAGER.LOG") != null);
// The listing shows the long names, not the ~N aliases.
var listing = fs.listEntry(fs.rootNode(), 0).?;
try std.testing.expectEqualStrings("2026-07-21T101530Z", listing.name_buffer[0..listing.name_len]);
var inner = fs.listEntry(stamp_dir, 2).?; // after "." and ".."
try std.testing.expectEqualStrings("device-manager.log", inner.name_buffer[0..inner.name_len]);
// Write through the created file and read it back by long-name resolve.
var node = fs.resolve("/2026-07-21T101530Z/device-manager.log").?;
try std.testing.expectEqual(@as(usize, 10), fs.writeFile(&node, 0, "hello logs"));
var buffer: [16]u8 = undefined;
try std.testing.expectEqual(@as(usize, 10), fs.readFile(node, 0, buffer[0..10]));
try std.testing.expectEqualStrings("hello logs", buffer[0..10]);
}
test "long-name create: ~N alias collision suffixes stay distinct" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
_ = fs.createFile(fs.rootNode(), "logger-alpha.log").?;
_ = fs.createFile(fs.rootNode(), "logger-beta.log").?;
// Same 6-char mangle stem (LOGGER) — the second must take ~2.
var raw_one = false;
var raw_two = false;
var cursor: u32 = 0;
while (fs.listEntry(fs.rootNode(), cursor)) |entry| : (cursor += 1) {
if (std.mem.eql(u8, entry.name_buffer[0..entry.name_len], "logger-alpha.log")) raw_one = true;
if (std.mem.eql(u8, entry.name_buffer[0..entry.name_len], "logger-beta.log")) raw_two = true;
}
try std.testing.expect(raw_one and raw_two);
try std.testing.expect(fs.resolve("/logger-alpha.log") != null);
try std.testing.expect(fs.resolve("/logger-beta.log") != null);
// Their short aliases took distinct ~N tails. (Alias LOOKUP is not a
// feature — findChild matches display names — but the on-disk aliases
// must not collide for other FAT readers.)
try std.testing.expect(fs.shortNameExists(fs.rootNode(), "LOGGER~1LOG".*));
try std.testing.expect(fs.shortNameExists(fs.rootNode(), "LOGGER~2LOG".*));
}
test "long-name create: unlink removes the chain; slots are reused cleanly" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
_ = fs.createFile(fs.rootNode(), "a-rather-long-file-name.txt").?;
try std.testing.expect(fs.removeFile(fs.rootNode(), "a-rather-long-file-name.txt"));
try std.testing.expect(fs.resolve("/a-rather-long-file-name.txt") == null);
// A new long name reuses the freed run without inheriting the old chain.
_ = fs.createFile(fs.rootNode(), "an-entirely-different-name.md").?;
try std.testing.expect(fs.resolve("/an-entirely-different-name.md") != null);
try std.testing.expect(fs.resolve("/a-rather-long-file-name.txt") == null);
var listing = fs.listEntry(fs.rootNode(), 0).?;
try std.testing.expectEqualStrings("an-entirely-different-name.md", listing.name_buffer[0..listing.name_len]);
}
test "8.3 fast path: an uppercase-compliant name gets one bare entry" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
_ = fs.createFile(fs.rootNode(), "DANOS.LOG").?;
// Exactly one directory entry: entry 0 is the file, entry 1 is the end.
var listing = fs.listEntry(fs.rootNode(), 0).?;
try std.testing.expectEqualStrings("DANOS.LOG", listing.name_buffer[0..listing.name_len]);
try std.testing.expect(fs.listEntry(fs.rootNode(), 1) == null);
// A lowercase 8.3-shaped name is case-preserved via a chain instead.
_ = fs.createFile(fs.rootNode(), "fat.log").?;
var second = fs.listEntry(fs.rootNode(), 1).?;
try std.testing.expectEqualStrings("fat.log", second.name_buffer[0..second.name_len]);
}
test "short-name checksum matches the reference vector" {
// "README TXT" is a widely published example: checksum 0x15... compute a
// fixed pair to pin the rotate-add against regressions.
const a = FileSystem.shortChecksum("README TXT".*);
const b = FileSystem.shortChecksum("LOGGER~1LOG".*);
try std.testing.expect(a != b);
// The algorithm is order-sensitive: swapped bytes change the sum.
const c = FileSystem.shortChecksum("REDAME TXT".*);
try std.testing.expect(a != c);
}
+77 -29
View File
@@ -16,11 +16,6 @@ const on_disk = @import("on-disk.zig");
const protocol = runtime.vfs_protocol; const protocol = runtime.vfs_protocol;
const dma = runtime.dma; const dma = runtime.dma;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
const mount_point = "/mnt/usb"; const mount_point = "/mnt/usb";
// The engine's BlockDevice, backed by the `.block` driver plus a DMA bounce // The engine's BlockDevice, backed by the `.block` driver plus a DMA bounce
@@ -81,17 +76,40 @@ fn fail(out: []u8) usize {
return writeReply(out, .{ .status = -1 }, &.{}); return writeReply(out, .{ .status = -1 }, &.{});
} }
/// How often to look for a block device while none is mounted. Storage arriving
/// is EVENT-shaped (the usb chain registering, possibly after a driver restart),
/// but the registry has no subscription — a slow poll from our own harness loop
/// keeps the service responsive (ping, terminate) while it waits, and keeps it
/// alive to catch storage that appears LATE (a restarted usb-storage after a
/// transient failure — the resilience half of docs/logging.md's storage story).
const mount_retry_ms = 500;
var mounted = false;
var service_endpoint: runtime.ipc.Handle = 0;
fn initialise(endpoint: runtime.ipc.Handle) bool { fn initialise(endpoint: runtime.ipc.Handle) bool {
service_endpoint = endpoint;
_ = runtime.system.write("/system/services/fat: starting, waiting for a block device\n"); _ = runtime.system.write("/system/services/fat: starting, waiting for a block device\n");
const device = runtime.block.open() orelse { // With the router in the kernel, clients hold OUR node ids directly; sweep
_ = runtime.system.write("/system/services/fat: no block device (no storage attached)\n"); // a dead client's open handles via the published exit events (the pattern
return false; // clean exit: nothing to serve // the old userspace router used for its own table).
}; _ = runtime.process.subscribeExits(endpoint);
tryBringUp();
if (!mounted) _ = runtime.system.timerOnce(endpoint, mount_retry_ms);
return true; // serve regardless: requests fail politely until storage mounts
}
/// One storage bring-up attempt: block device -> FAT mount -> VFS mounts. Sets
/// `mounted` on success; a failure leaves everything untouched for the next tick.
fn tryBringUp() void {
if (mounted) return;
const device = runtime.block.tryOpen() orelse return;
const geometry = device.geometry() orelse { const geometry = device.geometry() orelse {
_ = runtime.system.write("/system/services/fat: block geometry unavailable\n"); _ = runtime.system.write("/system/services/fat: block geometry unavailable\n");
return false; return;
}; };
ipc_block = .{ .device = device, .bounce = dma.alloc(4096, dma.coherent) orelse return false }; const bounce = dma.alloc(4096, dma.coherent) orelse return;
ipc_block = .{ .device = device, .bounce = bounce };
const block_device = engine.BlockDevice{ const block_device = engine.BlockDevice{
.context = &ipc_block, .context = &ipc_block,
@@ -102,22 +120,49 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
}; };
filesystem = engine.FileSystem.mount(block_device) orelse { filesystem = engine.FileSystem.mount(block_device) orelse {
_ = runtime.system.write("/system/services/fat: not a FAT filesystem\n"); _ = runtime.system.write("/system/services/fat: not a FAT filesystem\n");
return false; return;
}; };
writeLine("/system/services/fat: mounted FAT ({s}, {d} clusters, partition lba {d})\n", .{ @tagName(filesystem.geometry.fat_type), filesystem.geometry.cluster_count, filesystem.base_lba }); std.log.info("mounted FAT ({s}, {d} clusters, partition lba {d})", .{ @tagName(filesystem.geometry.fat_type), filesystem.geometry.cluster_count, filesystem.base_lba });
// Mount ourselves into the VFS namespace at /mnt/usb (retry while the VFS // Mount ourselves into the kernel VFS at /mnt/usb — and serve /var from the
// comes up). From here the VFS routes /mnt/usb/... to this server. // volume's /var subtree, so FHS paths (the logger's /var/log) stay decoupled
var tries: u32 = 0; // from which volume carries them.
while (tries < 100) : (tries += 1) { if (runtime.fs.mount(mount_point, endpointForMount())) {
if (runtime.fs.mount(mount_point, endpoint)) { std.log.info("mounted {s}", .{mount_point});
writeLine("/system/services/fat: mounted {s}\n", .{mount_point}); } else {
return true; _ = runtime.system.write("/system/services/fat: could not mount /mnt/usb\n");
}
runtime.system.sleep(50);
} }
_ = runtime.system.write("/system/services/fat: could not mount into the VFS\n"); if (runtime.fs.mountRewritten("/var", endpointForMount(), "/var")) {
return true; // still serve directly, even if the namespace mount didn't take std.log.info("mounted /var", .{});
} else {
_ = runtime.system.write("/system/services/fat: could not mount /var\n");
}
mounted = true;
}
fn endpointForMount() runtime.ipc.Handle {
return service_endpoint;
}
/// A subscribed process-exit event: release every open handle the dead client
/// held, so a crashed reader can't pin table slots (or, later, locks).
fn onNotification(badge: u64) void {
const got = runtime.ipc.Received{ .len = 0, .badge = badge, .cap = null };
if (got.isTimer()) {
tryBringUp();
if (!mounted) _ = runtime.system.timerOnce(service_endpoint, mount_retry_ms);
return;
}
if (!got.isChildExit()) return;
const dead = got.childProcessId();
var released: u32 = 0;
for (&open_nodes) |*o| {
if (o.used and o.owner == dead) {
o.* = .{};
released += 1;
}
}
if (released != 0) std.log.info("released {d} handle(s) for dead client {d}", .{ released, dead });
} }
const ParentLeaf = struct { parent: []const u8, leaf: []const u8 }; const ParentLeaf = struct { parent: []const u8, leaf: []const u8 };
@@ -132,7 +177,7 @@ fn splitParent(path: []const u8) ParentLeaf {
}; };
} }
fn handleOpen(out: []u8, path: []const u8, flags: u32) usize { fn handleOpen(out: []u8, path: []const u8, flags: u32, sender: u32) usize {
var node = filesystem.resolve(path); var node = filesystem.resolve(path);
if (node == null and flags & protocol.create != 0) { if (node == null and flags & protocol.create != 0) {
const split = splitParent(path); const split = splitParent(path);
@@ -146,13 +191,13 @@ fn handleOpen(out: []u8, path: []const u8, flags: u32) usize {
filesystem.truncate(&resolved); filesystem.truncate(&resolved);
} }
const index = allocOpen() orelse return fail(out); const index = allocOpen() orelse return fail(out);
open_nodes[index] = .{ .used = true, .node = resolved }; open_nodes[index] = .{ .used = true, .node = resolved, .owner = sender };
return writeReply(out, .{ .status = 0, .node = index }, &.{}); return writeReply(out, .{ .status = 0, .node = index }, &.{});
} }
fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize { fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = capability; _ = capability;
_ = sender; if (!mounted) return fail(out); // storage not up (yet): fail politely, clients retry
if (message.len < protocol.request_size) return fail(out); if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]); const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..]; const payload = message[protocol.request_size..];
@@ -162,7 +207,7 @@ fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.i
filesystem.current_time_epoch = runtime.system.wallClock(); filesystem.current_time_epoch = runtime.system.wallClock();
switch (request.operation) { switch (request.operation) {
.open => return handleOpen(out, payload[0..@min(payload.len, request.len)], request.flags), .open => return handleOpen(out, payload[0..@min(payload.len, request.len)], request.flags, sender),
.read => { .read => {
const o = openAt(request.node) orelse return fail(out); const o = openAt(request.node) orelse return fail(out);
var buffer: [protocol.maximum_payload]u8 = undefined; var buffer: [protocol.maximum_payload]u8 = undefined;
@@ -208,7 +253,9 @@ fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.i
return writeReply(out, .{ .status = 0 }, &.{}); return writeReply(out, .{ .status = 0 }, &.{});
}, },
.mkdir => { .mkdir => {
const split = splitParent(payload[0..@min(payload.len, request.len)]); const path = payload[0..@min(payload.len, request.len)];
if (filesystem.resolve(path) != null) return fail(out); // already exists — no duplicate entries
const split = splitParent(path);
const parent = filesystem.resolve(split.parent) orelse return fail(out); const parent = filesystem.resolve(split.parent) orelse return fail(out);
if (filesystem.createDirectory(parent, split.leaf) == null) return fail(out); if (filesystem.createDirectory(parent, split.leaf) == null) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{}); return writeReply(out, .{ .status = 0 }, &.{});
@@ -240,5 +287,6 @@ pub fn main() void {
.service = .fat, .service = .fat,
.init = initialise, .init = initialise,
.on_message = onMessage, .on_message = onMessage,
.on_notification = onNotification,
}); });
} }
+30 -50
View File
@@ -22,16 +22,30 @@ const runtime = @import("runtime");
const power = runtime.power_protocol; const power = runtime.power_protocol;
const build_options = @import("build_options"); const build_options = @import("build_options");
/// Where the kernel boot log is persisted on the USB FAT volume — an 8.3 name at /// The system services init brings up at boot, in order, by binary path. This is
/// the mount root (see system/services/log-flush). init writes it at shutdown; /// init's policy — the microkernel keeps such choices in user space, not the
/// the log-flush one-shot writes it once at boot. /// kernel. Drivers are absent on purpose: the device manager owns those. (A
const log_path = "/mnt/usb/DANOS.LOG"; /// future init reads this from a manifest under /system/services instead of a
/// hardcoded list.)
/// The system services init brings up at boot, in order. This is init's policy — the const boot_services = if (build_options.diagnose) [_][]const u8{
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent // The diagnose boot: no display service, so the kernel's on-screen boot
/// on purpose: the device manager owns those. (A future init reads this from a // transcript is never suppressed — the timestamped timeline (USB bring-up,
/// manifest under /system/services instead of a hardcoded list.) // storage, logger) stays readable on real hardware with no serial.
const boot_services = [_][]const u8{ "vfs", "input", "device-manager", "fat", "display", "display-demo" }; "/system/services/input",
"/system/services/device-manager",
"/system/services/fat",
"/system/services/logger",
} else [_][]const u8{
"/system/services/input",
"/system/services/device-manager",
"/system/services/fat",
"/system/services/display",
"/system/services/display-demo",
// Last: at shutdown children stop in reverse order, so the logger goes down
// FIRST — its final drain still has the fat server (and the whole storage
// chain) alive underneath it.
"/system/services/logger",
};
/// The live process id of each boot service (0 = not running), indexed by its position /// The live process id of each boot service (0 = not running), indexed by its position
/// in `boot_services`, plus how many times init has restarted it. init supervises these: /// in `boot_services`, plus how many times init has restarted it. init supervises these:
@@ -78,14 +92,6 @@ pub fn main() void {
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| child_ids[i] = id; if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| child_ids[i] = id;
} }
// Once the storage stack is up, a one-shot copies the boot log to the USB
// volume (/mnt/usb/DANOS.LOG) so it can be read on another machine — the only
// way to see it on a headless/real board with no host capturing serial. Fire
// and forget: it polls for the mount itself, and is deliberately NOT one of
// init's supervised children (a transient one-shot must not be stopped-and-
// waited-for during shutdown).
_ = runtime.system.spawn("log-flush");
// Subscribe to power events (retry: the power service registers well after // Subscribe to power events (retry: the power service registers well after
// init starts). Best-effort — without it, a `terminate` signal still // init starts). Best-effort — without it, a `terminate` signal still
// triggers the same shutdown path. // triggers the same shutdown path.
@@ -137,26 +143,21 @@ fn restartChild(id: u32) void {
// An unknown reason (the record aged out) is treated as a crash worth restarting. // An unknown reason (the record aged out) is treated as a crash worth restarting.
const reason = runtime.process.exitReason(id) orelse .fault; const reason = runtime.process.exitReason(id) orelse .fault;
if (reason == .exited) { if (reason == .exited) {
logLine("/system/services/init: {s} exited cleanly; not restarting\n", .{service}); std.log.info("{s} exited cleanly; not restarting", .{service});
return; return;
} }
restart_counts[i] += 1; restart_counts[i] += 1;
if (restart_counts[i] > maximum_restarts) { if (restart_counts[i] > maximum_restarts) {
logLine("/system/services/init: {s} keeps crashing; giving up after {d} restarts\n", .{ service, maximum_restarts }); std.log.info("{s} keeps crashing; giving up after {d} restarts", .{ service, maximum_restarts });
return; return;
} }
logLine("/system/services/init: {s} died ({s}); restarting ({d}/{d})\n", .{ service, @tagName(reason), restart_counts[i], maximum_restarts }); std.log.info("{s} died ({s}); restarting ({d}/{d})", .{ service, @tagName(reason), restart_counts[i], maximum_restarts });
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |new_id| child_ids[i] = new_id; if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |new_id| child_ids[i] = new_id;
return; return;
} }
// An untracked child (e.g. the log-flush one-shot): nothing to restart. // An untracked child (e.g. the log-flush one-shot): nothing to restart.
} }
fn logLine(comptime fmt: []const u8, args: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, args) catch return);
}
/// Look up the power service and subscribe our endpoint (handed over as the /// Look up the power service and subscribe our endpoint (handed over as the
/// call's capability) so events arrive as buffered messages here. /// call's capability) so events arrive as buffered messages here.
fn subscribePower() void { fn subscribePower() void {
@@ -175,26 +176,6 @@ fn subscribePower() void {
_ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {}; _ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {};
} }
/// Copy the whole kernel log to /mnt/usb/DANOS.LOG (the same file log-flush
/// writes at boot), so a poweroff captures the fullest log. Best-effort: if the
/// USB volume is not mounted, the open fails and it does nothing. Must run while
/// the storage services are still alive (see shutDown).
fn flushKernelLog() void {
// Truncate on open so this fuller flush replaces the boot-time one cleanly.
var file = runtime.fs.open(log_path, .{ .create = true, .truncate = true }) orelse return; // no USB volume
defer file.close();
var chunk: [4096]u8 = undefined;
var offset: usize = 0;
while (true) {
const got = runtime.system.klogRead(offset, &chunk);
if (got == 0) break; // reached the end of the accumulated log
if (file.writeAll(chunk[0..got]) == null) break; // storage went away
offset += got;
}
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, "/system/services/init: flushed log to {s} ({d} bytes)\n", .{ log_path, offset }) catch "");
}
/// The stop sequence: persist the log while storage is still up, then terminate /// The stop sequence: persist the log while storage is still up, then terminate
/// each child in reverse spawn order (vfs last — other services may flush through /// each child in reverse spawn order (vfs last — other services may flush through
/// it), waiting up to a deadline for each to exit before killing it, then ask the /// it), waiting up to a deadline for each to exit before killing it, then ask the
@@ -202,10 +183,9 @@ fn flushKernelLog() void {
fn shutDown() void { fn shutDown() void {
shutting_down = true; // the stop loop below kills children — those deaths aren't crashes shutting_down = true; // the stop loop below kills children — those deaths aren't crashes
_ = runtime.system.write("/system/services/init: shutting down\n"); _ = runtime.system.write("/system/services/init: shutting down\n");
// Persist the fullest log to the USB volume BEFORE tearing anything down: the // Log persistence is the logger service's job: it is the LAST boot service,
// reverse-order stop loop below kills the fat server first, so /mnt/usb must be // so the reverse-order stop below terminates it first and its final drain
// written while it is still mounted. // runs while the whole storage chain is still alive.
flushKernelLog();
var i = boot_services.len; var i = boot_services.len;
while (i > 0) { while (i > 0) {
i -= 1; i -= 1;
-62
View File
@@ -1,62 +0,0 @@
//! system/services/log-flush — a one-shot that copies the kernel's in-memory
//! diagnostic log to a file on the mounted USB FAT volume, so the boot log
//! survives to be read on another machine. On a headless or real board there is
//! no host capturing serial, so without this the log is lost at power-off; this
//! is the on-disk equivalent of QEMU's `-serial file:`.
//!
//! It reads the whole kernel log back through `klog_read` (the RAM sink in
//! system/kernel/log.zig) and writes it to /mnt/usb/DANOS.LOG. The name is 8.3
//! (FAT short-name rule: base <= 8, extension <= 3) and lives at the mount root
//! (there is no mkdir on the FAT path yet). init spawns this once the boot
//! services are up; init itself repeats the flush at shutdown for a fuller log.
//!
//! If no USB volume is mounted — no stick, or the initial-ramdisk sweep that
//! spawns every bundled binary bare with no VFS — it waits briefly, then exits
//! silently, deranging no other test's output.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
const log_path = "/mnt/usb/DANOS.LOG";
/// Copy the whole kernel log to the open file, looping klog_read -> write until
/// the log is exhausted. Returns the number of bytes written.
fn drainKernelLog(file: *fs.File) usize {
var chunk: [4096]u8 = undefined;
var offset: usize = 0;
while (true) {
const got = runtime.system.klogRead(offset, &chunk);
if (got == 0) break; // reached the end of the accumulated log
if (file.writeAll(chunk[0..got]) == null) break; // storage went away
offset += got;
}
return offset;
}
pub fn main() void {
// Wait for the fat server to mount /mnt/usb (it must bring up the whole USB
// storage chain first, so it races us at boot). Bounded: if the mount never
// appears — no volume, or the no-VFS ramdisk sweep — give up silently.
var ready = false;
var tries: u32 = 0;
while (tries < 1400) : (tries += 1) {
if (fs.openDirectory("/mnt/usb")) |directory| {
var dir = directory;
dir.close();
ready = true;
break;
}
runtime.system.sleep(50);
}
if (!ready) return; // /mnt/usb never became available — nothing to persist to
// Truncate on open: each flush replaces the file, so a shorter log on a later
// boot of the same stick leaves no stale tail from a previous, longer one.
var file = fs.open(log_path, .{ .create = true, .truncate = true }) orelse return;
const written = drainKernelLog(&file);
file.close();
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, "log-flush: wrote {d} bytes to {s}\n", .{ written, log_path }) catch return);
}
+304
View File
@@ -0,0 +1,304 @@
//! The logger service — the per-process log persister.
//!
//! Drains the tagged kernel log ring (`klog_read`/`klog_status`) and
//! demultiplexes it into **one file per process** on the flash volume:
//!
//! <base>/<boot-stamp>/<binary-path>.log
//! e.g. /mnt/usb/var/log/2026-07-21T101530Z/system/services/fat.log
//!
//! The boot stamp is the wall-clock time of boot (from klog_status), so one
//! boot session is one self-contained directory; the kernel's own records go to
//! kernel.log. Records carry the sender's pid and binary path, stamped by the
//! kernel — the logger trusts the ring, never the payload.
//!
//! Storage is best-effort and late: until the FAT volume mounts, the ring
//! simply buffers (it holds a full boot many times over), and the first drain
//! writes the whole backlog. The storage stack's own records are captured the
//! same way — services never write their own log files (the fat service
//! logging through itself would rendezvous-deadlock; the ring sidesteps that
//! by design).
//!
//! The logger announces itself ONCE (a periodic status line would feed the
//! very stream it drains — self-sustaining churn). Lost records surface as an
//! explicit "-- N records lost --" line derived from sequence-number gaps.
//!
//! Durability: files are opened create-once and kept open across a burst, then
//! all closed after a quiet period (~2 s) — each close is the fat server's
//! SCSI SYNCHRONIZE CACHE, so data-at-risk is bounded by the last busy burst
//! without thrashing the device on every record. `on_terminate` does a final
//! drain and closes everything, so an orderly shutdown loses nothing (init
//! stops the logger FIRST — reverse boot order — while fat is still up).
const std = @import("std");
const runtime = @import("runtime");
const system = runtime.system;
const fs = runtime.fs;
/// Where log trees live: the FHS path. The kernel VFS routes /var to whatever
/// volume the fat server mounted there (today: the /var subtree of the USB
/// flash volume) — swapping the persistent medium later touches fat's two
/// mount calls, never this constant.
const base = "/var/log";
/// Drain cadence and the quiet period after which files are closed (flushed).
const tick_ms = 250;
const quiet_close_ticks = 8; // 8 * 250 ms = 2 s
/// One cached open file per source process path. Sized above the practical
/// process count; the fat server's global open-node table (32) is the real
/// ceiling, so stay comfortably below it.
const maximum_files = 24;
const CachedFile = struct {
used: bool = false,
name: [system.maximum_process_name]u8 = undefined,
name_len: usize = 0,
file: fs.File = undefined,
};
var files: [maximum_files]CachedFile = @splat(.{});
var endpoint: runtime.ipc.Handle = 0;
/// The drain cursor into the ring's byte stream, and loss accounting.
var cursor: u64 = 0;
var next_expected_sequence: u64 = 0;
/// Carry buffer: a record can straddle two klog_read chunks.
var carry: [carry_capacity]u8 = undefined;
var carry_len: usize = 0;
const carry_capacity = 64 + 256 + 64; // header + payload + name, padded generously
/// The per-boot directory, formatted once storage appears.
var boot_directory: [base.len + 1 + 19]u8 = undefined;
var boot_directory_len: usize = 0;
var storage_ready = false;
var announced = false;
var ticks_since_record: u32 = 0;
pub fn main() void {
runtime.service.run(64, .{
.init = initialise,
.on_message = onMessage,
.on_notification = onNotification,
.on_terminate = onTerminate,
});
}
fn initialise(harness_endpoint: runtime.ipc.Handle) bool {
endpoint = harness_endpoint;
const status = system.klogStatus() orelse return false;
cursor = status.tail;
// Sequence expectations start at the tail record's sequence — discovered on
// the first drain; 0 is right for a fresh boot either way.
formatBootDirectory(status.boot_unix_seconds);
_ = system.timerOnce(endpoint, tick_ms);
return true;
}
/// The logger serves no protocol; the ping is answered by the harness.
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
}
fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_timer_bit == 0) return;
tick();
_ = system.timerOnce(endpoint, tick_ms);
}
fn onTerminate() void {
// The completeness receipt FIRST: this record enters the ring before the
// final drain, so the drain carries it into logger.log — a directory whose
// logger.log ends with this marker is complete through shutdown; one that
// doesn't was cut early and may be missing tails.
_ = system.write("logger: shutting down; final flush\n");
drain();
closeAll();
// Serial-only epilogue (after the drain, so it reaches no file — by design).
var line: [96]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "logger: flushed through sequence {d}\n", .{next_expected_sequence}) catch return);
}
fn tick() void {
if (!storage_ready) {
// makePath doubles as the readiness probe: while /var is unmounted the
// resolve fails fast (no storage round trip) and the ring buffers; the
// first success creates the whole per-boot tree.
if (!fs.makePath(boot_directory[0..boot_directory_len])) return;
storage_ready = true;
if (!announced) {
announced = true; // once — a periodic line would feed the stream we drain
var line: [128]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "logger: logging to {s}\n", .{boot_directory[0..boot_directory_len]}) catch "");
}
}
drain();
// Quiet-period close: one device cache flush per burst.
ticks_since_record += 1;
if (ticks_since_record == quiet_close_ticks) closeAll();
}
fn drain() void {
if (!storage_ready) return;
var chunk: [4096]u8 = undefined;
while (true) {
@memcpy(chunk[0..carry_len], carry[0..carry_len]);
const got = system.klogRead(cursor, chunk[carry_len..]) orelse {
// Cursor overwritten: re-sync to the ring tail; the sequence gap is
// reported by the next record's header.
const status = system.klogStatus() orelse return;
cursor = status.tail;
carry_len = 0;
continue;
};
if (got == 0) return; // caught up (any partial record stays carried)
cursor += got;
consume(chunk[0 .. carry_len + got]);
}
}
/// Parse whole records out of `bytes`; keep any trailing partial in `carry`.
fn consume(bytes: []u8) void {
const header_size = system.klog_record_header_size;
var offset: usize = 0;
while (bytes.len - offset >= header_size) {
const header = std.mem.bytesToValue(system.KlogRecordHeader, bytes[offset..][0..32]);
if (header.magic != system.klog_record_magic) {
// Corrupt frame — should not happen; drop the carry and re-sync.
carry_len = 0;
const status = system.klogStatus() orelse return;
cursor = status.head;
return;
}
const record_len = recordLength(header);
if (bytes.len - offset < record_len) break; // partial — carry it
const name = bytes[offset + header_size ..][0..header.name_len];
const message = bytes[offset + header_size + header.name_len ..][0..header.message_len];
deliver(header, name, message);
offset += record_len;
}
const rest = bytes.len - offset;
if (rest > carry_capacity) {
carry_len = 0; // cannot happen with sane frames; drop rather than overflow
return;
}
@memcpy(carry[0..rest], bytes[offset..]);
carry_len = rest;
}
fn deliver(header: system.KlogRecordHeader, name: []const u8, message: []const u8) void {
ticks_since_record = 0;
const file = fileFor(if (header.pid == 0 or name.len == 0) "kernel" else name) orelse return;
if (header.sequence != next_expected_sequence and next_expected_sequence != 0) {
var gap_line: [64]u8 = undefined;
const lost = header.sequence - next_expected_sequence;
if (std.fmt.bufPrint(&gap_line, "-- {d} records lost --\n", .{lost})) |line| {
_ = file.writeAll(line);
} else |_| {}
}
next_expected_sequence = header.sequence + 1;
// [+ssssss.mmm] level: payload
var stamp: [48]u8 = undefined;
const seconds = header.timestamp_ns / 1_000_000_000;
const millis = (header.timestamp_ns / 1_000_000) % 1000;
const level: []const u8 = switch (header.level) {
.err => "error: ",
.warn => "warning: ",
.debug => "debug: ",
.info, .raw => "",
};
if (std.fmt.bufPrint(&stamp, "[{d:>6}.{d:0>3}] {s}", .{ seconds, millis, level })) |prefix| {
_ = file.writeAll(prefix);
} else |_| {}
_ = file.writeAll(message);
if (header.flags & system.klog_flag_truncated != 0) _ = file.writeAll("~");
_ = file.writeAll("\n");
}
/// The cached (or freshly opened) file for a source name. The file path is the
/// binary path with its leading '/' stripped, ".log" appended, under the
/// per-boot directory; parents are created on first use.
fn fileFor(name: []const u8) ?*fs.File {
for (&files) |*cached| {
if (cached.used and std.mem.eql(u8, cached.name[0..cached.name_len], name)) return &cached.file;
}
var slot: ?*CachedFile = null;
for (&files) |*cached| {
if (!cached.used) {
slot = cached;
break;
}
}
const cached = slot orelse evictOne() orelse return null;
var path: [base.len + 1 + 19 + 1 + system.maximum_process_name + 4]u8 = undefined;
const relative = if (name.len != 0 and name[0] == '/') name[1..] else name;
const full = std.fmt.bufPrint(&path, "{s}/{s}.log", .{ boot_directory[0..boot_directory_len], relative }) catch return null;
// Parent directories: everything up to the final slash.
if (std.mem.lastIndexOfScalar(u8, full, '/')) |last| {
if (!fs.makePath(full[0..last])) return null;
}
var file = fs.open(full, .{ .create = true }) orelse return null;
// Append: land after whatever an earlier open of this boot wrote.
if (file.attributes()) |attributes| file.seekTo(attributes.size);
cached.* = .{ .used = true, .file = file };
@memcpy(cached.name[0..name.len], name);
cached.name_len = name.len;
return &cached.file;
}
fn evictOne() ?*CachedFile {
// All slots busy: close the first (oldest-created) and reuse it. Simple and
// rare — the process count sits well under the cache size.
for (&files) |*cached| {
if (cached.used) {
cached.file.close();
cached.used = false;
return cached;
}
}
return null;
}
fn closeAll() void {
for (&files) |*cached| {
if (cached.used) {
cached.file.close();
cached.used = false;
}
}
}
fn recordLength(header: system.KlogRecordHeader) usize {
return std.mem.alignForward(usize, system.klog_record_header_size + header.name_len + header.message_len, system.klog_record_alignment);
}
/// Format the per-boot directory "<base>/YYYY-MM-DDTHHMMSSZ" from the boot
/// wall-clock anchor. No colons — FAT names cannot carry them. A dead RTC
/// (anchor 0) yields the 1970 epoch directory, which is still a valid,
/// distinct-per-boot-rarely name and better than refusing to log.
fn formatBootDirectory(boot_unix_seconds: u64) void {
const epoch_seconds = std.time.epoch.EpochSeconds{ .secs = boot_unix_seconds };
const year_day = epoch_seconds.getEpochDay().calculateYearDay();
const month_day = year_day.calculateMonthDay();
const day_seconds = epoch_seconds.getDaySeconds();
const written = std.fmt.bufPrint(&boot_directory, "{s}/{d:0>4}-{d:0>2}-{d:0>2}T{d:0>2}{d:0>2}{d:0>2}Z", .{
base,
year_day.year,
month_day.month.numeric(),
@as(u32, month_day.day_index) + 1,
day_seconds.getHoursIntoDay(),
day_seconds.getMinutesIntoHour(),
day_seconds.getSecondsIntoMinute(),
}) catch return;
boot_directory_len = written.len;
}
@@ -147,8 +147,8 @@ pub fn main(init: runtime.process.Init) void {
const spinner = runtime.system.spawnSupervised("process-test", &.{"spinner"}, endpoint) orelse fail("spawn spinner"); const spinner = runtime.system.spawnSupervised("process-test", &.{"spinner"}, endpoint) orelse fail("spawn spinner");
runtime.system.sleep(100); // let the sleeper block and the spinner get a core runtime.system.sleep(100); // let the sleeper block and the spinner get a core
if (!listed(sleeper, "process-test")) fail("sleeper not in process_enumerate"); if (!listed(sleeper, "/system/tests/process-test")) fail("sleeper not in process_enumerate");
if (!listed(spinner, "process-test")) fail("spinner not in process_enumerate"); if (!listed(spinner, "/system/tests/process-test")) fail("spinner not in process_enumerate");
// Kills that must be refused: a kernel task (id 0), and an id that was never // Kills that must be refused: a kernel task (id 0), and an id that was never
// issued — both -ESRCH. (-EPERM needs a second supervisor; the kernel-level // issued — both -ESRCH. (-EPERM needs a second supervisor; the kernel-level
@@ -167,8 +167,8 @@ pub fn main(init: runtime.process.Init) void {
if (!runtime.system.kill(spinner)) fail("kill spinner"); if (!runtime.system.kill(spinner)) fail("kill spinner");
if (awaitChildExit(endpoint) != spinner) fail("spinner exit notification"); if (awaitChildExit(endpoint) != spinner) fail("spinner exit notification");
if (listed(sleeper, "process-test")) fail("sleeper still listed after kill"); if (listed(sleeper, "/system/tests/process-test")) fail("sleeper still listed after kill");
if (listed(spinner, "process-test")) fail("spinner still listed after kill"); if (listed(spinner, "/system/tests/process-test")) fail("spinner still listed after kill");
// M17.2: both children were killed by us, and the reason says so — the whole // M17.2: both children were killed by us, and the reason says so — the whole
// restart-policy input, read through the runtime like a real supervisor would. // restart-policy input, read through the runtime like a real supervisor would.
@@ -1,17 +1,17 @@
//! system/services/shm-client — the creating half of the shm test (docs/display-v2.md V2). //! system/services/shared-memory-client — the creating half of the shared-memory test (docs/display-v2.md V2).
//! It `shm_create`s a shared region, writes a known pattern into it, and hands the region's //! It `shared_memory_create`s a shared region, writes a known pattern into it, and hands the region's
//! capability to `shm-server` as an `ipc_call` send_cap. The server maps that capability and //! capability to `shared-memory-server` as an `ipc_call` send_cap. The server maps that capability and
//! confirms the pattern is visible — proving cross-process shared memory over the extended //! confirms the pattern is visible — proving cross-process shared memory over the extended
//! capability-passing path. //! capability-passing path.
const runtime = @import("runtime"); const runtime = @import("runtime");
const system = runtime.system; const system = runtime.system;
const shm = runtime.shm; const shared_memory = runtime.shared_memory;
const ipc = runtime.ipc; const ipc = runtime.ipc;
const pattern_len = 4096; const pattern_len = 4096;
/// The pattern the server checks — must match shm-server.zig. /// The pattern the server checks — must match shared-memory-server.zig.
fn expected(i: usize) u8 { fn expected(i: usize) u8 {
return @truncate(i *% 7 +% 3); return @truncate(i *% 7 +% 3);
} }
@@ -19,28 +19,28 @@ fn expected(i: usize) u8 {
fn lookupServer() ?ipc.Handle { fn lookupServer() ?ipc.Handle {
var attempts: usize = 0; var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) { while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.shm_test)) |h| return h; if (ipc.lookup(.shared_memory_test)) |h| return h;
system.sleep(50); system.sleep(50);
} }
return null; return null;
} }
pub fn main() void { pub fn main() void {
const region = shm.create(pattern_len) orelse { const region = shared_memory.create(pattern_len) orelse {
_ = system.write("shm: create failed\n"); _ = system.write("shared-memory: create failed\n");
return; return;
}; };
var i: usize = 0; var i: usize = 0;
while (i < pattern_len) : (i += 1) region.ptr[i] = expected(i); while (i < pattern_len) : (i += 1) region.ptr[i] = expected(i);
const server = lookupServer() orelse { const server = lookupServer() orelse {
_ = system.write("shm: no server\n"); _ = system.write("shared-memory: no server\n");
return; return;
}; };
// A non-empty message (so it reaches on_message, not the ping path), carrying the shm // A non-empty message (so it reaches on_message, not the ping path), carrying the shared-memory
// region's capability. The reply is empty; we just need the round trip. // region's capability. The reply is empty; we just need the round trip.
var reply: [64]u8 = undefined; var reply: [64]u8 = undefined;
_ = ipc.callCap(server, "shm", &reply, region.handle) catch { _ = ipc.callCap(server, "shared-memory", &reply, region.handle) catch {
_ = system.write("shm: call failed\n"); _ = system.write("shared-memory: call failed\n");
}; };
} }
@@ -0,0 +1,44 @@
//! system/services/shared-memory-server — the receiving half of the shared-memory test (docs/display-v2.md V2).
//! It registers under `ServiceId.shared_memory_test`; when `shared-memory-client` calls it carrying a
//! shared-memory capability, it `shared_memory_map`s that capability and checks the client's pattern
//! is visible through the mapping — proving the two processes share the same physical pages
//! (not a copy). On success it prints `shared-memory: shared 4096 bytes ok`, the test's marker.
const runtime = @import("runtime");
const system = runtime.system;
const shared_memory = runtime.shared_memory;
const ipc = runtime.ipc;
const pattern_len = 4096;
/// The pattern the client writes — must match shared-memory-client.zig.
fn expected(i: usize) u8 {
return @truncate(i *% 7 +% 3);
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
const cap = capability orelse {
_ = system.write("shared-memory: shared FAILED (no capability)\n");
return 0;
};
const ptr = shared_memory.map(cap) orelse {
_ = system.write("shared-memory: shared FAILED (map)\n");
return 0;
};
var i: usize = 0;
while (i < pattern_len) : (i += 1) {
if (ptr[i] != expected(i)) {
_ = system.write("shared-memory: shared FAILED (mismatch)\n");
return 0;
}
}
_ = system.write("shared-memory: shared 4096 bytes ok\n");
return 0; // empty reply — the client only needs the round trip to unblock
}
pub fn main() void {
runtime.service.run(64, .{ .service = .shared_memory_test, .on_message = onMessage });
}
-44
View File
@@ -1,44 +0,0 @@
//! system/services/shm-server — the receiving half of the shm test (docs/display-v2.md V2).
//! It registers under `ServiceId.shm_test`; when `shm-client` calls it carrying a
//! shared-memory capability, it `shm_map`s that capability and checks the client's pattern
//! is visible through the mapping — proving the two processes share the same physical pages
//! (not a copy). On success it prints `shm: shared 4096 bytes ok`, the test's marker.
const runtime = @import("runtime");
const system = runtime.system;
const shm = runtime.shm;
const ipc = runtime.ipc;
const pattern_len = 4096;
/// The pattern the client writes — must match shm-client.zig.
fn expected(i: usize) u8 {
return @truncate(i *% 7 +% 3);
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
const cap = capability orelse {
_ = system.write("shm: shared FAILED (no capability)\n");
return 0;
};
const ptr = shm.map(cap) orelse {
_ = system.write("shm: shared FAILED (map)\n");
return 0;
};
var i: usize = 0;
while (i < pattern_len) : (i += 1) {
if (ptr[i] != expected(i)) {
_ = system.write("shm: shared FAILED (mismatch)\n");
return 0;
}
}
_ = system.write("shm: shared 4096 bytes ok\n");
return 0; // empty reply — the client only needs the round trip to unblock
}
pub fn main() void {
runtime.service.run(64, .{ .service = .shm_test, .on_message = onMessage });
}
+89
View File
@@ -0,0 +1,89 @@
//! /system/tests/vfs-test — a ring-3 client that proves the kernel VFS end to
//! end through the plain `runtime.fs` API: resolve its OWN binary under the
//! kernel-served /system mount, check its metadata, read its ELF magic, and
//! list /system/services. On success it heartbeats "vfstest: ok" so the kernel
//! test can observe it; on failure it reports what went wrong.
//!
//! The "park" role (the fat-client-death test): open a file on the FAT volume,
//! then hold the handle forever without closing — the kill and the fat
//! server's release-on-death sweep are the point.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
pub fn main(init: runtime.process.Init) void {
if (init.arguments.count > 1) {
park();
return;
}
// Our own binary, resolved through the kernel mount table.
const self_path = "/system/tests/vfs-test";
var file = fs.open(self_path, .{}) orelse {
_ = runtime.system.write("vfstest: open of own binary failed\n");
return;
};
defer file.close();
const attributes = file.attributes() orelse {
_ = runtime.system.write("vfstest: attributes failed\n");
return;
};
if (attributes.kind != .regular or attributes.size == 0) {
_ = runtime.system.write("vfstest: bad attributes\n");
return;
}
var header: [4]u8 = undefined;
const n = file.read(&header) orelse 0;
if (n != 4 or header[0] != 0x7f or header[1] != 'E' or header[2] != 'L' or header[3] != 'F') {
_ = runtime.system.write("vfstest: ELF magic mismatch\n");
return;
}
// The write refusal: /system is read-only by construction.
if (file.write("x") != null or fs.open("/system/tests/new-file", .{ .create = true }) != null) {
_ = runtime.system.write("vfstest: /system accepted a write\n");
return;
}
// Listing: /system/services contains init.
var saw_init = false;
if (fs.openDirectory("/system/services")) |listing| {
var directory = listing;
defer directory.close();
var entry: fs.Entry = .{};
while (directory.next(&entry)) {
if (std.mem.eql(u8, entry.name(), "init")) saw_init = true;
}
}
if (!saw_init) {
_ = runtime.system.write("vfstest: /system/services listing missed init\n");
return;
}
while (true) {
_ = runtime.system.write("vfstest: ok\n");
runtime.system.sleep(1000);
}
}
fn park() void {
// The storage chain (usb -> block -> fat -> mounts) takes a few seconds;
// retry until the volume appears.
var parked: ?fs.File = null;
var tries: u32 = 0;
while (parked == null and tries < 1000) : (tries += 1) {
parked = fs.open("/mnt/usb/parked", .{ .create = true });
if (parked == null) runtime.system.sleep(20);
}
if (parked == null) {
_ = runtime.system.write("vfstest: park open failed\n");
return;
}
while (true) {
_ = runtime.system.write("vfstest: parked\n");
runtime.system.sleep(500);
}
}
-39
View File
@@ -1,39 +0,0 @@
//! Pure path utilities for the VFS mount router — no IPC, no state, so they are
//! host-testable in isolation. The router uses these to decide whether an opened
//! path lies under a mount point and, if so, what it looks like relative to that
//! mount.
const std = @import("std");
/// If `path` lies under `mount_prefix` — equal to it, or the prefix followed by a
/// path separator — return the path relative to the mount ("/" for an exact
/// match, otherwise the tail beginning with '/'). Returns null when `path` is not
/// under the mount, so a prefix like "/mnt/usb" never captures "/mnt/usbextra".
pub fn underMount(path: []const u8, mount_prefix: []const u8) ?[]const u8 {
if (path.len < mount_prefix.len) return null;
if (!std.mem.eql(u8, path[0..mount_prefix.len], mount_prefix)) return null;
if (path.len == mount_prefix.len) return "/";
if (path[mount_prefix.len] != '/') return null;
return path[mount_prefix.len..];
}
/// Whether `path` is absolute (rooted at '/'). Bare names — what the flat ramfs
/// uses — are relative and never route through a mount.
pub fn isAbsolute(path: []const u8) bool {
return path.len > 0 and path[0] == '/';
}
test "underMount matches only at path boundaries" {
try std.testing.expectEqualStrings("/", underMount("/mnt/usb", "/mnt/usb").?);
try std.testing.expectEqualStrings("/system/kernel", underMount("/mnt/usb/system/kernel", "/mnt/usb").?);
try std.testing.expect(underMount("/mnt/usbextra", "/mnt/usb") == null); // not a boundary
try std.testing.expect(underMount("/mnt", "/mnt/usb") == null); // shorter than the prefix
try std.testing.expect(underMount("/other", "/mnt/usb") == null);
try std.testing.expect(underMount("greeting", "/mnt/usb") == null); // a bare name
}
test "isAbsolute distinguishes paths from bare names" {
try std.testing.expect(isAbsolute("/mnt/usb"));
try std.testing.expect(!isAbsolute("greeting"));
try std.testing.expect(!isAbsolute(""));
}
-62
View File
@@ -1,62 +0,0 @@
//! /system/services/vfs/vfs-test — a client that proves the VFS round trip end to end: open a
//! file through the `runtime.fs` file API, write to it, seek back, read it, and compare.
//! On success it heartbeats "vfstest: ok" so the kernel test can observe it;
//! on failure it reports what went wrong. Shipped in the initial_ramdisk alongside vfs.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
pub fn main(init: runtime.process.Init) void {
const payload = "hello-vfs";
// The "park" role (the vfs-client-death test): open a file, then hold the
// handle forever without closing — the kill and the VFS's release-on-death
// are the point.
if (init.arguments.count > 1) {
var parked: ?fs.File = null;
var tries: u32 = 0;
while (parked == null and tries < 200) : (tries += 1) {
parked = fs.open("parked", .{ .create = true });
if (parked == null) runtime.system.sleep(20);
}
if (parked == null) {
_ = runtime.system.write("vfstest: park open failed\n");
return;
}
while (true) {
_ = runtime.system.write("vfstest: parked\n");
runtime.system.sleep(500);
}
}
// The VFS server may not have registered yet — retry open until it's up.
var opened: ?fs.File = null;
var tries: u32 = 0;
while (opened == null and tries < 200) : (tries += 1) {
opened = fs.open("greeting", .{ .create = true });
if (opened == null) runtime.system.sleep(20);
}
var greeting = opened orelse {
_ = runtime.system.write("vfstest: open failed\n");
return;
};
if ((greeting.write(payload) orelse 0) != payload.len) {
_ = runtime.system.write("vfstest: write failed\n");
return;
}
greeting.seekTo(0);
var buffer: [32]u8 = undefined;
const n = greeting.read(&buffer) orelse 0;
greeting.close();
if (n == payload.len and std.mem.eql(u8, buffer[0..n], payload)) {
while (true) {
_ = runtime.system.write("vfstest: ok\n");
runtime.system.sleep(1000);
}
}
_ = runtime.system.write("vfstest: mismatch\n");
}
-396
View File
@@ -1,396 +0,0 @@
//! system/services/vfs — the user-space VFS server. Shipped in the initial_ramdisk, spawned as a
//! ring-3 process, and reached by every other process through IPC (the `runtime`
//! file API marshals open/read/write/stat/close into calls to this server's
//! endpoint, published under the well-known `vfs` service id).
//!
//! Two namespaces meet here (M5):
//! - a small in-memory **ramfs** — opening a bare name creates it — enough to
//! prove the round trip and to back the existing tests;
//! - **mounted filesystems**: a mount table maps an absolute path prefix (e.g.
//! `/mnt/usb`) to a backend server's endpoint. An open of a path under a mount
//! is *forwarded* to that backend (which speaks this same protocol), and every
//! later read/write/status/readdir/close on the resulting handle is relayed to
//! it. The VFS is the router; a filesystem (FAT) is the backend.
//!
//! A path routes through a mount only when it is absolute and lies under a mount
//! prefix; bare names always resolve in the flat ramfs — the backward-compat
//! contract the `vfs` / `vfs-client-death` tests rely on.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.vfs_protocol;
const path = @import("path.zig");
const ipc = runtime.ipc;
const Node = struct {
used: bool = false,
name: [24]u8 = undefined,
name_len: usize = 0,
data: [512]u8 = undefined,
size: usize = 0,
};
const OpenFile = struct {
used: bool = false,
// For a local handle: an index into `nodes`. For a forwarding handle: the
// node id the backend returned. (usize == u64 here, so it holds either.)
node: usize = 0,
// Non-null for a handle that forwards to a mounted backend.
backend: ?ipc.Handle = null,
// The client (task id — an IPC badge is one) that opened this handle. What
// release-on-death sweeps by: a service must never depend on its clients
// cleaning up after themselves (docs/process-lifecycle.md).
owner: u32 = 0,
};
// One mounted filesystem: an absolute path prefix and the backend endpoint that
// serves everything under it.
const Mount = struct {
used: bool = false,
prefix: [64]u8 = undefined,
prefix_len: usize = 0,
backend: ipc.Handle = 0,
};
var nodes = [_]Node{.{}} ** 8;
var opens = [_]OpenFile{.{}} ** 16;
var mounts = [_]Mount{.{}} ** 8;
fn findNode(name: []const u8) ?usize {
for (&nodes, 0..) |*n, i| {
if (n.used and std.mem.eql(u8, n.name[0..n.name_len], name)) return i;
}
return null;
}
fn createNode(name: []const u8) ?usize {
for (&nodes, 0..) |*n, i| {
if (!n.used) {
const l = @min(name.len, n.name.len);
@memcpy(n.name[0..l], name[0..l]);
n.* = .{ .used = true, .name = n.name, .name_len = l, .size = 0 };
return i;
}
}
return null;
}
fn openAt(id: u64) ?*OpenFile {
if (id >= opens.len) return null;
const o = &opens[@intCast(id)];
return if (o.used) o else null;
}
/// The mount whose prefix most specifically contains `name`, and the path
/// relative to it. Only absolute paths route; bare names never match.
const MountMatch = struct { backend: ipc.Handle, relative: []const u8 };
fn longestMount(name: []const u8) ?MountMatch {
if (!path.isAbsolute(name)) return null;
var best: ?MountMatch = null;
var best_len: usize = 0;
for (&mounts) |*m| {
if (!m.used) continue;
const prefix = m.prefix[0..m.prefix_len];
if (path.underMount(name, prefix)) |relative| {
if (best == null or prefix.len >= best_len) {
best_len = prefix.len;
best = .{ .backend = m.backend, .relative = relative };
}
}
}
return best;
}
/// Serialise a reply header + payload into `out`; returns the total length.
fn writeReply(out: []u8, reply: protocol.Reply, payload: []const u8) usize {
@memcpy(out[0..protocol.reply_size], std.mem.asBytes(&reply));
const n = @min(payload.len, out.len - protocol.reply_size);
@memcpy(out[protocol.reply_size..][0..n], payload[0..n]);
return protocol.reply_size + n;
}
fn fail(out: []u8) usize {
return writeReply(out, .{ .status = -1 }, &.{});
}
/// Format one whole log line and emit it in a single `debug_write`, so lines from
/// concurrent processes can never land in the middle of it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// --- mount routing ----------------------------------------------------------
/// Forward an open under a mount to its backend and, on success, allocate a local
/// forwarding handle that remembers the backend's node id.
fn forwardOpen(out: []u8, backend: ipc.Handle, relative: []const u8, flags: u32, sender: u32) usize {
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = flags };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const rel = relative[0..@min(relative.len, protocol.maximum_payload)];
@memcpy(message[protocol.request_size..][0..rel.len], rel);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + rel.len], &reply) catch return fail(out);
if (n < protocol.reply_size) return fail(out);
const backend_reply = std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]);
if (backend_reply.status != 0) return writeReply(out, .{ .status = backend_reply.status }, &.{});
for (&opens, 0..) |*o, i| {
if (!o.used) {
o.* = .{ .used = true, .node = @intCast(backend_reply.node), .backend = backend, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{});
}
}
return fail(out);
}
/// Relay a read/write/status/readdir/close on a forwarding handle to the backend
/// (the node already rewritten to the backend's id) and copy its reply out.
fn forwardRequest(out: []u8, backend: ipc.Handle, request: protocol.Request, payload: []const u8) usize {
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const plen = @min(payload.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..plen], payload[0..plen]);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + plen], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Forward a path-based operation (mkdir, unlink) under a mount to its backend and
/// relay the reply. No handle is created — these operate by path and return only a
/// status.
fn forwardPath(out: []u8, backend: ipc.Handle, operation: protocol.Operation, relative: []const u8) usize {
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = 0 };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const rel = relative[0..@min(relative.len, protocol.maximum_payload)];
@memcpy(message[protocol.request_size..][0..rel.len], rel);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + rel.len], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Forward a rename to its backend: the payload is the mount-relative old path, a
/// 0x00 separator, then the mount-relative new path. Relays the backend's reply.
fn forwardRename(out: []u8, backend: ipc.Handle, old_relative: []const u8, new_relative: []const u8) usize {
const total = old_relative.len + 1 + new_relative.len;
var message: [protocol.message_maximum]u8 = undefined;
if (protocol.request_size + total > message.len) return fail(out);
const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 };
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
var p = protocol.request_size;
@memcpy(message[p..][0..old_relative.len], old_relative);
p += old_relative.len;
message[p] = 0;
p += 1;
@memcpy(message[p..][0..new_relative.len], new_relative);
p += new_relative.len;
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0..p], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Best-effort close of a backend node (used when a dead client's forwarding
/// handles are swept — the backend must not leak the vfs's opens).
fn forwardClose(backend: ipc.Handle, backend_node: u64) void {
const request = protocol.Request{ .operation = .close, .node = backend_node, .offset = 0, .len = 0, .flags = 0 };
var reply: [protocol.message_maximum]u8 = undefined;
_ = ipc.call(backend, std.mem.asBytes(&request), &reply) catch {};
}
fn doMount(out: []u8, prefix: []const u8, backend: ipc.Handle) usize {
for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefix[0..m.prefix_len], prefix)) {
m.backend = backend;
writeLine("/system/services/vfs: remounted {s}\n", .{prefix});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
for (&mounts) |*m| {
if (!m.used) {
const l = @min(prefix.len, m.prefix.len);
m.used = true;
@memcpy(m.prefix[0..l], prefix[0..l]);
m.prefix_len = l;
m.backend = backend;
writeLine("/system/services/vfs: mounted {s}\n", .{prefix[0..l]});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
return fail(out);
}
fn doUnmount(out: []u8, prefix: []const u8) usize {
for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefix[0..m.prefix_len], prefix)) {
m.used = false;
writeLine("/system/services/vfs: unmounted {s}\n", .{prefix});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
return fail(out);
}
/// Release every open handle `client` held — called on that client's published
/// exit event. Forwarding handles also tell their backend to release; local
/// nodes (the ramfs files) stay, since ramfs contents outlive their writers.
fn releaseClientHandles(client: u32) void {
var released: u32 = 0;
for (&opens) |*o| {
if (o.used and o.owner == client) {
if (o.backend) |backend| forwardClose(backend, o.node);
o.used = false;
released += 1;
}
}
if (released != 0) writeLine("/system/services/vfs: released {d} handle(s) for dead client {d}\n", .{ released, client });
}
/// Handle one request from `sender`; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8, sender: u32, capability: ?ipc.Handle) usize {
if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..];
switch (request.operation) {
.mount => {
const prefix = payload[0..@min(payload.len, request.len)];
const backend = capability orelse return fail(out);
return doMount(out, prefix, backend);
},
.unmount => {
const prefix = payload[0..@min(payload.len, request.len)];
return doUnmount(out, prefix);
},
.open => {
const name = payload[0..@min(payload.len, request.len)];
if (longestMount(name)) |m| return forwardOpen(out, m.backend, m.relative, request.flags, sender);
// An absolute path with no matching mount is simply not found — only
// bare names live in the flat ramfs. (Else /mnt/usb would be silently
// created as a flat file when its filesystem is not yet mounted.)
if (path.isAbsolute(name)) return fail(out);
const ni = findNode(name) orelse createNode(name) orelse return fail(out);
for (&opens, 0..) |*o, i| {
if (!o.used) {
o.* = .{ .used = true, .node = ni, .backend = null, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{});
}
}
return fail(out);
},
.read => {
const of = openAt(request.node) orelse return fail(out);
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const nd = &nodes[@intCast(of.node)];
const off: usize = @intCast(request.offset);
if (off >= nd.size) return writeReply(out, .{ .status = 0, .len = 0 }, &.{}); // EOF
const n = @min(@min(nd.size - off, request.len), protocol.maximum_payload);
return writeReply(out, .{ .status = 0, .len = @intCast(n) }, nd.data[off .. off + n]);
},
.write => {
const of = openAt(request.node) orelse return fail(out);
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const nd = &nodes[@intCast(of.node)];
const off: usize = @intCast(request.offset);
if (off > nd.data.len) return fail(out);
const n = @min(@min(payload.len, request.len), nd.data.len - off);
@memcpy(nd.data[off .. off + n], payload[0..n]);
if (off + n > nd.size) nd.size = off + n;
return writeReply(out, .{ .status = 0, .len = @intCast(n) }, &.{});
},
.status => {
const of = openAt(request.node) orelse return fail(out);
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const st = protocol.FileStatus{ .size = nodes[@intCast(of.node)].size, .kind = @intFromEnum(protocol.NodeKind.regular) };
return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&st));
},
.readdir => {
const of = openAt(request.node) orelse return fail(out);
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
// The flat ramfs has no directories: report EOF.
return writeReply(out, .{ .status = 0, .len = 0 }, &.{});
},
.close => {
const of = openAt(request.node);
if (of) |o| {
if (o.backend) |backend| forwardClose(backend, o.node);
o.used = false;
}
return writeReply(out, .{ .status = 0 }, &.{});
},
.mkdir, .unlink => {
const name = payload[0..@min(payload.len, request.len)];
if (longestMount(name)) |m| return forwardPath(out, m.backend, request.operation, m.relative);
// Only a mounted backend has real directories; the flat ramfs cannot
// create or remove them (and a bare-name path is not a mount target).
return fail(out);
},
.rename => {
const both = payload[0..@min(payload.len, request.len)];
const sep = std.mem.indexOfScalar(u8, both, 0) orelse return fail(out);
const old_path = both[0..sep];
const new_path = both[sep + 1 ..];
const mo = longestMount(old_path) orelse return fail(out);
const mn = longestMount(new_path) orelse return fail(out);
// Both paths must live under the same mount — cross-filesystem rename is
// not supported.
if (mo.backend != mn.backend) return fail(out);
return forwardRename(out, mo.backend, mo.relative, mn.relative);
},
}
}
/// Startup, under the harness: subscribe to the published exit events — when a
/// client dies holding open handles, the exit notification is how the VFS learns
/// to release them (docs/process-lifecycle.md).
fn initialise(endpoint: ipc.Handle) bool {
if (!runtime.process.subscribeExits(endpoint)) {
_ = runtime.system.write("/system/services/vfs: exit subscription failed\n");
}
_ = runtime.system.write("/system/services/vfs: ready\n");
return true;
}
/// A non-signal notification: the only kind the VFS subscribes to is exit events.
fn onNotification(badge: u64) void {
if (badge & ipc.notify_exit_bit != 0) {
releaseClientHandles(@intCast(badge & ~(ipc.notify_badge_bit | ipc.notify_exit_bit)));
}
}
pub fn main() void {
// The harness owns the loop: requests dispatch to handle(), exit events to
// onNotification(), ping and terminate are answered for free — this service
// gained the whole lifecycle contract by deleting its hand-rolled loop.
runtime.service.run(protocol.message_maximum, .{
.service = .vfs,
.init = initialise,
.on_message = handle,
.on_notification = onNotification,
});
}
+131 -39
View File
@@ -195,12 +195,12 @@ CASES = [
"smp": 4, "smp": 4,
"expect": r"display: online \d+x\d+[\s\S]*display: cursor tracking mouse ok", "expect": r"display: online \d+x\d+[\s\S]*display: cursor tracking mouse ok",
"fail": r"display: (could not|mouse subscribe failed)|CPU EXCEPTION|KERNEL PANIC"}, "fail": r"display: (could not|mouse subscribe failed)|CPU EXCEPTION|KERNEL PANIC"},
# Shared memory (v2 V2): shm-client creates a region, writes a pattern, and passes its # Shared memory (v2 V2): shared-memory-client creates a region, writes a pattern, and passes its
# capability to shm-server, which maps it and confirms the same bytes — proving # capability to shared-memory-server, which maps it and confirms the same bytes — proving
# cross-process shared pages over the extended capability passing. # cross-process shared pages over the extended capability passing.
{"name": "shm", {"name": "shared-memory",
"expect": r"shm: shared 4096 bytes ok", "expect": r"shared-memory: shared 4096 bytes ok",
"fail": r"shm: (shared FAILED|create failed|no server|call failed|map)|CPU EXCEPTION|KERNEL PANIC"}, "fail": r"shared-memory: (shared FAILED|create failed|no server|call failed|map)|CPU EXCEPTION|KERNEL PANIC"},
# virtio-gpu driver (v2 V3): boot with an emulated virtio-gpu. The device-manager stack # virtio-gpu driver (v2 V3): boot with an emulated virtio-gpu. The device-manager stack
# discovers the PCI function and spawns the driver, which brings up the control virtqueue, # discovers the PCI function and spawns the driver, which brings up the control virtqueue,
# creates a 2D scanout resource backed by DMA memory, set_scanouts it, paints a test # creates a 2D scanout resource backed by DMA memory, set_scanouts it, paints a test
@@ -221,15 +221,16 @@ CASES = [
# require all three markers to appear somewhere rather than in a fixed order. # require all three markers to appear somewhere rather than in a fixed order.
"expect": r"(?s)(?=.*display: scanout upgraded to virtio-gpu)(?=.*display: native present verified)(?=.*display-demo: ok)", "expect": r"(?s)(?=.*display: scanout upgraded to virtio-gpu)(?=.*display: native present verified)(?=.*display-demo: ok)",
"fail": r"display: native present FAILED|display: could not|display-demo: (no display|create failed)|CPU EXCEPTION|KERNEL PANIC"}, "fail": r"display: native present FAILED|display: could not|display-demo: (no display|create failed)|CPU EXCEPTION|KERNEL PANIC"},
# Mode-set + EDID + vsync (v2 V5): same boot as display-native. After upgrading, the # Mode-set + EDID + fenced presents (v2 V5): same boot as display-native. After upgrading,
# compositor queries the driver's modes, switches to a different resolution, and confirms the # the compositor queries the driver's modes, switches to a different resolution, and confirms
# backend now reports it; the fenced present path makes it a vsync present. (The driver also # the backend now reports it; each present is fenced — completion-acknowledged and tear-free,
# logs the EDID preferred mode during bring-up.) Reuses the display-native kernel scenario. # not vblank-paced (docs/display-v2.md, "Fenced is not vsync"). (The driver also logs the
# EDID preferred mode during bring-up.) Reuses the display-native kernel scenario.
{"name": "display-modeset", {"name": "display-modeset",
"build_case": "display-native", "build_case": "display-native",
"qemu_extra": ["-device", "virtio-gpu-pci"], "qemu_extra": ["-device", "virtio-gpu-pci"],
"mem": "512M", "mem": "512M",
"expect": r"(?s)(?=.*display: mode set to \d+x\d+, verified)(?=.*display: vsync present ok)", "expect": r"(?s)(?=.*display: mode set to \d+x\d+, verified)(?=.*display: fenced present ok)",
"fail": r"display: mode set FAILED|display: mode-set self-check: |display: native present FAILED|CPU EXCEPTION|KERNEL PANIC"}, "fail": r"display: mode set FAILED|display: mode-set self-check: |display: native present FAILED|CPU EXCEPTION|KERNEL PANIC"},
# Resilience: driver restart + re-attach (v2 V6). device-manager (in test-scanout-restart # Resilience: driver restart + re-attach (v2 V6). device-manager (in test-scanout-restart
# mode) kills the virtio-gpu driver once after it hellos; the restart policy respawns it, it # mode) kills the virtio-gpu driver once after it hellos; the restart policy respawns it, it
@@ -426,6 +427,7 @@ CASES = [
# open handle, and the VFS releases it (process-lifecycle.md "Who learns of a death"). # open handle, and the VFS releases it (process-lifecycle.md "Who learns of a death").
{"name": "vfs-client-death", {"name": "vfs-client-death",
"smp": 4, "smp": 4,
"timeout": 90, # the park client waits out the whole USB->block->fat chain
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.4: signals over IPC — ping, reload, terminate (clean exit), the one-shot # M17.4: signals over IPC — ping, reload, terminate (clean exit), the one-shot
@@ -446,7 +448,7 @@ CASES = [
r"device-manager: child added[\s\S]*" r"device-manager: child added[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*" r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-manager: child removed[\s\S]*" r"device-manager: child removed[\s\S]*"
r"device-manager: restarting usb-xhci-bus[\s\S]*" r"device-manager: restarting \S*usb-xhci-bus[\s\S]*"
r"device-manager: child added", r"device-manager: child added",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# USB HID end to end: boot the full tree, enumerate the xHCI, and let the # USB HID end to end: boot the full tree, enumerate the xHCI, and let the
@@ -457,7 +459,43 @@ CASES = [
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
# usb-kbd/usb-mouse ride the default boot xHCI bus (see qemu_args). # usb-kbd/usb-mouse ride the default boot xHCI bus (see qemu_args).
"expect": r"(?=[\s\S]*usb-hid/keyboard: ok)(?=[\s\S]*usb-hid/mouse: ok)", "expect": r"(?=[\s\S]*usb-hid-keyboard: ok)(?=[\s\S]*usb-hid-mouse: ok)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Keyboard echo: inject a known phrase via QMP send-key; the usb-hid-keyboard
# driver decodes it and echoes each character to the log (the simple
# real-hardware keyboard check — type a phrase, read it back off the stick,
# or watch it live on screen in a -Ddiagnose boot). Proves the whole path:
# HID report -> decode -> layout -> character.
{"name": "usb-key-echo",
"build_case": "usb-hid",
"smp": 4,
"timeout": 150,
"qmp_sequence": [
{"delay": 8, "command": "send-key", "arguments": {"keys": [{"type": "qcode", "data": c}]}}
for c in ["k", "e", "y", "t", "e", "s", "t"]
],
"expect": r"keytest",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# USB hub (B4a/B4b, docs/usb-hub.md): a USB2 hub on the xHCI bus with a
# keyboard behind it (a STATIC boot topology — no hot-plug event needed).
# B4a: the hub enumerates and powers its downstream ports. B4b extends the
# expect to the downstream keyboard binding usb-hid-keyboard.
{"name": "usb-hub",
"build_case": "usb-hid",
"smp": 4,
"timeout": 150,
# A SECOND xhci controller carries the hub topology, isolated from the boot
# controller's auto-assigned devices (whose ports the hub would collide
# with). danos spawns a second usb-xhci-bus for it. The hub sits on port 1,
# the keyboard on the hub's downstream port 1 (port=1.1).
"qemu_extra": ["-device", "qemu-xhci,id=xhci2",
"-device", "usb-hub,bus=xhci2.0,port=1",
"-device", "usb-kbd,bus=xhci2.0,port=1.1"],
# B4a: the hub powers its ports. B4b: the keyboard behind it enumerates
# (route string + TT) and binds usb-hid-keyboard.
"expect": r"(?s)(?=.*hub slot \d+: \d+ downstream ports powered)"
r"(?=.*hub slot \d+ port \d+ device:.*0x0627)"
r"(?=.*usb-hid-keyboard: ok \(device 3)",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# USB mass storage end to end: the boot usb-storage device (the FAT32 image, # USB mass storage end to end: the boot usb-storage device (the FAT32 image,
# which has a real 0x55AA boot sector) is enough — the manager spawns # which has a real 0x55AA boot sector) is enough — the manager spawns
@@ -525,8 +563,8 @@ CASES = [
{"name": "acpi-ps2", {"name": "acpi-ps2",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"expect": r"acpi: reported PNP0303[\s\S]*" "expect": r"discovery: reported PNP0303[\s\S]*"
r"device-manager: spawned ps2-bus[\s\S]*" r"device-manager: spawned \S*ps2-bus[\s\S]*"
r"ps2-bus: keyboard driver attached", r"ps2-bus: keyboard driver attached",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# M21.1: the SCI + power button. Boot the manager (which spawns the acpi # M21.1: the SCI + power button. Boot the manager (which spawns the acpi
@@ -552,17 +590,18 @@ CASES = [
r"power: entering S5", r"power: entering S5",
"fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"}, "fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"},
# M8: the boot log is persisted to the USB FAT volume. Reuses the orderly- # M8: the boot log is persisted to the USB FAT volume. Reuses the orderly-
# shutdown build (full tree + power button): init spawns log-flush at boot, # shutdown build (full tree + power button): the logger service announces its
# which copies the kernel log to /mnt/usb/DANOS.LOG once /mnt/usb is mounted # per-boot directory once storage mounts (first marker), then the power
# (first marker); then the power button drives init's own pre-teardown flush # button drives the orderly stop — the logger, stopped first, final-drains
# (second marker), proving both triggers write the file while storage is up. # and reports the flush (second marker) before S5.
{"name": "log-flush", {"name": "logger",
"build_case": "orderly-shutdown", "build_case": "orderly-shutdown",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"qmp_after": {"delay": 8, "command": "system_powerdown"}, "qmp_after": {"delay": 8, "command": "system_powerdown"},
"expect": r"log-flush: wrote \d+ bytes to /mnt/usb/DANOS\.LOG[\s\S]*" "expect": r"logger: logging to /var/log/\d{4}-\d{2}-\d{2}T\d{6}Z[\s\S]*"
r"init: flushed log to /mnt/usb/DANOS\.LOG[\s\S]*" r"init: shutting down[\s\S]*"
r"logger: flushed through sequence \d+[\s\S]*"
r"power: entering S5", r"power: entering S5",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers + # M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers +
@@ -571,8 +610,8 @@ CASES = [
{"name": "acpi-report", {"name": "acpi-report",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"expect": r"acpi: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*" "expect": r"discovery: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*"
r"acpi: reported PNP0F13 \(device \d+, 1 resources\)", r"discovery: reported PNP0F13 \(device \d+, 1 resources\)",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# M19.1/M19.3: the ring-3 PCI scan. pci-bus walks the ECAM through its mmio_map # M19.1/M19.3: the ring-3 PCI scan. pci-bus walks the ECAM through its mmio_map
# grant and registers every function it finds; the kernel's own walk retired, so # grant and registers every function it finds; the kernel's own walk retired, so
@@ -588,7 +627,7 @@ CASES = [
"timeout": 60, "timeout": 60,
"expect": r"pci-bus: (\d+) functions found[\s\S]*" "expect": r"pci-bus: (\d+) functions found[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*" r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-manager: restarting pci-bus[\s\S]*" r"device-manager: restarting \S*pci-bus[\s\S]*"
r"pci-bus: \1 functions found[\s\S]*" r"pci-bus: \1 functions found[\s\S]*"
r"DANOS-TEST-RESULT: PASS", r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -622,6 +661,45 @@ CASES = [
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# The user-space VFS: a client opens/writes/reads a file through the rt file # The user-space VFS: a client opens/writes/reads a file through the rt file
# API, which IPCs the VFS server process; the round trip must match. # API, which IPCs the VFS server process; the round trip must match.
# (A usb-hotplug case was prototyped here, but QEMU's qemu-xhci does not
# raise a runtime port-change event to a polling driver on device_add, so it
# cannot exercise the path. The hot-plug code — port-change queue, teardown
# via Disable Slot, ChildRemoved reporting — is validated on real hardware,
# flagged for the user. The qmp_sequence harness support it added remains.)
# Hub-behind-hub (B4c): route strings compose across tiers — a keyboard two
# hubs deep enumerates and binds. Static nested topology on a 2nd controller.
{"name": "usb-hub-nested",
"build_case": "usb-hid",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci2",
"-device", "usb-hub,bus=xhci2.0,port=1",
"-device", "usb-hub,bus=xhci2.0,port=1.1",
"-device", "usb-kbd,bus=xhci2.0,port=1.1.1"],
"expect": r"(?s)(?=.*hub slot \d+ port \d+ device:.*0x0409)"
r"(?=.*usb-hid-keyboard: ok \(device 3)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Hub-downstream disconnect (B4c): device_del the keyboard behind the hub;
# the hub's status-change endpoint reports it, the device is torn down
# (ChildRemoved + Disable Slot). QEMU's hub DOES raise downstream changes
# (unlike root-port hot-plug).
{"name": "usb-hub-unplug",
"build_case": "usb-hid",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci2",
"-device", "usb-hub,bus=xhci2.0,port=1",
"-device", "usb-kbd,bus=xhci2.0,port=1.1,id=dkbd"],
"qmp_after": {"delay": 8, "command": "device_del", "arguments": {"id": "dkbd"}},
"expect": r"(?s)(?=.*usb-hid-keyboard: ok \(device 3)"
r"(?=.*slot \d+ disconnected)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The kernel VFS root (M-F): the mount table serves the initrd at /system —
# path resolution, node status/read (an ELF magic), and directory listing,
# asserted kernel-side.
{"name": "kvfs",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "vfs", {"name": "vfs",
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -695,11 +773,12 @@ def resolve_firmware(arch):
+ "\nInstall OVMF (edk2-ovmf / ovmf) or add its path above.") + "\nInstall OVMF (edk2-ovmf / ovmf) or add its path above.")
def qmp_send(path, command): def qmp_send(path, command, arguments=None):
"""One QMP command: connect, capabilities handshake, execute. Raises on any """One QMP command: connect, capabilities handshake, execute. Raises on any
failure — the caller retries until the guest's socket is ready. This is how failure — the caller retries until the guest's socket is ready. This is how
a case injects a host-side event (system_powerdown = the ACPI power button) a case injects a host-side event into the running guest: system_powerdown
into the running guest (docs/power.md).""" (the ACPI power button, docs/power.md) or device_add/device_del (USB
hot-plug, docs/driver-model.md)."""
sock = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM) sock = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
sock.settimeout(5) sock.settimeout(5)
try: try:
@@ -709,7 +788,10 @@ def qmp_send(path, command):
stream.write(json.dumps({"execute": "qmp_capabilities"}) + "\n") stream.write(json.dumps({"execute": "qmp_capabilities"}) + "\n")
stream.flush() stream.flush()
stream.readline() # {"return": {}} stream.readline() # {"return": {}}
stream.write(json.dumps({"execute": command}) + "\n") message = {"execute": command}
if arguments:
message["arguments"] = arguments
stream.write(json.dumps(message) + "\n")
stream.flush() stream.flush()
stream.readline() stream.readline()
finally: finally:
@@ -725,7 +807,14 @@ def run_case(arch, case):
# The bootable FAT32 USB image the build produced (tools/make-fat-image.py), # The bootable FAT32 USB image the build produced (tools/make-fat-image.py),
# presented to the guest as a usb-storage device (see qemu_args). # presented to the guest as a usb-storage device (see qemu_args).
boot_volume = os.path.join(REPO, "zig-out", "danos-usb.img") # Boot a per-run COPY of the image: the guest MUTATES its boot volume (the
# fat tests create/delete files; the logger writes /var/log), and QEMU is
# hard-killed after a match — booting the build artifact in place let one
# run's leftovers fail the next (a stale TESTDIR trips the mkdir-duplicate
# refusal) and dirtied the build cache's own output.
built_volume = os.path.join(REPO, "zig-out", "danos-usb.img")
boot_volume = os.path.join(WORK, "boot-volume.img")
shutil.copy(built_volume, boot_volume)
vars_fd = os.path.join(WORK, "vars.fd") vars_fd = os.path.join(WORK, "vars.fd")
shutil.copy(arch["ovmf_vars"], vars_fd) shutil.copy(arch["ovmf_vars"], vars_fd)
serial = os.path.join(WORK, "serial.log") serial = os.path.join(WORK, "serial.log")
@@ -751,8 +840,10 @@ def run_case(arch, case):
if os.path.exists(qmp_path): if os.path.exists(qmp_path):
os.remove(qmp_path) os.remove(qmp_path)
cmd += ["-qmp", f"unix:{qmp_path},server,nowait"] cmd += ["-qmp", f"unix:{qmp_path},server,nowait"]
qmp_after = case.get("qmp_after") # {"delay": seconds, "command": "..."} # Hooks: a single qmp_after {"delay","command"} or a qmp_sequence list of
qmp_sent = False # {"delay","command","arguments"} — every hook must deliver before a pass.
qmp_hooks = case.get("qmp_sequence") or ([case["qmp_after"]] if case.get("qmp_after") else [])
qmp_pending = [dict(hook, sent=False) for hook in qmp_hooks]
started = time.monotonic() started = time.monotonic()
qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
try: try:
@@ -760,12 +851,13 @@ def run_case(arch, case):
deadline = time.monotonic() + timeout deadline = time.monotonic() + timeout
while time.monotonic() < deadline: while time.monotonic() < deadline:
time.sleep(0.2) time.sleep(0.2)
if qmp_after and not qmp_sent and time.monotonic() - started >= qmp_after["delay"]: for hook in qmp_pending:
try: if not hook["sent"] and time.monotonic() - started >= hook["delay"]:
qmp_send(qmp_path, qmp_after["command"]) try:
qmp_sent = True qmp_send(qmp_path, hook["command"], hook.get("arguments"))
except OSError: hook["sent"] = True
pass # socket not up yet; retry next tick except OSError:
pass # socket not up yet; retry next tick
text = "" text = ""
if os.path.exists(serial): if os.path.exists(serial):
with open(serial, "r", errors="replace") as f: with open(serial, "r", errors="replace") as f:
@@ -773,8 +865,8 @@ def run_case(arch, case):
if fail and fail.search(text): if fail and fail.search(text):
return False, "hit failure marker" return False, "hit failure marker"
if expect.search(text): if expect.search(text):
if qmp_after and not qmp_sent: if any(not hook["sent"] for hook in qmp_pending):
continue # the hook must deliver before the case may pass continue # every hook must deliver before the case may pass
return True, "matched " + repr(case["expect"]) return True, "matched " + repr(case["expect"])
if qemu.poll() is not None: # QEMU exited on its own if qemu.poll() is not None: # QEMU exited on its own
if expect.search(text): if expect.search(text):
+8 -5
View File
@@ -179,14 +179,17 @@ def short_name_for(name, used):
else: else:
base, ext = name, "" base, ext = name, ""
upper_base, upper_ext = base.upper(), ext.upper() upper_base, upper_ext = base.upper(), ext.upper()
# A name fits 8.3 if it is short enough and uses valid characters; a lowercase # A name fits 8.3 if it is short enough and uses valid characters. The raw
# name is simply stored uppercased (FAT is case-insensitive, so the bootloader # 8.3 entry is always uppercase; if that loses the real name's case (e.g.
# and the danos driver still find it). Only genuinely non-8.3 names (too long, # "init" -> "INIT"), a long-name chain carries the exact name. This matters
# e.g. initial-ramdisk.img) get a mangled short name plus LFN entries. # because the EFI loader *enumerates* /system to build the ramdisk — it gets
# back whatever the directory stores, so the stored name must be exact, not
# merely case-insensitively findable.
fits = (1 <= len(base) <= 8 and len(ext) <= 3 fits = (1 <= len(base) <= 8 and len(ext) <= 3
and all(c in VALID_83 for c in upper_base + upper_ext)) and all(c in VALID_83 for c in upper_base + upper_ext))
if fits: if fits:
return (upper_base.ljust(8) + upper_ext.ljust(3)).encode("ascii"), False exact = base == upper_base and ext == upper_ext
return (upper_base.ljust(8) + upper_ext.ljust(3)).encode("ascii"), not exact
# Mangle to STEM~N.EXT. # Mangle to STEM~N.EXT.
stem = "".join(c for c in upper_base if c in VALID_83 and c != " ")[:6] or "FILE" stem = "".join(c for c in upper_base if c in VALID_83 and c != " ")[:6] or "FILE"
index = 1 index = 1
-48
View File
@@ -1,48 +0,0 @@
#!/usr/bin/env python3
"""Build-time initial_ramdisk packer. Concatenates user binaries into one image the
bootloader ferries to the kernel.
Usage: make-initial-ramdisk.py <out.img> [<name> <file>]...
Image layout (little-endian), mirroring src/user/proto/initial-ramdisk.zig:
Header : magic u32 ("DNRD"=0x444E5244), count u32
Entry*N : name [32]u8 (NUL-padded), offset u64, len u64
blobs : each entry's file bytes at its offset
"""
import struct
import sys
MAGIC = 0x444E5244
HEADER = struct.Struct("<II") # magic, count
ENTRY = struct.Struct("<32sQQ") # name[32], offset, len
def main() -> int:
out_path = sys.argv[1]
rest = sys.argv[2:]
if len(rest) % 2 != 0:
sys.stderr.write("usage: make-initial-ramdisk.py <out.img> [<name> <file>]...\n")
return 2
items = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)]
table_end = HEADER.size + len(items) * ENTRY.size
entries = b""
blobs = []
off = table_end
for name, path in items:
with open(path, "rb") as f:
data = f.read()
entries += ENTRY.pack(name.encode()[:31], off, len(data))
blobs.append(data)
off += len(data)
with open(out_path, "wb") as f:
f.write(HEADER.pack(MAGIC, len(items)))
f.write(entries)
for b in blobs:
f.write(b)
return 0
if __name__ == "__main__":
sys.exit(main())
+53
View File
@@ -0,0 +1,53 @@
#!/usr/bin/env python3
"""Pack the bundled binaries into the boot capsule (boot/system.img) — the
single file the EFI loader reads in one sequential pass, which is the only
shape firmware file I/O is fast at (a per-file tree walk measured minutes on
real firmware). The format is the v2 initial_ramdisk (system/initial-ramdisk.zig):
entries named by full FHS path, so the running system is identical whether the
loader read the capsule or walked the tree.
Usage: pack-system-image.py <out.img> [<path> <file>]...
Layout (little-endian): Header{magic "DNR2", count}, Entry{name[64], offset, len}*N, blobs.
"""
import struct
import sys
MAGIC = 0x32524E44 # "DNR2"
HEADER = struct.Struct("<II")
ENTRY = struct.Struct("<64sQQ")
def main() -> int:
out_path = sys.argv[1]
rest = sys.argv[2:]
if len(rest) % 2 != 0:
sys.stderr.write("usage: pack-system-image.py <out.img> [<path> <file>]...\n")
return 2
items = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)]
table_end = HEADER.size + len(items) * ENTRY.size
entries = b""
blobs = []
offset = table_end
for path, source in items:
name = path if path.startswith("/") else "/" + path
encoded = name.encode("ascii")
if len(encoded) > 63:
sys.stderr.write(f"pack-system-image: path too long (>63): {name}\n")
return 2
with open(source, "rb") as f:
data = f.read()
entries += ENTRY.pack(encoded, offset, len(data))
blobs.append(data)
offset += len(data)
with open(out_path, "wb") as f:
f.write(HEADER.pack(MAGIC, len(items)))
f.write(entries)
for blob in blobs:
f.write(blob)
return 0
if __name__ == "__main__":
sys.exit(main())