Author SHA1 Message Date
Daniel Samson 3fc2d5b083 display: finalize v1 — docs + plan to DONE (D5)
The three integration cases (display, display-service, display-demo) plus the
pure host tests all pass; the default build is clean. Update docs/display.md's
"Verifying it" to name the real cases, and mark docs/display-plan.md D1-D5 done.

The display service v1 is complete: a framebuffer compositor that owns the
write-combining framebuffer, composites a z-ordered layer stack into a cacheable
back buffer, presents only the damaged region, and is driven over IPC by the
runtime.display client — proven end to end by the display-demo process. Deferred
by design (docs/display.md): runtime mode-setting and true vsync.
2026-07-14 01:58:18 +01:00
Daniel Samson 105203b447 display: layer client API + the display-demo client (D4)
Drive the compositor from a separate process, proving the pipeline end to end.

- runtime.display: a Layer handle (fill / blitTile / configure / damage /
  destroy), createLayer, and a color(r,g,b) helper that caches the mode and
  packs via protocol.pack. Coordinates are signed over the u32 wire fields
  (@bitCast both ways), so a layer may sit or move partly off-screen.
- system/services/display-demo: the input-source analog for the compositor — a
  full-screen wallpaper, a rectangle it slides back and forth (moved by
  configure each frame, so the damage-driven present repaints old + new), and a
  cursor. Presents in a loop paced by runtime.time. Wired into build + initrd.

Fix this surfaced: protocol.message_maximum was 4096, but the kernel caps every
IPC message at MESSAGE_MAXIMUM = 256, so replyWait rejected the oversized receive
buffer with -E2BIG and the serve loop had been spinning since D2 (invisibly, as
those gates matched init-time heartbeats). Set it to 256; blit_tile is now
explicitly a small-tile path (larger bitmaps are the deferred shm surface).

Gate: `python3 test/qemu_test.py display-demo` — the demo drives frames of
motion through the layer client API and logs `display-demo: ok`. Regression:
zig build test, display (D1), display-service (D2/D3), and default zig build.
2026-07-14 01:56:37 +01:00
Daniel Samson f9cf0007c5 display: layer stack + damage-driven compositor (D3)
Turn the service into a real compositor. A layer is a server-owned surface (its
own mmap'd cacheable buffer) with a screen position, z-order, and visibility.
Clients create layers, draw into them by command, mark damage, and present; the
compositor repaints only the damaged region — clear to the wallpaper, paint the
visible layers bottom-to-top (z-sorted), flush that rectangle back -> front (WC).

- compositor.zig: the pure, host-tested core — Rect (intersect/unite), Surface,
  fillRect, composite (opaque, clipped to a damage rect), blitTile (unaligned-
  safe read of a client tile). No syscall/runtime dependency.
- protocol.zig: pack(format, r, g, b) — native pixel encoding for rgbx/bgrx, the
  shared colour vocabulary of client and server. Host-tested.
- display.zig: the layer table + create/configure/destroy/fill_rect/blit_tile/
  damage/present ops wired onto the compositor, plus a damage-accumulating present.
- Startup self-check: two overlapping layers composited on the real framebuffer,
  read back to confirm the overlap shows the top layer and outside shows the
  bottom — logs `display: compositor self-check ok`.

Gate: `zig build test` green (compositor + pack), and the display-service case's
self-check passes on hardware. Both new pure modules added to the test loop.
2026-07-14 01:39:04 +01:00
Daniel Samson 69b018cc32 display: the compositor service — claim, double-buffer, present (D2)
Stand up /system/services/display: a ring-3 process that claims the framebuffer
D1 seeded, maps it write-combining as the front buffer, allocates a cacheable
back buffer, and presents composed frames. The GUI track's compositor, reached
by name over ServiceId.display (= 9).

- protocol.zig: the display wire protocol (info/create_layer/configure_layer/
  destroy_layer/fill_rect/blit_tile/damage/present). `info` and a whole-screen
  `present` are live; the layer ops fail-stub until D3.
- display.zig: enumerate -> claim -> mmio_map(WC) the LFB, mmap a cacheable back
  buffer, clear it and present it (proving the double-buffer path), then serve.
- runtime.display + barrel exports (display, display_protocol): a cached
  `.display` client with info()/present(), the runtime.block shape.
- init spawns "display" in boot_services; build.zig wires the protocol onto the
  runtime, builds the exe, packs it into the initial-ramdisk, installs it.

mmap fix the back buffer forced: systemMmap was capped at 256 pages (1 MiB) by a
fixed kernel-stack scratch array. Rewrote it to map page-by-page with rollback
(no array) and raised the cap to 8192 pages (32 MiB) — enough for a 4K back
buffer. A real limitation met.

Gate: `python3 test/qemu_test.py display-service` matches the service's own
serial heartbeats (display: online WxH / presented frame 0), printed only after
the full claim -> WC-map -> back-buffer -> present chain. Regression-checked
usermem, heap, init, and D1's display.
2026-07-14 01:26:25 +01:00
Daniel Samson cd812cc00e display: framebuffer handoff primitive + service design (D1)
Kick off the display service track (docs/display.md, docs/display-plan.md): a
user-space compositor that owns the framebuffer. GOP and the PCI display device
are two views of one controller; GOP dies at ExitBootServices, so the portable
base is the boot-handoff linear framebuffer.

D1 makes that framebuffer reachable from user space over the existing device
claim/mmio_map path rather than a bespoke syscall:

- device-abi: a `display` DeviceClass, a DisplayInfo{w,h,pitch,format} on the
  descriptor, and a flags field on resources with a write-combining bit.
- devices-broker: seedDisplay() publishes the loader's framebuffer as a
  root-level `display` node (one WC-flagged memory resource); kmain seeds it
  after discovery. displayDevice()/displayClaimed() track the claim.
- paging/mmio_map: mapUserDeviceInto gains a write_combining bool — a WC-flagged
  resource maps through PAT entry 4 instead of strong-uncacheable (an
  uncacheable framebuffer blit is glacial).
- console: falls silent while a display service holds the framebuffer, and is
  forced back on by the panic/exception paths.

Gate: the `display` kernel test asserts the seeded node's shape and that the
claim + mmio_map leaf is genuinely write-combining (PAT bit set, PCD/PWT clear).
Regression-checked discovery/ioport/claim-release/supervision/device-list/
device-manager with the +1 device in the table.
2026-07-14 00:54:56 +01:00
Daniel Samson f157a93c9c fat/usb-storage: flush the device cache on close so writes survive power-off
init writes /mnt/usb/DANOS.LOG at shutdown and then enters S5 — but the write sat in the USB flash controller's write cache and was lost when power was cut, because nothing issued SCSI SYNCHRONIZE CACHE. Invisible in QEMU (its backing file commits immediately); real on hardware. The boot-time flush survived only because the machine kept running afterward and the cache drained on its own.

- block protocol gains a flush op; usb-storage serves it with SYNCHRONIZE CACHE (10); runtime.block gains Device.flush()
- the FAT server tracks whether blocks were written and, on a file close, commits the device cache (durable-on-close — the right default for removable media, and it makes init's existing shutdown close() persist the log before S5, no init/VFS change needed)
- verified the SYNCHRONIZE CACHE actually reaches the driver on each dirty close; usb-storage/fat-mount/mutations/rename/mtime/log-flush all green
2026-07-13 23:29:15 +01:00
Daniel Samson 88644e57d6 acpi: stop parsing AML in the kernel; power management is userspace's
Kernel discovery interpreted the whole DSDT/SSDTs (~0.5 MB -> ~9700 nodes) solely to extract the _S5 sleep type for a kernel-side soft-off — ~1-2s of work on the single-core, pre-scheduler critical path, with nothing to overlap it. The ring-3 acpi service already parses the same blobs and owns S5 end-to-end (reads pm1a_cnt from the FADT, writes SLP_TYP itself; orderly-shutdown tests it). So drop the kernel parse entirely.

- all *static*-table parsing stays (MADT/HPET/FADT/MCFG): CPUs, timers, PCIe, and the power register map are still detected from ACPI, not legacy/compat addresses (timer still calibrates via HPET/PM-timer/CPUID, PIT only as last resort)
- kernel keeps reboot (FADT reset register + legacy fallbacks — no AML); soft-off is userspace-only now
- removed PowerInformation.s5/s3, the kernel AML namespace, aml_stats, amlDeviceCount, and the aml import; power.shutdown/enable/sleepS3 gone
- tests: poweroff case retired (S5 covered by orderly-shutdown); discovery drops its S5/AML checks; acpi-parse self-verifies against a device-count floor since there's no kernel count to match
2026-07-13 23:09:30 +01:00
Daniel Samson 10c11d1806 paging: map the framebuffer write-combining, and clear it after paging
The console's one-time full-screen clear was millions of individual uncached word-writes to the GPU BAR: fine in QEMU, but on real hardware firmware MTRRs force that region uncacheable, so the clear crawls. Program PAT entry 4 to write-combining (setupPat, on the BSP and every AP) and map the framebuffer window with the PAT bit, so those writes batch into bursts.

The clear ran on the *loader's* uncached mapping because console.init happened before enablePaging. Move it to just after paging, so it uses the kernel's write-combining mapping instead. The cost: on-screen output is absent before paging (an early panic still lands in the serial/RAM log). Verified: 11 QEMU cases (incl. SMP AP-PAT, USB/FAT/ACPI) plus a screendump showing the framebuffer cleared to black with the status text rendered correctly.
2026-07-13 22:40:39 +01:00
Daniel Samson c4595700ba paging: map the physmap with 2 MiB huge pages
The physmap (the kernel's permanent window onto all physical RAM) was built one 4 KiB page at a time: on a 64 GiB machine that is 16.7M mapPage calls and ~128 MiB of page tables (the bulk of the 'Kernel footprint' line). Map the 2 MiB-aligned interior with huge PD leaves instead — one entry per 2 MiB, no PT beneath — and only the unaligned head/tail with 4 KiB. 512x fewer entries and table frames; footprint at 8 GiB drops from ~16 MiB of tables to ~1 MiB total, and it scales.

translateIn/isExecutable now stop at a 2 MiB leaf, and descend() panics rather than walking through one (a 4 KiB map inside a huge page would corrupt RAM; the physmap and 4 KiB regions live in disjoint PML4 slots, so it can't happen — the guard just makes a bug loud). Verified: 20 QEMU cases (vmm/heap/wx/usermem/dma/ipc/smp/usb/fat/acpi/faults) green.
2026-07-13 22:22:40 +01:00
Daniel Samson 28b4dabbaa serial: make the log sink a build option, off by default
Serial is now a QEMU/dev aid, not a real-hardware necessity: a legacy-free board often has no live COM1, and the boot log is kept in RAM (klog) and flushed to disk. So the serial sink is compiled in only under -Dserial (default false).

- kernel.zig gates serialInit + the log sink on build_options.serial
- boot/efi.zig gates its EFI: progress breadcrumbs (con_out) via progress(); fatal-error messages stay always-on so a failed boot still explains itself
- run-x86-64 boots a serial-enabled image variant (factored addKernel/addBootImage helpers) so a dev boot always captures serial0, without baking serial into the flashable image
- test/qemu_test.py builds -Dserial=true (it asserts on serial markers)
- the loopback probe stays as a real-HW safety net for -Dserial images
2026-07-13 22:05:08 +01:00
Daniel Samson 4cb4f2a80f serial: skip a dead COM1 via a loopback probe (fixes ~57s real-HW boot)
On a legacy-free board COM1 is decoded but has no live UART: its LSR reads 0x00, so the transmit-holding-empty bit never sets and writeByte spun its full 100k-iteration guard on every logged byte (~15ms/byte x ~3.7k bytes to 'initialised' ~= 57s). QEMU always has a working UART, so this only bit real hardware; the boot log survived via log.ramSink.

- probe() loopback-tests the UART in init(); write() is a no-op when absent, so a dead port costs nothing per byte
- guard cut 100k -> 5k (backstop for a live-but-stalled UART only)
- serialPresent() + a boot line making the absent-UART case visible
2026-07-13 21:15:56 +01:00
Daniel Samson 67702fa250 Phase 2d (ii): filesystem modification time (mtime)
Completes Phase 2: the FAT filesystem now stamps and reports a real modification
time, built on the Phase 2d(i) kernel wall-clock. This is the last stat field the
compiler's build cache needs to reason about (source vs cached output).

- on-disk.zig: fatToEpoch / epochToFatDateTime convert between the two 16-bit DOS
  date/time fields and Unix epoch seconds (UTC — FAT has no timezone). Host-tested
  round-trip + an absolute check (1577836800 == 2020-01-01).
- engine: a settable current_time_epoch that create/write stamp into the entry's
  write (and creation) date/time; Node/Listing gained an mtime decoded from those
  fields on read. Host test: a create stamps the mtime, read back through resolve
  and listEntry.
- vfs protocol FileStatus + runtime.fs.Attributes gained an mtime field; the fat
  server sets current_time_epoch from runtime.system.wallClock() per request and
  returns mtime from stat. The flat ramfs reports 0 (it has no timestamps).
- fat-test reads the created file's mtime through stat and checks it is a real
  current time, behind a new `fat-mtime` QEMU case.

Verified against the host: the guest stamped mtime 1783971676 while the host clock
was 1783971680 (boot+test lag) — the file's mtime is real current time. zig build,
zig build test (the epoch<->DOS conversions + the engine mtime test),
zig build check-fat-image, and a sequential QEMU sweep — fat-mount, fat-mutations,
fat-rename, fat-mtime, vfs, vfs-client-death, log-flush, orderly-shutdown,
initial-ramdisk, smoke, wall-clock, usb-storage — all green. mode/inode remain.
2026-07-13 20:43:40 +01:00
Daniel Samson 8a38540312 Phase 2d (i): a kernel wall-clock from the CMOS RTC
Adds real (calendar) time, the foundation for filesystem timestamps. Monotonic
time (`clock`) says how long since boot; this says what time it actually is.

- cpu.zig (x86_64): readRtcUnixSeconds() reads the CMOS real-time clock (ports
  0x70/0x71) — waits out an update-in-progress, reads twice until stable, handles
  BCD-vs-binary and 12-vs-24-hour per status register B — and converts to Unix
  epoch seconds (UTC).
- kernel/wall-clock.zig: reads the RTC once at boot and anchors it to the monotonic
  clock, so a query is a cheap arithmetic offset — no per-call CMOS poll, no lock,
  no SMP hazard on the shared ports. kmain calls init() once the monotonic clock is
  final and logs the epoch.
- wall_clock() syscall (33) -> Unix epoch seconds, wrapped by runtime.system
  .wallClock(). Wall-clock *seconds* are mechanism the kernel owns like the
  monotonic clock; calendars/timezones are user-space policy (the stale comment on
  systemClock that called wall-clock a "user-space service" is updated in spirit by
  the new handler's doc).
- A `wall-clock` kernel test asserts the boot RTC read is a plausible current epoch.

Verified against the host: the guest read epoch 1783971244 while `date -u +%s` gave
1783971245 (one second of boot lag) — the CMOS read + epoch conversion are correct to
the second. zig build, zig build test, and smoke/clock/init are green.
2026-07-13 20:35:03 +01:00
Daniel Samson 54635eecf5 Phase 2c: rename, wired through the VFS to runtime.fs
Completes the Phase 2 FAT mutation set (truncate, mkdir, unlink, rename).

- engine: rename(dir, old_name, new_name) rewrites an existing entry's 8.3 name in
  place within the same directory. Refuses a missing source, a non-8.3 target, or a
  name that already exists; drops any long-name entries on the old file (it takes
  its new 8.3 name), LFN-aware like removeFile. Cross-directory and long-name-
  preserving rename are noted limitations. Host-tested (rename keeps contents;
  collision, non-8.3, and missing-source are refused).
- vfs protocol: a `rename` operation whose payload is old-path, a 0x00 separator,
  then new-path.
- VFS router: a forwardRename helper + a `.rename` case that requires both paths
  under the same mount (cross-filesystem rename is refused) and forwards the
  mount-relative old+new.
- fat server: a `.rename` handler that requires the same parent directory and calls
  engine.rename.
- runtime.fs: rename(old_path, new_path).
- fat-test now renames the file it created (before removing it) and asserts the old
  name is gone, behind a new `fat-rename` QEMU case.

Verified: zig build, zig build test (the engine rename unit test), zig build
check-fat-image, and a sequential QEMU sweep — fat-mount, fat-mutations, fat-rename,
vfs, vfs-client-death, log-flush, orderly-shutdown, initial-ramdisk, smoke — green.
2026-07-13 20:22:32 +01:00
Daniel Samson 184d90c2c6 Phase 2b: wire mkdir + unlink through the VFS to runtime.fs
The engine gained mkdir/unlink in Phase 2a; this exposes them as first-class
filesystem operations so programs can use them.

- vfs protocol: two new path-based operations, mkdir and unlink (appended, so
  existing opcodes/offsets are unchanged).
- VFS router: a forwardPath helper relays a path-based op under a mount to its
  backend; the mkdir/unlink cases forward to the mounted filesystem (the flat
  ramfs refuses them — it has no directories).
- fat server: mkdir -> engine.createDirectory, unlink -> engine.removeFile, each
  resolving the parent via a shared splitParent helper (also used by open-create).
- runtime.fs: makeDirectory(path) and remove(path).
- fat-test now exercises the whole path — mkdir /mnt/usb/TESTDIR, create + write +
  read a file inside it, then remove it — behind a new `fat-mutations` QEMU case.

Verified: zig build, zig build test, zig build check-fat-image, and a sequential
QEMU sweep — fat-mount, fat-mutations (mkdir/write/read/unlink through the mount),
vfs, vfs-client-death, log-flush, orderly-shutdown, initial-ramdisk, smoke — green.
2026-07-13 20:12:07 +01:00
Daniel Samson a32eed877d Phase 2a: FAT engine mutations + O_TRUNC (fix the overwrite corruption)
The FAT engine could create, read, write, and grow files, but never free clusters
or make directories — so overwriting a shorter file left a stale tail (a real bug:
corrupt boot-log re-flushes, and later corrupt compiler cache/.o files). This adds
the mutation half of the engine, with the corruption fix wired all the way through.

Engine (system/services/fat/engine.zig), all host-tested:
- freeChain: return a cluster chain to the pool (bounded against a corrupt cycle) —
  the shared primitive under truncate and remove.
- truncate: free the chain and zero the entry's size/first-cluster (O_TRUNC).
- createDirectory (mkdir): allocate + initialise a cluster with "." and ".." and add
  the directory entry to the parent.
- removeFile (unlink): free the chain and mark the 8.3 entry plus any preceding
  long-name entries deleted, so a reused slot can't inherit an orphaned long name.
  Refuses directories.
- createFile now shares a common addEntry helper with createDirectory.

O_TRUNC wired end to end: a truncate open-flag (vfs protocol) that the router already
forwards; runtime.fs.OpenOptions.truncate; and fat's handleOpen calls engine.truncate
on an existing file. The boot-log flush (log-flush + init) now opens with truncate, so
a shorter log on a later boot of the same stick leaves no stale tail — closing the
caveat from the boot-log work.

mkdir/unlink are engine-complete and host-tested but not yet exposed as VFS
operations / runtime.fs methods (they are new path-based ops needing router cases);
rename and the richer stat (mtime/mode, blocked on wall-clock) remain. See
docs/zig-self-hosting.md (Phase 2).

Verified: zig build, zig build test (8 engine host tests, incl. truncate, the
overwrite-no-stale-tail regression, remove, and mkdir), zig build check-fat-image, and
a sequential QEMU sweep — fat-mount, log-flush (DANOS.LOG read back at 11804 bytes,
clean, with the truncate-based final flush), vfs, vfs-client-death, usb-storage,
orderly-shutdown, initial-ramdisk, smoke — all green.
2026-07-13 20:02:01 +01:00
Daniel Samson 347a041d85 Phase 1a: add the native runtime.fs, retire the posix shim
The first step of the Zig self-hosting roadmap (docs/zig-self-hosting.md): give
danos programs a danos-native file API and remove the premature POSIX compatibility
shim. This also resolves the earlier misplacement of a full-write helper into the
compat layer — that behaviour now lives natively in runtime.fs.File.writeAll.

- library/runtime/fs.zig: the danos-native file client over the VFS (open/read/
  write/writeAll/seekTo/attributes/close, directory listing, mount). Handles are
  *values* — a File/Directory owns its VFS node id and byte offset — so there is no
  per-process fd table or descriptor limit, unlike the POSIX fd model the shim
  emulated. This is where the operations that later become std.os.danos are staged.
- Retire library/posix/ (unistd, stdio): only five call sites used it, all file
  operations, all migrated to runtime.fs — fat (mount), the vfs-test and fat-test
  clients, and init/log-flush (the boot-log flush). stdio was already dead.
- build.zig: drop the posix module, its addUserBinary parameter, the per-binary
  import, and the ~26 call-site arguments.
- Docs: the VFS protocol's client is now runtime.fs; the docs index and
  coding-standards note posix is retired and the foreign-ABI naming exception now
  applies to the future std.os.danos seam; the process-lifecycle note points the
  future musl layer at that same seam rather than the deleted directory.

Deferred by design (see the roadmap): the C-ABI runtime.os errno seam is built at
fork time (its shape must match std/os/danos.zig); truncate/mkdir/rename are
Phase 2; stdio-byte fds and cwd are later slices.

Verified: zig build, zig build test, zig build check-fat-image, and a sequential
QEMU sweep — vfs, vfs-client-death (the park/hold-handle path), fat-mount, log-flush,
orderly-shutdown, initial-ramdisk (log-flush silent in the bare sweep), smoke, init,
usb-storage, device-manager — all green.
2026-07-13 19:43:56 +01:00
Daniel Samson 53e42837e0 Docs: add the Zig self-hosting roadmap
A forward-looking design note on making danos a real Zig target
(-target x86_64-danos) and eventually running the compiler on it, focused on
the standard-library surface (not the editor/terminal).

The core realisation: Zig 0.16 (post-writergate) collapses an OS port to ONE
seam — std.fs is gone, everything routes through the std.Io vtable, and
std.posix is generic over a single per-OS `system` module (std.os.<tag>). So the
port is "write std.os.danos once" and the whole fs/process/Io tower lights up,
rather than reimplementing the namespaces.

Records the decisions this shapes now: build runtime.os (the seam, promoted into
a forked std/os/danos.zig later) plus a thin runtime.fs; retire the premature
library/posix shim (only 5 unistd call sites); do NOT hand-mirror the high-level
std namespaces; do NOT emulate the Linux ABI; defer musl. Covers the host/target/
self-host roles and the four-part compiler fork, a coverage table of what danos
has vs the gaps (mkdir/unlink/rename/truncate, richer stat, wall-clock, env, cwd,
entropy, stdio bytes), a phased plan (target -> read-side+retire-posix ->
fs-mutation+stat -> single-threaded self-linked compiler), and the risks
(fork rebase treadmill, -fsingle-threaded and -fno-llvm/-fno-lld being
load-bearing, the "w"-does-not-truncate corruption bug). Linked from the docs index.
2026-07-13 19:26:51 +01:00
Daniel Samson f52c591f5e Persist the kernel boot log to the USB FAT volume
On a headless or real board nothing captures serial, so the boot log — the whole
diagnostic stream — is lost at power-off. This retains it in the kernel and copies
it to the boot USB volume as /mnt/usb/DANOS.LOG, the on-disk equivalent of QEMU's
`-serial file:`. Pull the stick, read DANOS.LOG on another machine.

How it fits together:
- Kernel RAM sink (log.zig): a fixed 256 KiB in-image buffer registered as a log
  sink in kmain, right after serial. Because userspace debug_write funnels through
  log.write, it captures the entire stream — kernel lines and every service's
  output — from the first line. Fills linearly and stops when full (earliest boot
  output, the most valuable, is kept); no allocation, so it is panic-safe.
- klog_read syscall (32): copies that buffer out to a user buffer, the mirror of
  debug_write — same overflow-safe user-half bounds check, kernel -> user copy,
  under the kernel lock so the snapshot can't grow mid-copy. Wrapped by
  runtime.system.klogRead.
- log-flush (new one-shot, in the initial-ramdisk): waits for the fat server to
  mount /mnt/usb, then copies the whole log to /mnt/usb/DANOS.LOG. init spawns it
  once the boot services are up (fire-and-forget; it polls the mount itself). If
  no volume is mounted — no stick, or the no-VFS ramdisk sweep — it exits silently.
- init shutdown flush: init repeats the copy inline at the top of shutDown(),
  BEFORE it tears down the storage services (the fat server is stopped first), so
  a clean poweroff captures the fullest log while /mnt/usb is still writable.
- unistd.writeAll: loops write() past the 224-byte VFS payload cap; both flush
  paths use it.

The filename is 8.3 (DANOS.LOG) at the mount root — the FAT short-name rule, and
there is no mkdir on the FAT path yet. Extend-only writes mean the two same-session
flushes never leave stale bytes (the shutdown log is a superset of the boot log);
a shorter log on a later boot of the same stick can leave a stale tail — a noted,
cosmetic limitation, not worth pulling O_TRUNC into the FAT write path for now.

Verified end to end under QEMU: a new `log-flush` case (reusing the orderly-shutdown
build) asserts both markers then S5, and DANOS.LOG is read back out of the image
afterwards (11804 bytes, containing the kernel init line, the FAT mount line, and
the boot flush marker). Regression stays green: zig build, zig build test, zig
build check-fat-image, and a sequential QEMU sweep — smoke, init, initial-ramdisk
(log-flush silent in the bare sweep), orderly-shutdown, fat-mount, usb-storage,
usb-hid, vfs, input, device-manager, process, signals, dma, fault-pf.
2026-07-13 18:18:34 +01:00
Daniel Samson 77d2e22ed1 M7: in-repo FAT32 image builder + boot the whole system off a USB stick
The system now boots off a real FAT32 filesystem on a USB mass-storage device
instead of QEMU's synthesized VVFAT drive. A new in-repo image builder formats
that filesystem from the FHS boot tree, and both QEMU call sites (the run step
and the test harness) attach it as a usb-storage device on the xHCI bus, so
every boot exercises the full USB path OVMF -> BOOTX64.efi -> kernel.

- tools/make-fat-image.py: a Python 3 stdlib-only FAT32 formatter (mirrors
  tools/make-initial-ramdisk.py — no external host dependencies). It lays down
  the boot sector + BPB/EBPB32, FSInfo, backup boot sector, two FATs, and the
  root/subdir/file cluster chains, emitting long-name entries where a name is
  not 8.3. Packs the four boot inputs (EFI/BOOT/BOOTX64.efi, system/kernel,
  system/services/init, boot/initial-ramdisk.img) into their boot paths. A
  --verify subcommand re-checks the 0xAA55 signature, recomputes the cluster
  count -> FAT32, and resolves EFI/BOOT/BOOTX64.efi, all with no dependencies.

- build.zig: a mk_fat step builds zig-out/danos-usb.img from the four boot
  artifacts (so changing -Dtest-case rebuilds the image with that kernel), a
  check-fat-image step runs --verify, and run-x86-64 boots the image on a
  usb-storage device (if=none,id=bootusb + usb-storage,bus=xhci.0,bootindex=0),
  keeping usb-kbd/usb-mouse on the same controller.

- test/qemu_test.py: the default boot config now boots off danos-usb.img on a
  usb-storage device (xHCI + usb-kbd + usb-mouse + the boot stick). The seven
  per-case qemu_extra blocks that added their own qemu-xhci/usb-kbd/usb-mouse
  (or a VVFAT stick) collided on id=xhci and are removed — the default provides
  the bus and the boot device. usb-storage and fat-mount now exercise the real
  FAT32 boot image (usb-storage reads its 0x55AA boot sector; fat mounts it at
  /mnt/usb). A build_case override lets a case reuse another's kernel, used by a
  new usb-boot case: an explicit, named boot-from-USB regression guard.

Verified: zig build, zig build test, and zig build check-fat-image are green
(FAT32, 128992 clusters, BOOTX64.efi present); a broad sequential QEMU sweep
passes — smoke, init, vfs, input, device-manager, usb-report, usb-hid,
usb-storage, fat-mount, device-list, driver-restart, acpi-report, iommu,
orderly-shutdown, usb-boot, dma, msi, initial-ramdisk, args, process — proving
the boot switch holds across kernel tests, the full init tree, the USB stack,
the FAT mount, and orderly shutdown.
2026-07-13 15:34:18 +01:00
Daniel Samson a64a01a6a9 M6: FAT read/write filesystem server, mounted into the VFS
Add a FAT12/16/32 filesystem the VFS mounts at /mnt/usb, reading and writing a
USB stick through the block device. Verified end to end under QEMU: the fat
server mounts the volume, the VFS routes /mnt/usb to it, and a client lists the
root and reads a file (the ELF magic of /mnt/usb/system/kernel).

- engine.zig: the FAT engine over a BlockDevice interface — mount (a bare FAT or,
  as QEMU's VVFAT and most real sticks present it, an MBR-partitioned disk), FAT
  chain walk (12/16/32), cluster allocation, directory traversal with long-name
  read, and file read / write / create. Host-tested against a RAM-backed FAT16
  image (create, cluster-spanning write, mid-file overwrite, read-back, list).
- on-disk.zig: the align(1) boot-sector / directory / long-name / FSInfo structs
  and the cluster-count FAT-type detection.
- fat.zig: the server — wraps the .block device (a DMA bounce buffer) in a
  BlockDevice, mounts the FAT, serves the vfs-protocol as a backend, and mounts
  itself into the VFS at /mnt/usb. Spawned by init as a boot service.
- runtime.block: the block-device client (geometry / read / write by physical
  address, so whole sectors never cross IPC).
- Raise the kernel service-name registry (maximum_services) 8 -> 16: it is
  indexed directly by ServiceId, and fat = 8 was being rejected, so the fat
  server exited before registering.
- VFS: an absolute path with no matching mount is now not-found rather than
  silently created in the flat ramfs — so /mnt/usb fails cleanly until mounted.

Tests: fat-mount (the full stack: block -> FAT -> VFS mount -> list + file read)
passes; host units cover the engine and on-disk structs; the vfs, shutdown, and
USB regression suite stays green (10/10).
2026-07-13 15:05:53 +01:00
Daniel Samson 35e8921de8 M5: VFS mount support — mount table + forwarding router
Turn the flat-ramfs VFS into a router: a mount table maps an absolute path prefix
(e.g. /mnt/usb) to a backend server's endpoint, and open/read/write/status/
readdir/close on a path under a mount are forwarded to that backend, which speaks
the same vfs-protocol. This is what a FAT filesystem mounts into.

- protocol: append readdir / mount / unmount operations, a NodeKind enum (the FSH
  file types) that now fills FileStatus.kind, a DirectoryEntry record, and a
  directory open flag. Appended values keep existing clients and tests unchanged.
- vfs.zig: a mount table, longest-prefix routing, forwarding of every op on a
  backend handle, mount/unmount handlers (the backend arrives as the call's
  capability), and release-on-death that also closes the backend's handles.
- path.zig: pure, host-tested mount-prefix matching that never captures a
  non-boundary like /mnt/usbextra.
- unistd: mount(), opendir / readdir / closedir clients.

Bare names still resolve in the flat ramfs — the backward-compat contract; the
vfs and vfs-client-death tests pass unchanged. End-to-end mount+read is exercised
by the FAT server (M6). Host units cover path matching and protocol sizes.
2026-07-13 14:21:30 +01:00
Daniel Samson 3fb9d5936a USB driver stack: xHCI transfers, HID keyboard/mouse, mass storage
Flesh out the xHCI host-controller driver into a full transfer engine and build
the three USB class drivers on top, all verified end to end under QEMU.

- xHCI engine (usb-xhci-library.zig): controller reset, command/event rings with
  cycle-bit bookkeeping (gated on a No-Op-command proof), device slots, Address
  Device, control transfers, full chapter-9 enumeration, Configure Endpoint, and
  interrupt/bulk transfers. Each interface is device_registered with its
  (class,subclass,protocol) identity, unique per (port,interface).
- Bus<->class transfer protocol (usb-transfer-protocol.zig + runtime.usb): open /
  control / interrupt-subscribe (async report pump on a poll timer) / bulk-by-
  physical-address, so sector data never crosses the 256-byte IPC limit.
- USB HID keyboard + mouse (usb-hid/): decode boot-protocol reports and publish
  to the input service. A USB usage is already the input protocol's keycode.
- USB mass storage (usb-storage/): Bulk-Only Transport + transparent SCSI,
  serving a block device under the new .block service id (block-protocol).
- device-manager matches USB interfaces to class drivers (usbDriverForIdentity).
- usb-abi / usb-ids made importable modules; add HID and mass-storage class
  requests, packTriple, and a usb_device DeviceClass.
- Fix test/qemu_test.py on macOS: the QMP unix-socket path was built from the
  deep worktree path and exceeded the 104-byte sun_path limit, so QEMU exited
  before booting. It now lives under a short temp path.

Tests: usb-report, usb-hid, usb-storage pass under python3 test/qemu_test.py;
host units (usb-abi, usb-ids, hid-report, bulk-only-transport, scsi) green.
2026-07-13 14:11:00 +01:00
66 changed files with 8774 additions and 808 deletions
+12 -3
View File
@@ -2,6 +2,7 @@ const std = @import("std");
const uefi = std.os.uefi;
const elf = std.elf;
const boot_handoff = @import("boot-handoff");
const build_options = @import("build_options");
const BootInformation = boot_handoff.BootInformation;
const GraphicsOutput = uefi.protocol.GraphicsOutput;
const EdidActive = uefi.protocol.edid.Active;
@@ -84,7 +85,7 @@ fn boot() !noreturn {
// the map and exiting would invalidate the map key.
const cr3 = try buildBootstrapTables(bs, &boot_information);
log("EFI: kernel loaded, exiting boot services\r\n");
progress("EFI: kernel loaded, exiting boot services\r\n");
boot_information.memory_map = try exitBootServices(bs);
// Switch onto our tables and jump to the kernel in one uninterruptible step.
@@ -395,7 +396,7 @@ fn loadInit(bs: *uefi.tables.BootServices, boot_information: *BootInformation) !
const image = try loadFile(bs, init_file_name);
boot_information.init_base = @intFromPtr(image.ptr);
boot_information.init_len = image.len;
log("EFI: /system/services/init loaded\r\n");
progress("EFI: /system/services/init loaded\r\n");
}
/// Ferry the initial_ramdisk (the VFS server + drivers) to the kernel, same as init.
@@ -403,7 +404,7 @@ fn loadInitialRamdisk(bs: *uefi.tables.BootServices, boot_information: *BootInfo
const image = try loadFile(bs, initial_ramdisk_file_name);
boot_information.initial_ramdisk_base = @intFromPtr(image.ptr);
boot_information.initial_ramdisk_len = image.len;
log("EFI: initial_ramdisk loaded\r\n");
progress("EFI: initial_ramdisk loaded\r\n");
}
/// Validate the ELF, copy every PT_LOAD segment to its physical address, and
@@ -561,6 +562,14 @@ fn log(comptime message: []const u8) void {
_ = out.outputString(std.unicode.utf8ToUtf16LeStringLiteral(message)) catch {};
}
/// A boot-progress breadcrumb: like `log`, but compiled out unless `-Dserial`
/// (off by default), so a real-hardware boot stays silent. Fatal errors use
/// `log` directly and always show, so a failed boot still explains itself.
fn progress(comptime message: []const u8) void {
if (!build_options.serial) return;
log(message);
}
/// Write a runtime ASCII byte string (e.g. an @errorName) by widening to UTF-16.
fn logBytes(bytes: []const u8) void {
const out = uefi.system_table.con_out orelse return;
+270 -76
View File
@@ -58,7 +58,6 @@ fn addUserBinary(
b: *std.Build,
target: std.Build.ResolvedTarget,
runtime_module: *std.Build.Module,
posix_module: *std.Build.Module,
mmio_module: *std.Build.Module,
xkeyboard_config_module: *std.Build.Module,
acpi_ids_module: *std.Build.Module,
@@ -78,9 +77,6 @@ fn addUserBinary(
.stack_protector = false,
.imports = &.{
.{ .name = "runtime", .module = runtime_module },
// POSIX/C compatibility layer, available to any program that wants it
// (danos-native code uses `runtime` directly). See library/posix/.
.{ .name = "posix", .module = posix_module },
// Typed volatile MMIO + memory barriers, for drivers. See library/mmio/.
.{ .name = "mmio", .module = mmio_module },
// Keyboard layouts (keycode + modifiers -> keysym/character), available
@@ -100,6 +96,105 @@ fn addUserBinary(
return exe;
}
/// The modules the kernel imports, gathered once so both kernel variants (the
/// installed one and the serial-enabled one `run-x86-64` boots) are built from
/// the same set. `build_options` is *not* here — it carries `serial`/`test_case`,
/// which differ per variant, so `addKernel` builds it fresh each time.
const KernelModules = struct {
boot_handoff: *std.Build.Module,
abi: *std.Build.Module,
device_abi: *std.Build.Module,
architecture: *std.Build.Module,
platform: *std.Build.Module,
parameters: *std.Build.Module,
initial_ramdisk: *std.Build.Module,
};
/// Build the freestanding x86_64 kernel ELF. Factored so we can build it twice
/// from one recipe: the installed/flashable image (serial off by default) and the
/// serial-enabled variant `run-x86-64` boots — they differ only in the `serial`
/// build option baked into `build_options`.
fn addKernel(
b: *std.Build,
kernel_target: std.Build.ResolvedTarget,
optimize: std.builtin.OptimizeMode,
modules: KernelModules,
test_case: ?[]const u8,
serial: bool,
) *std.Build.Step.Compile {
// Compile-time configuration the kernel reads as `@import("build_options")`:
// the QEMU harness's -Dtest-case, and whether the serial log sink is compiled
// in (see the -Dserial option). Built per variant since `serial` differs.
const build_options = b.addOptions();
build_options.addOption(?[]const u8, "test_case", test_case);
build_options.addOption(bool, "serial", serial);
const build_options_module = build_options.createModule();
const exe = b.addExecutable(.{
.name = "kernel",
.root_module = b.createModule(.{
.root_source_file = b.path("system/kernel/kernel.zig"),
.target = kernel_target,
.optimize = optimize,
.code_model = .kernel, // kernel runs in the top 2 GiB (higher half)
.red_zone = false, // interrupts would corrupt the SystemV red zone
.single_threaded = false, // SMP: the big kernel lock's atomics must be real across cores
.sanitize_c = .off, // the UBSan runtime needs f128/SSE support we don't provide
.stack_check = false, // stack-probe calls have no runtime to land in
.stack_protector = false,
.imports = &.{
.{ .name = "boot-handoff", .module = modules.boot_handoff },
.{ .name = "abi", .module = modules.abi },
.{ .name = "device-abi", .module = modules.device_abi },
.{ .name = "architecture", .module = modules.architecture },
.{ .name = "platform", .module = modules.platform },
.{ .name = "parameters", .module = modules.parameters },
.{ .name = "build_options", .module = build_options_module },
.{ .name = "initial-ramdisk", .module = modules.initial_ramdisk },
},
}),
});
exe.setLinkerScript(b.path("system/kernel/architecture/x86_64/linker.ld"));
exe.entry = .{ .symbol_name = "_start" };
// The self-hosted linker ignores parts of the linker script (PHDRS,
// /DISCARD/, AT(), section order); the higher-half layout depends on the
// script being authoritative, so pin the kernel to LLVM + LLD.
exe.use_llvm = true;
exe.use_lld = true;
// Higher-half virtual base (matches KERNEL_VIRT_BASE in linker.ld); the
// linker's AT() clauses give each segment a low physical load address
// (.text at 1 MiB), which the loader allocates and copies into.
exe.image_base = 0xFFFFFFFF80100000;
return exe;
}
/// Assemble the bootable FAT32 image (the in-repo Python builder) holding what
/// the firmware and loader need off the ESP: the EFI stub, `kernel`, `init`, and
/// the initial-ramdisk. Factored so the serial-enabled `run-x86-64` variant can
/// bundle its own kernel while sharing the (serial-independent) loader, init, and
/// ramdisk. Returns the image's LazyPath.
fn addBootImage(
b: *std.Build,
kernel_bin: std.Build.LazyPath,
efi_bin: std.Build.LazyPath,
init_bin: std.Build.LazyPath,
initial_ramdisk_img: std.Build.LazyPath,
) std.Build.LazyPath {
const mk_fat = b.addSystemCommand(&.{"python3"});
mk_fat.addFileArg(b.path("tools/make-fat-image.py"));
const fat_image = mk_fat.addOutputFileArg("danos-usb.img");
mk_fat.addArg("64"); // MiB
mk_fat.addArg("EFI/BOOT/BOOTX64.efi");
mk_fat.addFileArg(efi_bin);
mk_fat.addArg("system/kernel");
mk_fat.addFileArg(kernel_bin);
mk_fat.addArg("system/services/init");
mk_fat.addFileArg(init_bin);
mk_fat.addArg("boot/initial-ramdisk.img");
mk_fat.addFileArg(initial_ramdisk_img);
return fat_image;
}
pub fn build(b: *std.Build) void {
ensureZigVersion();
@@ -142,6 +237,28 @@ pub fn build(b: *std.Build) void {
.root_source_file = b.path("system/devices/acpi-ids.zig"),
});
// The USB device-framework wire ABI (chapter-9 set-up packets, standard +
// class requests, descriptors) and the USB class-code taxonomy — the flat
// reference the xHCI bus driver, the USB class drivers, and the device
// manager's identity matcher all share. Pure data, like pci-class/acpi-ids.
const usb_abi_module = b.addModule("usb-abi", .{
.root_source_file = b.path("system/devices/usb-abi.zig"),
});
const usb_ids_module = b.addModule("usb-ids", .{
.root_source_file = b.path("system/devices/usb-ids.zig"),
});
// The USB transfer protocol: what a USB class driver says to the xHCI bus
// driver to drive its device (open / control / interrupt / bulk). A protocol
// module like vfs-protocol, shared by the bus driver and every class driver.
const usb_transfer_protocol_module = b.addModule("usb-transfer-protocol", .{
.root_source_file = b.path("system/drivers/usb-xhci-bus/usb-transfer-protocol.zig"),
});
// The block-device protocol: read/write of fixed-size blocks, spoken between a
// filesystem and a block driver (usb-storage). A protocol module like the rest.
const block_protocol_module = b.addModule("block-protocol", .{
.root_source_file = b.path("system/services/block/protocol.zig"),
});
// Kernel tunables (maximum_cpus, stack sizes, tick rate). A dependency-free module of
// compile-time constants, imported wherever a knob is read; keeps the trade-offs
// in one place instead of scattered across the tree. See system/parameters.zig.
@@ -225,6 +342,18 @@ pub fn build(b: *std.Build) void {
.root_source_file = b.path("system/services/device-manager/device-manager-protocol.zig"),
});
runtime_module.addImport("device-manager-protocol", device_manager_protocol_module);
// The USB transfer protocol, so runtime.usb (the class-driver client) can speak
// it, the way runtime.input speaks the input protocol.
runtime_module.addImport("usb-transfer-protocol", usb_transfer_protocol_module);
// The block protocol, so runtime.block (the block-device client) can speak it.
runtime_module.addImport("block-protocol", block_protocol_module);
// The display protocol, so runtime.display (the compositor client) and the display
// service both speak it through the runtime, like the other protocol modules.
const display_protocol_module = b.addModule("display-protocol", .{
.root_source_file = b.path("system/services/display/protocol.zig"),
});
runtime_module.addImport("display-protocol", display_protocol_module);
// The power protocol: system power's domain-named surface (docs/power.md).
const power_protocol_module = b.addModule("power-protocol", .{
@@ -253,18 +382,6 @@ pub fn build(b: *std.Build) void {
},
});
// The POSIX / C compatibility layer, a separate library layered strictly over the
// runtime (it calls the runtime's IPC/heap, never system calls directly). This is
// the one place POSIX/C spellings are allowed verbatim — see docs/coding-standards.md
// and library/posix/posix.zig.
const posix_module = b.addModule("posix", .{
.root_source_file = b.path("library/posix/posix.zig"),
.imports = &.{
.{ .name = "runtime", .module = runtime_module },
.{ .name = "vfs-protocol", .module = vfs_protocol_module },
},
});
// The initial_ramdisk container format, shared by the kernel (unpacks it) and the
// build-time packer tools/make-initial-ramdisk.py (produces it). No dependencies.
const initial_ramdisk_module = b.addModule("initial-ramdisk", .{
@@ -274,9 +391,12 @@ pub fn build(b: *std.Build) void {
// Compile-time configuration the kernel reads as `@import("build_options")`. The
// QEMU test harness sets -Dtest-case=<name> to run one self-test at boot.
const test_case = b.option([]const u8, "test-case", "Kernel self-test case to run at boot (see system/kernel/tests.zig)");
const build_options = b.addOptions();
build_options.addOption(?[]const u8, "test_case", test_case);
const build_options_module = build_options.createModule();
// The serial-console log sink. Off by default: a real machine often has no
// working legacy COM1, and the boot log is kept in RAM (klog) and flushed to
// disk instead — serial is now only a QEMU convenience. `run-x86-64` and the
// QEMU test harness (test/qemu_test.py, which asserts on serial markers) turn
// it on; a flashable `zig build` image leaves it out. See serial.zig.
const serial = b.option(bool, "serial", "Compile the serial-console log sink into the kernel (default: off; run-x86-64 and the test harness enable it)") orelse false;
// --- Kernel: freestanding x86_64 ELF, jumped to by the bootloader ---
// SSE2 is part of the x86_64 baseline and UEFI leaves it enabled at handoff,
@@ -288,41 +408,17 @@ pub fn build(b: *std.Build) void {
.abi = .none,
});
const exe = b.addExecutable(.{
.name = "kernel",
.root_module = b.createModule(.{
.root_source_file = b.path("system/kernel/kernel.zig"),
.target = kernel_target,
.optimize = optimize,
.code_model = .kernel, // kernel runs in the top 2 GiB (higher half)
.red_zone = false, // interrupts would corrupt the SystemV red zone
.single_threaded = false, // SMP: the big kernel lock's atomics must be real across cores
.sanitize_c = .off, // the UBSan runtime needs f128/SSE support we don't provide
.stack_check = false, // stack-probe calls have no runtime to land in
.stack_protector = false,
.imports = &.{
.{ .name = "boot-handoff", .module = boot_handoff_module },
.{ .name = "abi", .module = abi_module },
.{ .name = "device-abi", .module = device_abi_module },
.{ .name = "architecture", .module = architecture_module },
.{ .name = "platform", .module = platform_module },
.{ .name = "parameters", .module = parameters_module },
.{ .name = "build_options", .module = build_options_module },
.{ .name = "initial-ramdisk", .module = initial_ramdisk_module },
},
}),
});
exe.setLinkerScript(b.path("system/kernel/architecture/x86_64/linker.ld"));
exe.entry = .{ .symbol_name = "_start" };
// The self-hosted linker ignores parts of the linker script (PHDRS,
// /DISCARD/, AT(), section order); the higher-half layout depends on the
// script being authoritative, so pin the kernel to LLVM + LLD.
exe.use_llvm = true;
exe.use_lld = true;
// Higher-half virtual base (matches KERNEL_VIRT_BASE in linker.ld); the
// linker's AT() clauses give each segment a low physical load address
// (.text at 1 MiB), which the loader allocates and copies into.
exe.image_base = 0xFFFFFFFF80100000;
const kernel_modules = KernelModules{
.boot_handoff = boot_handoff_module,
.abi = abi_module,
.device_abi = device_abi_module,
.architecture = architecture_module,
.platform = platform_module,
.parameters = parameters_module,
.initial_ramdisk = initial_ramdisk_module,
};
// The installed/flashable kernel: serial follows -Dserial (off by default).
const exe = addKernel(b, kernel_target, optimize, kernel_modules, test_case, serial);
// Everything installs into a FHS-shaped zig-out: it IS the danos filesystem *and*
// the boot volume. Each binary lands at its addressed, leaf-collapsed path — the
@@ -336,7 +432,7 @@ pub fn build(b: *std.Build) void {
// Built by the shared user-binary recipe (see addUserBinary): freestanding,
// linked into the kernel's user region against the `runtime` runtime library, and
// started in ring 3 by the kernel's user-ELF loader.
const init_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "init", "system/services/init/init.zig");
const init_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "init", "system/services/init/init.zig");
const init_install = b.addInstallArtifact(init_exe, .{ .dest_dir = .{ .override = .{ .custom = "system/services" } } });
b.getInstallStep().dependOn(&init_install.step);
@@ -344,20 +440,43 @@ pub fn build(b: *std.Build) void {
// Each is built by the same user-binary recipe, then packed into one image by
// the host-side make-initial-ramdisk tool. The bootloader ferries the image to the kernel,
// which unpacks it and spawns each program (system/initial-ramdisk.zig).
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig");
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig");
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
// The xHCI bus driver builds chapter-9 requests and decodes descriptors from
// usb-abi, and reports each interface's (class,subclass,protocol) identity via
// usb-ids.packTriple.
usb_xhci_bus_exe.root_module.addImport("usb-abi", usb_abi_module);
usb_xhci_bus_exe.root_module.addImport("usb-ids", usb_ids_module);
usb_xhci_bus_exe.root_module.addImport("usb-transfer-protocol", usb_transfer_protocol_module);
// The USB HID class drivers: keyboard and mouse. They own no hardware — each
// opens its device through runtime.usb (the transfer protocol) and publishes to
// the input service. They build chapter-9 class requests from usb-abi.
const usb_hid_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-keyboard", "system/drivers/usb-hid/keyboard.zig");
usb_hid_keyboard_exe.root_module.addImport("usb-abi", usb_abi_module);
const usb_hid_mouse_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-mouse", "system/drivers/usb-hid/mouse.zig");
usb_hid_mouse_exe.root_module.addImport("usb-abi", usb_abi_module);
// The USB mass-storage class driver: opens its device via runtime.usb, drives it
// with Bulk-Only Transport + SCSI, and serves the block protocol under `.block`.
const usb_storage_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-storage", "system/drivers/usb-storage/usb-storage.zig");
usb_storage_exe.root_module.addImport("block-protocol", block_protocol_module);
// The FAT filesystem server: mounts the block device and serves it into the VFS
// at /mnt/usb. Its engine (engine.zig / on-disk.zig) is imported relatively.
const fat_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat", "system/services/fat/fat.zig");
const display_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display", "system/services/display/display.zig");
const display_demo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display-demo", "system/services/display-demo/display-demo.zig");
const fat_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat-test", "system/services/fat/fat-test.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
// The PCI bus driver decodes each function's class triple to human names in its
// boot log (class/subclass/prog-IF), so pull in the shared pci-class reference.
pci_bus_exe.root_module.addImport("pci-class", pci_class_module);
// A test fixture, not a real driver: hellos to the device manager, then faults —
// what the driver-restart scenario drives the crash-loop cap with.
const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig");
const device_list_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig");
const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig");
const device_list_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig");
// The discovery service: one swappable process per firmware
// (docs/discovery.md), bundled under the neutral ramdisk name
// "discovery" so the device manager never learns which firmware it is on.
@@ -371,18 +490,22 @@ pub fn build(b: *std.Build) void {
.acpi => "system/services/acpi/acpi.zig",
.fdt => "system/services/fdt/fdt.zig",
};
const discovery_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source);
const discovery_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source);
if (discovery == .acpi) discovery_exe.root_module.addImport("aml", aml_module);
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// Names the xHCI PCI class triple from the shared taxonomy instead of a bare 0x0C0330.
device_manager_exe.root_module.addImport("pci-class", pci_class_module);
// The manager matches reported USB interfaces by their (class,subclass,protocol)
// triple (usbDriverForIdentity), built from the named usb-ids codes.
device_manager_exe.root_module.addImport("usb-ids", usb_ids_module);
// The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
const input_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig");
const input_source_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-source", "system/services/input-source/input-source.zig");
const input_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig");
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig");
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig");
const input_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig");
const input_source_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-source", "system/services/input-source/input-source.zig");
const input_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig");
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig");
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig");
const log_flush_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "log-flush", "system/services/log-flush/log-flush.zig");
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool
// (the container format is trivial, and Python sidesteps std API churn). Args:
@@ -402,6 +525,20 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(ps2_mouse_exe.getEmittedBin());
mk_run.addArg("usb-xhci-bus");
mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin());
mk_run.addArg("usb-hid-keyboard");
mk_run.addFileArg(usb_hid_keyboard_exe.getEmittedBin());
mk_run.addArg("usb-hid-mouse");
mk_run.addFileArg(usb_hid_mouse_exe.getEmittedBin());
mk_run.addArg("usb-storage");
mk_run.addFileArg(usb_storage_exe.getEmittedBin());
mk_run.addArg("fat");
mk_run.addFileArg(fat_exe.getEmittedBin());
mk_run.addArg("fat-test");
mk_run.addFileArg(fat_test_exe.getEmittedBin());
mk_run.addArg("display");
mk_run.addFileArg(display_exe.getEmittedBin());
mk_run.addArg("display-demo");
mk_run.addFileArg(display_demo_exe.getEmittedBin());
mk_run.addArg("pci-bus");
mk_run.addFileArg(pci_bus_exe.getEmittedBin());
mk_run.addArg("crash-test");
@@ -422,6 +559,8 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(args_echo_exe.getEmittedBin());
mk_run.addArg("process-test");
mk_run.addFileArg(process_test_exe.getEmittedBin());
mk_run.addArg("log-flush");
mk_run.addFileArg(log_flush_exe.getEmittedBin());
// Also install the packed binaries to their FHS homes, so zig-out is a true image
// of the filesystem — even though at boot they arrive inside the initial-ramdisk.
@@ -433,6 +572,12 @@ pub fn build(b: *std.Build) void {
.{ ps2_keyboard_exe, "system/drivers" },
.{ ps2_mouse_exe, "system/drivers" },
.{ usb_xhci_bus_exe, "system/drivers" },
.{ usb_hid_keyboard_exe, "system/drivers" },
.{ usb_hid_mouse_exe, "system/drivers" },
.{ usb_storage_exe, "system/drivers" },
.{ fat_exe, "system/services" },
.{ display_exe, "system/services" },
.{ log_flush_exe, "system/services" },
}) |entry| {
const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } });
b.getInstallStep().dependOn(&step.step);
@@ -445,6 +590,13 @@ pub fn build(b: *std.Build) void {
// Boot methods live in boot/, one per way of getting the kernel running.
// Each is its own binary/entry (a loader is built for its own target); today
// that's UEFI for x86-64, with room for e.g. a device-tree path for the Pis.
// The loader reads -Dserial too, so its boot-progress breadcrumbs (con_out,
// which firmware may mirror to a serial console) are silenced by default — a
// real-hardware boot stays quiet. Fatal-error messages ignore this and always
// show, so a failed boot still explains itself on screen. See boot/efi.zig.
const loader_options = b.addOptions();
loader_options.addOption(bool, "serial", serial);
const loader_options_module = loader_options.createModule();
const efiexe = b.addExecutable(.{
.name = "BOOTX64",
.root_module = b.createModule(.{
@@ -457,6 +609,7 @@ pub fn build(b: *std.Build) void {
.imports = &.{
// The bootloader speaks only the handoff contract — never the user ABI.
.{ .name = "boot-handoff", .module = boot_handoff_module },
.{ .name = "build_options", .module = loader_options_module },
},
}),
});
@@ -466,6 +619,32 @@ pub fn build(b: *std.Build) void {
const efi_install = b.addInstallArtifact(efiexe, .{ .dest_dir = .{ .override = .{ .custom = "EFI/BOOT" } } });
b.getInstallStep().dependOn(&efi_install.step);
// --- danos-usb.img: the bootable FAT32 USB image ---
// Format a real FAT32 image (the in-repo Python builder, no external tools)
// holding exactly what the firmware and bootloader need off the ESP: the EFI
// stub, the kernel, init, and the initial-ramdisk. QEMU presents this image as
// a USB mass-storage device the guest boots from (see run-x86-64 and the test
// harness), and the danos fat driver mounts the same image at /mnt/usb.
const fat_image = addBootImage(b, exe.getEmittedBin(), efiexe.getEmittedBin(), init_exe.getEmittedBin(), initial_ramdisk_img);
const fat_image_install = b.addInstallFile(fat_image, "danos-usb.img");
b.getInstallStep().dependOn(&fat_image_install.step);
// The image `run-x86-64` boots: identical to the flashable one but with the
// serial log sink compiled in, so a developer always gets the machine-readable
// log captured to serial0 — without baking serial into the image users flash.
// Built lazily (only when `run-x86-64` is requested), and never installed.
const exe_serial = addKernel(b, kernel_target, optimize, kernel_modules, test_case, true);
const fat_image_serial = addBootImage(b, exe_serial.getEmittedBin(), efiexe.getEmittedBin(), init_exe.getEmittedBin(), initial_ramdisk_img);
// `zig build check-fat-image` — validate the produced image is a real FAT32
// with the EFI stub present (the builder's own --verify, no external tools).
const check_fat = b.addSystemCommand(&.{"python3"});
check_fat.addFileArg(b.path("tools/make-fat-image.py"));
check_fat.addArg("--verify");
check_fat.addFileArg(fat_image);
const check_fat_step = b.step("check-fat-image", "Verify the FAT32 USB image is valid and bootable");
check_fat_step.dependOn(&check_fat.step);
// --- run-x86-64: boot the x86-64 kernel in QEMU via UEFI/OVMF ---
// Firmware lives in different places per OS/distro, so probe the known
// layouts (Architecture, Debian/Ubuntu, Fedora, macOS Homebrew) and use the first
@@ -526,10 +705,14 @@ pub fn build(b: *std.Build) void {
});
run_efi.addArg("-drive");
run_efi.addPrefixedFileArg("if=pflash,format=raw,file=", vars_out);
// Present the FHS zig-out to the guest as a FAT drive — it is the boot volume.
// Boot off the FAT32 USB image: a mass-storage device on the same xHCI bus as
// the keyboard and mouse. OVMF finds \EFI\BOOT\BOOTX64.efi on it and boots.
// The serial-enabled variant, so serial0 carries the log for this dev boot.
run_efi.addArg("-drive");
run_efi.addPrefixedFileArg("if=none,id=bootusb,format=raw,file=", fat_image_serial);
run_efi.addArgs(&.{
"-drive",
b.fmt("format=raw,file=fat:rw:{s}", .{b.install_path}),
"-device",
"usb-storage,bus=xhci.0,drive=bootusb,removable=on,bootindex=0",
"-net",
"none",
// Emulated display advertising 1280x720 as its native (EDID preferred)
@@ -548,8 +731,10 @@ pub fn build(b: *std.Build) void {
const make_log_dir = b.addSystemCommand(&.{ "mkdir", "-p", log_dir });
const serial_log = b.fmt("{s}/run-x86-64-serial0-{s}.log", .{ log_dir, timestamp(b) });
run_efi.addArgs(&.{ "-serial", b.fmt("file:{s}", .{serial_log}) });
// The whole FHS zig-out must be installed (and the scratch dir created) before we mount it.
run_efi.step.dependOn(b.getInstallStep());
// We boot the self-contained `fat_image_serial` (added as a file arg above, so
// it's already a dependency) — not the installed FHS zig-out — so `run-x86-64`
// builds only the serial kernel, never the flashable one. Just make the serial
// scratch dir first.
run_efi.step.dependOn(&make_log_dir.step);
const run_efi_step = b.step("run-x86-64", "Boot the x86-64 kernel in QEMU (UEFI/OVMF); serial0 is logged to zig-out/qemu-test/run-x86-64-serial0-<timestamp>.log");
@@ -581,6 +766,15 @@ pub fn build(b: *std.Build) void {
"library/mmio/mmio.zig", // barriers assemble + registers round-trip
"system/drivers/ps2-bus/scancode.zig", // set-2 decode + keyboard state machine
"system/drivers/ps2-bus/mouse-packet.zig", // 3-byte mouse packet assembly
"system/drivers/usb-hid/hid-report.zig", // HID boot-report keyboard/mouse decode
"system/drivers/usb-storage/bulk-only-transport.zig", // CBW/CSW wrapper sizes
"system/drivers/usb-storage/scsi.zig", // SCSI CDB encodings (big-endian)
"system/services/vfs/path.zig", // mount-prefix path matching
"system/services/vfs/protocol.zig", // NodeKind / DirectoryEntry sizes + op values
"system/services/fat/on-disk.zig", // FAT on-disk struct sizes + type detection
"system/services/fat/engine.zig", // FAT read/write over a RAM-backed image
"system/services/display/compositor.zig", // Rect math + fill/composite/blit-tile
"system/services/display/protocol.zig", // pack(): native pixel encoding per format
}) |root| {
const mod_tests = b.addTest(.{
.root_module = b.createModule(.{
+23 -12
View File
@@ -69,7 +69,12 @@ rather than restate it. Roughly in the order things happen at runtime:
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
service layered on top.
19. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
19. **[display.md](display.md) — the display service.** The display half of the GUI
track: a user-space compositor that owns the framebuffer, composes a layer stack into
a double buffer, and presents it. Why GOP and the PCI display device are two views of
one controller, the device-node + write-combining handoff, and what flicker-free buys
that tear-free doesn't. Plan: [display-plan.md](display-plan.md).
20. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Start with the north star:
@@ -82,6 +87,12 @@ Start with the north star:
- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on
fault isolation + live restart — the reincarnation-server + capability model that
makes "if I break it, I can restart it" real. danos's core motivation.
- **[zig-self-hosting.md](zig-self-hosting.md) — running Zig on danos.** A design note
(not built yet) on making danos a real Zig target (`-target x86_64-danos`) and
eventually running the compiler on it. The key realisation: Zig 0.16 reduces an OS
port to **one seam** (`std.os.danos`), so we build `runtime.os` (→ that seam) plus a
thin `runtime.fs`, retire the `posix` shim, and follow a phased path to
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
Cutting across all of these:
@@ -191,21 +202,22 @@ system/ → /system danos's own internals (the self-representation)
services/ init/ vfs/ device-manager/ system servers → /system/services (vfs/ holds
vfs.zig, vfs-test.zig, protocol.zig)
library/ → /lib libraries, one sub-directory each
runtime/ the danos-native runtime — the stable application ABI
posix/ POSIX/C compatibility, layered over runtime
runtime/ the danos-native runtime + file API (fs) — the stable application ABI
boot/ → /boot the loaders
tools/ test/ host-side build + QEMU test harness
```
A sub-project exposes its **public interface as a module**: `system/services/vfs/` owns
the VFS wire protocol (`protocol.zig`, the `vfs-protocol` module), which the POSIX
layer imports by name. `usb`/`block` drivers will expose their protocols the same way.
the VFS wire protocol (`protocol.zig`, the `vfs-protocol` module), which the runtime's
file API (`runtime.fs`) imports by name. `usb`/`block` drivers expose their protocols the
same way.
`library/posix/` is special: it is the **one place** POSIX/C spellings are allowed
verbatim (`stat`, `O_CREAT`, `fopen`, `errno`). Everywhere else follows the danos
naming rule with no exception — see [coding-standards.md](coding-standards.md). The
POSIX layer calls the runtime, never the kernel's system calls directly, so it never
appears in the private-ABI path.
There is **no POSIX/C compatibility layer today**: danos programs do file I/O through the
danos-native `runtime.fs` (open/read/write/list over the VFS). A hand-rolled POSIX shim
(`library/posix/`) was retired as premature — the real POSIX/C surface will come later
from the `std.os.danos` seam (and, eventually, musl) when danos becomes a Zig target (see
[zig-self-hosting.md](zig-self-hosting.md)). When it does, the foreign-ABI naming
exception in [coding-standards.md](coding-standards.md) applies to that seam.
## Source map
@@ -229,8 +241,7 @@ appears in the private-ABI path.
| Framebuffer text console (mirrors to serial) | `system/kernel/console.zig` |
| In-kernel test cases | `system/kernel/tests.zig` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/IO-APIC/timer, serial, linker script) | `system/kernel/architecture/x86_64/` |
| danos-native runtime (`runtime`): syscall wrappers, heap, IPC, device access — the stable application ABI | `library/runtime/` |
| POSIX/C compatibility (`posix`): unistd, stdio — the one place POSIX names are allowed | `library/posix/` |
| danos-native runtime (`runtime`): syscall wrappers, heap, IPC, device access, the file API (`fs`) — the stable application ABI | `library/runtime/` |
| System services (init, the VFS server + `protocol`, the device-manager) | `system/services/` |
| Device drivers, one sub-project each (`pci-bus`, `ps2-bus`, `usb-xhci-bus` bus drivers) | `system/drivers/` |
| Build + `run-x86-64` (QEMU/OVMF) | `build.zig` |
+12 -9
View File
@@ -65,15 +65,18 @@ Three, and only three.
`errno`, `O_CREAT`. We don't get to rename `fwrite` to `fileWrite` — it wouldn't be
`fwrite` any more.
**This exception is scoped to one place: `library/posix/`.** A file under
`library/posix/` *is* the foreign ABI, so it keeps the ABI's spellings — that is the
whole rule for that directory. **Everywhere else, Zig/danos naming applies with no
POSIX exception**, so there is nothing to get wrong: if you're not in
`library/posix/`, expand it. A concept POSIX also has gets a danos name outside that
layer — the VFS wire protocol carries a `FileStatus`, not a `Stat`, and a `create`
flag, not `O_CREAT`; `library/posix/` is what maps `stat`→`status` and
`O_CREAT`→`create` at the boundary. (The `syscall` *wrappers* elsewhere are not an
exception to this — they wrap the private danos ABI, so they use danos names.)
**This exception is scoped to a file that *is* a foreign ABI, and nothing else.**
danos has no such file today: the old `library/posix/` compatibility shim was retired
once its callers moved to the danos-native `runtime.fs`, since a hand-rolled POSIX
layer is premature until danos actually needs it (see
[zig-self-hosting.md](zig-self-hosting.md)). The exception will apply again to the
`std.os.danos` seam when danos becomes a real Zig target — that module *is* the C-ABI
`system` interface, so it keeps `open`/`read`/`errno`/`O_CREAT`. **Everywhere else,
Zig/danos naming applies with no exception**: a concept POSIX also has gets a danos
name — the VFS wire protocol carries a `FileStatus`, not a `Stat`, and a `create`
flag, not `O_CREAT`; the boundary is where `stat`→`status` and `O_CREAT`→`create` get
mapped. (The `syscall` *wrappers* elsewhere are not an exception — they wrap the
private danos ABI, so they use danos names.)
2. **Zig idioms are spelled the way Zig spells them.** Three names are the language's,
not ours, and are left alone:
+181
View File
@@ -0,0 +1,181 @@
# Display service — build plan (v1: the dumb-framebuffer compositor)
The ordered, checkpointable build-out for [display.md](display.md). Each milestone is
small, lands on its own, and ends in a **verifiable gate** — shaped for a `/loop` run.
Read [display.md](display.md) first for the *why*; this is the *what* and the *order*.
## Locked decisions (do not relitigate)
- **Handoff = device node + write-combining `mmio_map`.** The kernel seeds a synthetic
`display0` node from `BootInformation.framebuffer`; the service claims + WC-maps it.
(Not a bespoke `framebuffer_map` syscall — the device route inherits ownership,
release-on-death, and re-claim-on-restart.)
- **v1 = the full compositor pipeline on the dumb framebuffer.** One `display` service
owns the LFB + a cacheable back buffer + a layer stack; double-buffer + damage-driven
present; clients draw via server-side commands. **No** runtime mode-setting, **no**
shared-memory surfaces — both deferred (see display.md, "What v1 does not do").
## Conventions
Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations in
full, kebab-case file names, no `Co-Authored-By` trailers on commits. New user binaries
go through `addUserBinary` in [build.zig](../build.zig) and get packed into the
initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported into the
`runtime` module.
## How to verify along the way
- `zig build test` — host unit tests (compositor math: layer clipping, damage merge,
pitch/format blits are all host-testable with a fake framebuffer).
- `python3 test/qemu_test.py <case>` — boots the real kernel in QEMU; assert on the
serial log ([tests.zig](../system/kernel/tests.zig) is the registry).
- The `run-efi` target renders to QEMU's display (`-device VGA,edid=on,xres=1280,yres=720`)
— a screenshot confirms pixels for the milestones whose gate is visual.
---
## D1 — The handoff primitive (kernel) ✅
Make the boot framebuffer reachable and mappable **write-combining** from user space.
- [x] [device-abi.zig](../system/devices/device-abi.zig): added `DeviceClass.display`; a
`DisplayInfo{ width, height, pitch, format }` carried on the descriptor; a
`flags` field on `ResourceDescriptor` + `resource_flag_write_combining`.
- [x] [devices-broker.zig](../system/kernel/devices-broker.zig): `seedDisplay(base, w, h,
pitch, format)` publishes a root-level `display` node with one WC-flagged `memory`
resource `[base, height*pitch]` + the `DisplayInfo`; `displayDevice()` /
`displayClaimed()`. Seeded from `kmain` after `devices_broker.init`.
- [x] [process.zig](../system/kernel/process.zig) `systemMmioMap` + paging
(`mapUserDeviceInto` gains a `write_combining` bool): a resource's WC flag maps it
through the WC PAT slot (`setupPat`) instead of strong-uncacheable.
- [x] [console.zig](../system/kernel/console.zig): `setSuppressed` quiesces `write` while
the display device is claimed (driven from `systemDeviceClaim` / release); the
terminal panic + exception paths clear it first so a dying machine still draws.
**Gate (met, automated):** the `display` kernel test (`python3 test/qemu_test.py display`,
`displayTest` in [tests.zig](../system/kernel/tests.zig)) asserts the seeded node's shape
and geometry, then walks the real claim + `mmio_map` path into a throwaway address space
and verifies the leaf is **write-combining** (PAT entry 4: PAT bit set, PCD/PWT clear) —
with an uncacheable-still-uncacheable regression guard. Chosen over the original
screenshot-of-a-fill gate because it proves the *actual* WC property headlessly; the
visible fill folds into D2's gate (the service clears the screen through the back buffer).
Regression-checked: `discovery`, `ioport`, `claim-release`, `supervision`, `device-list`,
`device-manager` all still pass with the +1 device in the table.
## D2 — Service skeleton, protocol, runtime module ✅
Stand up the named service and the double-buffer, no layers yet.
- [x] `system/services/display/protocol.zig`: `Operation{ info, create_layer,
configure_layer, destroy_layer, fill_rect, blit_tile, damage, present }`; `extern`
`Request`/`Reply`; size + `maximum_payload` consts. (Model: block/protocol.zig.)
- [x] [abi.zig](../system/abi.zig): `ServiceId.display = 9`.
- [x] `system/services/display/display.zig`: `main` → enumerate + claim + WC-map the LFB
(front) → `mmap` a cacheable back buffer of `height*pitch` → `runtime.service.run`.
`info` and a whole-screen `present` (back → front) are live; layer ops fail-stub
until D3. Init clears the back buffer and presents it — the double-buffer path.
- [x] [library/runtime/display.zig](../library/runtime/runtime.zig) (+ barrel export of
`display` and `display_protocol`): `info()` and `present()`, cached `.display`
lookup with retry (model: block.zig).
- [x] [init.zig](../system/services/init/init.zig): `"display"` added to `boot_services`.
- [x] [build.zig](../build.zig): `display-protocol` module on the runtime; `display` exe
via `addUserBinary`; packed into the initial-ramdisk; installed to
`/system/services/display`.
- [x] **Kernel fix the back buffer surfaced:** `mmap` was capped at 256 pages (1 MiB) by
a fixed kernel-stack `frames` array. Rewrote `systemMmap` to map page-by-page with
rollback (no scratch array) and raised the cap to 8192 pages (32 MiB) — enough for a
4K back buffer. A real limitation met, exactly the kind this project chases.
**Gate (met, automated):** `python3 test/qemu_test.py display-service` spawns the
compositor and matches its own serial heartbeats — `display: online {w}x{h} pitch …`
followed by `display: presented frame 0` — which it prints only after the whole
claim → WC-map → back-buffer → clear → present chain succeeds (matched on serial like the
fault cases, since a lone blocking service can't reschedule the in-kernel test context to
poll). Regression-checked: `usermem`, `heap` (the `mmap` rewrite), `init` (the boot-list
addition), and D1's `display` all still pass.
## D3 — Layer stack + compositor + damage present ✅
The heart: composite an ordered layer stack, present only what changed.
- [x] A layer table (16 slots): each `Layer` = position, z, visible, a server-owned
`mmap`'d surface (freed on `destroy_layer`). `damage` accumulates the dirty screen
region since the last present.
- [x] `create_layer` / `configure_layer` (damages old + new footprints) / `destroy_layer`,
`fill_rect`, `blit_tile` (reads the inline tile from the IPC payload, unaligned-safe),
`damage`, `present`.
- [x] Pure, host-tested [compositor.zig](../system/services/display/compositor.zig): `Rect`
(intersect/unite), `Surface`, `fillRect`, `composite` (opaque, clipped to a damage
rect), `blitTile`. `present` clears the damaged region to the wallpaper, paints the
visible layers bottom-to-top (z-sorted), and flushes just that rect back → front (WC).
Colour packing (rgbx/bgrx) is `protocol.pack`, also host-tested.
- [x] Host tests (`zig build test`, green): rect intersect/unite, `fillRect` clipping +
`stride > width` padding, `composite` overlap-shows-top + damage clipping, `blitTile`
unaligned read + clipping, and `pack` for both pixel formats.
**Gate (met):** `zig build test` green for the compositor + pack unit tests, **and** the
`display-service` case's startup self-check composites two overlapping layers on the real
framebuffer and reads back the composited pixels — overlap = top layer, outside = bottom
layer — logging `display: compositor self-check ok` (matched by the harness).
## D4 — Client API + the demo client ✅
Prove the pipeline end-to-end from a separate process.
- [x] Finished [runtime/display.zig](../library/runtime/runtime.zig): a `Layer` handle with
`fill` / `blitTile` (inline tile) / `configure` (move/restack/show) / `damage` /
`destroy`, `createLayer`, and a `color(r,g,b)` helper (caches the mode, packs via
`protocol.pack`). Coordinates are signed over the wire (`@bitCast` both ways).
- [x] `system/services/display-demo/`: a hardware-free client (the `input-source` analog)
— a full-screen wallpaper layer, a rectangle that slides back and forth (moved by
`configure` each frame, so the compositor repaints old + new), and a cursor layer;
presents in a loop paced by `runtime.time`. Wired into build + initial-ramdisk.
- [x] **Bug this surfaced:** `protocol.message_maximum` was 4096, but the kernel caps
every IPC message at `MESSAGE_MAXIMUM` = 256 — so `replyWait` rejected the oversized
receive buffer with `-E2BIG` and the serve loop had been *spinning* since D2 (unseen,
as D2/D3 matched init-time heartbeats). Set it to 256; `blit_tile` is now explicitly
a small-tile path (≤ 54 px inline), larger bitmaps being the deferred shm surface.
**Gate (met):** `python3 test/qemu_test.py display-demo` spawns the service + `display-demo`;
the demo drives a run of frames of motion through the layer client API and logs
`display-demo: ok` (the visible motion is a screenshot via `zig build run-x86-64`).
Regression-checked: `zig build test`, `display` (D1), and `display-service` (D2/D3) all
still pass, and the default `zig build` is clean.
## D5 — Test cases + docs ✅
- [x] The three integration cases exist and pass: `display` (D1 handoff, kernel),
`display-service` (D2/D3 compositor + self-check), and `display-demo` (D4 full
pipeline: spawn `display` + `display-demo`, match `display-demo: ok`) —
[tests.zig](../system/kernel/tests.zig) + [qemu_test.py](../test/qemu_test.py). Plus
the pure host tests (`zig build test`).
- [x] [display.md](display.md) updated to the built state (the "Verifying it" section names
the real cases); [README index](README.md) entry present (#19); the `display-track`
memory marked DONE with the commits.
**Gate (met):** `python3 test/qemu_test.py display display-service display-demo` all pass,
`zig build test` is green, and the default `zig build` is clean.
---
## v1 status: complete
D1–D5 done. The display service is a working framebuffer compositor: it owns the
framebuffer (write-combining), composites a z-ordered layer stack into a cacheable back
buffer, presents only the damaged region, and is driven over IPC by the `runtime.display`
client — proven end-to-end by a separate demo process. Two limitations are deliberate and
documented (docs/display.md): no runtime mode-setting (native backend) and no true vsync
(no vblank on a dumb framebuffer). Next steps are the Deferred items below.
---
## Deferred (explicitly not in this plan)
- **Shared-memory surfaces** — generalize M13 capability passing to memory objects
(`shm_create`/`shm_map`), so bitmap clients hand the compositor a rendered surface
instead of drawing commands. The compositor's layer model already anticipates it.
- **Native backend (Bochs DISPI, then virtio-gpu)** — behind the same internal backend
interface as the dumb framebuffer: EDID mode list + runtime resolution/bpp change +
(eventually) a vblank/flip path for true vsync.
- **Driver/compositor process split** — only when a second backend or a second head makes
the abstraction pay for itself.
+264
View File
@@ -0,0 +1,264 @@
# The display service: a framebuffer compositor
The [framebuffer](framebuffer.md) the loader hands over is a flat block of pixel
memory, and the kernel's [bootstrap console](../system/kernel/console.zig) draws text
into it directly. That console is a stop-gap. The **display service**
(`system/services/display/`) is the real thing: an ordinary ring-3 process that *owns*
the framebuffer, composes a stack of **layers** into an off-screen back buffer, and
**presents** finished frames to the screen — the display half of the GUI track
([vision.md](vision.md)), the sibling of the [input service](input.md).
This note is the architecture and the reasoning behind it. The concrete build order
lives in [display-plan.md](display-plan.md).
## First, a distinction that shapes everything: GOP vs. the PCI device
It is tempting to think "the GOP framebuffer" and "the VGA-compatible display
controller in the PCIe tree" are two different things. They are not — they are **two
interfaces to the same silicon, at different times and different levels**, and knowing
which one you're holding decides what you can do.
- **GOP is firmware's *temporary* driver** for the display controller. It gives you a
linear framebuffer pointer and can set video modes — but only until
`ExitBootServices`. The loader already leans on this: [`queryFramebuffer`](../boot/efi.zig)
reads the monitor's EDID, picks the native mode, and calls `set_mode` **before**
exiting ([gop.md](gop.md)). Once the kernel runs, GOP is **gone** — no `set_mode`, no
mode list, no EDID. What survives is the frozen snapshot in
[`BootInformation.framebuffer`](../system/boot-handoff.zig): `{base, width, height,
pitch, format}`, and nothing more.
- **The PCI class-0x03 device is the raw controller** — BARs, config space, registers,
IO ports. It is what you actually *own* after boot. On QEMU's emulated adapter
([`-device VGA,edid=on`](../build.zig), the Bochs VBE/DISPI model) the `base` GOP handed
you *is* that device's linear-framebuffer BAR — the same physical memory, seen through
a different door. On a real discrete GPU, GOP's `base` is an aperture inside the GPU's
VRAM BAR. danos already decodes this device
([pci-class.zig](../system/devices/pci-class.zig) has the full `display` namespace, and
`pci-bus` already reports it to the [device manager](device-manager.md) with its class
triple) — but nothing binds it yet.
What that difference costs you, concretely:
| You want to… | Dumb GOP framebuffer (boot handoff) | Native device driver (PCI 0x03) |
|-------------------------------------------|-------------------------------------|------------------------------------------|
| **Report** the current mode | ✅ from the handoff | ✅ |
| **Change resolution / bpp at runtime** | ❌ GOP is gone | ✅ program DISPI regs / virtio-gpu queue |
| **Re-read EDID, enumerate monitor modes** | ❌ | ✅ the device exposes an EDID block |
| **Refresh rate** | ❌ (virtual anyway) | only a real KMS driver — far future |
| **vblank / tear-free present** | ❌ no vblank signal | ✅ vblank IRQ + page-flip (real GPUs) |
| **Works on the Pi (no PCI VGA)** | ✅ VideoCore hands a simple FB | ✗ per-device |
The lesson: the **portable base for the whole GUI stack is the GOP / boot-handoff linear
framebuffer**. Runtime mode-setting is a *per-device upgrade* layered on top — and on
the Raspberry Pis there is no PCI VGA at all, so the neutral framebuffer is the only
thing all three target machines share. That is why the display service is built on the
dumb framebuffer first, with the native backend as an optional module behind the same
interface.
## Two constraints this service exists to meet
Like the input service — which existed partly to motivate the asynchronous
[`ipc_send`](ipc.md) primitive — the display service runs straight into two limits the
rest of the system hasn't had to face:
1. **The framebuffer is kernel-only today.** It arrives through the boot handoff, is
mapped into the kernel's physmap, and is touched only by
[`console.zig`](../system/kernel/console.zig). It is *not* a
[devices-broker](../system/kernel/devices-broker.zig) node, so
`device.claim`/`mmio_map` cannot reach it, and there is no framebuffer
[syscall](syscall.md). A user-space display service needs a **new mechanism just to
touch the pixels**. (See "The handoff" below — this is built.)
2. **danos has no cross-process shared memory.** The memory syscalls are `mmap`
(private, zeroed), `mmio_map` (a *claimed device's* MMIO), and `dma_alloc` (new
pinned physical). The block driver's "pass a buffer by physical address" trick
([block/protocol.zig](../system/services/block/protocol.zig)) works *only because its
consumer is DMA hardware*. A compositor that CPU-reads and blends client layers can't
use it — it would have to *map* another process's memory, which nothing allows. This
is deferred (see "What v1 does not do"), because v1 sidesteps it entirely.
## Architecture
```
kernel ── owns the boot framebuffer; bootstrap console only
│ seeds a "display0" device node from BootInformation.framebuffer
│ (ResourceKind.memory = [base, height*pitch], write-combining hint,
│ plus DisplayInfo{width, height, pitch, format})
▼
display service (system/services/display/, ServiceId.display) ← the compositor
│ device.claim(display0) → mmio_map(WRITE-COMBINING) = FRONT buffer (the LFB)
│ mmap(cacheable) a BACK buffer of the same geometry
│ owns: an ordered LAYER STACK + a per-frame DAMAGE list
│ loop: composite dirty layers → back buffer → present dirty rects → front
│ backend is an INTERNAL interface: {gop-fb} today; {bochs-dispi, virtio-gpu} later
▼ reached by name (ipc_lookup); clients drive it over the display protocol
┌────────────────────────────────────┬──────────────────────────────────────┐
drawing clients (v1) surface clients (deferred)
runtime.display commands: runtime.display surfaces:
create_layer / configure_layer shm_create → pass as a capability →
fill_rect / blit_tile / damage the compositor maps & composites the
present client-rendered bitmap directly
```
The bring-up sequence mirrors a hardware driver's — it is the
[`usb-xhci-bus` `initialise`](../system/drivers/usb-xhci-bus/usb-xhci-bus.zig) shape
(claim → `mmio_map` → run loop) — and the request/reply service shell is the
[FAT](../system/services/fat/fat.zig) / [input](../system/services/input/input.zig) shape
([`runtime.service.run`](../library/runtime/service.zig) with a `protocol.zig` of
`extern struct` messages and an `Operation` tag).
**One process, for now.** v1 is a *single* service that both owns the framebuffer and
composites — it does not split a "framebuffer driver" from a "compositor" the way input
splits `ps2-bus` from the input service. The backend (dumb FB vs. a native GPU) is an
*internal* interface, not a process boundary. That boundary earns its keep only when a
second backend or a second monitor appears; until then it is complexity with no payoff.
## The handoff: a device node + a write-combining map
The framebuffer crosses into user space through the machinery that already exists for
every other device, rather than a bespoke syscall — so it inherits ownership,
release-on-death, and re-claim-on-restart for free (the [resilience](resilience.md)
story: a crashed display service returns the LFB to the kernel, and its restart
re-claims it).
- The kernel seeds a synthetic **`display0`** node into the
[devices-broker](../system/kernel/devices-broker.zig) at init, from
`BootInformation.framebuffer`: one `ResourceKind.memory` resource spanning
`[base, height*pitch]`, tagged **write-combining**, plus a small
`DisplayInfo{width, height, pitch, format}` (the memory resource says *where* and *how
big*; `DisplayInfo` says how to *interpret* the bytes).
- The service `device.claim`s it and `mmio_map`s the resource. The map is
**write-combining**, not the strong-uncacheable that `mmio_map` uses for register
MMIO. The kernel already programs a WC PAT slot for its own console
([`setupPat`](../system/kernel/architecture/x86_64/paging.zig)); this reaches it from
the user mapping path. **This matters:** an uncacheable framebuffer makes the
back→front blit unusably slow.
- On `claim`, the kernel's bootstrap console goes quiet, so the two never fight over the
LFB. A panic is the one exception — by then the service is likely dead anyway, and a
panic on screen wins.
The display service is a **named boot service**: `init` spawns it by name alongside
`vfs`/`input`/`device-manager` ([init.zig](../system/services/init/init.zig)), and it
self-discovers `display0` with `device.enumerate`. The [device manager](device-manager.md)
matching path (PCI class 0x03 → a driver) is reserved for the future *native* backend, not
this singleton synthetic node.
## Double buffering and the write-combining discipline
Two buffers, with deliberately different memory types:
- The **front buffer** is the LFB — **write-combining**: fast to *write*, slow to
*read*. The rule is therefore **never read the front buffer**. Only ever stream into
it, sequentially.
- The **back buffer** is ordinary **cacheable** RAM (`mmap`), the same geometry. All
compositing happens here, where reads and read-modify-write blends are cheap.
So a frame is: compose every dirty layer into the cacheable back buffer, then **present**
— copy the changed regions back→front in sequential, WC-friendly writes. Two details the
[framebuffer](framebuffer.md) note already establishes carry over: step rows by `pitch`,
not `width*4`; and handle both `rgbx` and `bgrx` [pixel formats](gop.md).
## Flicker vs. tearing — what double buffering does and doesn't buy
These are two different artifacts, and the dumb framebuffer fixes exactly one of them:
- **Flicker** is the user seeing intermediate, half-drawn states (a clear-then-redraw
flash). Double buffering **eliminates it completely** — the screen only ever receives
whole, finished frames.
- **Tearing** is a present landing while the display's scanout beam is mid-frame, so the
top of the screen shows the new frame and the bottom the old. Avoiding it requires
presenting during the vertical blank (**vsync**) — which needs a vblank signal. **A
dumb GOP framebuffer has no vblank.**
So v1 is **flicker-free**, and it *minimizes* the tear window by presenting only damaged
rectangles (less to copy → a smaller window in which the beam can catch a half-updated
frame), but it is **not tear-free**. Genuine vsync waits for a backend with a vblank IRQ
or a flush/flip path — a native-device capability, not something the firmware
framebuffer can offer. Stated plainly here so the limitation is understood, not
discovered.
## Layers and the client protocol
The compositor holds an **ordered stack of layers**. Each layer has a rectangle, a
z-order, a visibility flag, and a surface. Presenting walks the stack bottom-to-top,
painting each dirty layer into the back buffer, then flushes the damage to the front.
In v1 the surfaces are **server-owned**, and clients draw into them with a small
immediate-mode command protocol — essentially the model early X used, and enough for a
shell, a terminal, a cursor, and a wallpaper:
| Operation | Meaning |
|--------------------|---------------------------------------------------------------|
| `info` | report `{width, height, pitch, format}` of the display |
| `create_layer` | allocate a server-owned surface, return a layer handle |
| `configure_layer` | set a layer's rect, z-order, visibility |
| `destroy_layer` | release a layer |
| `fill_rect` | fill a rectangle of a layer with a colour |
| `blit_tile` | copy a small client-supplied pixel tile into a layer (inline) |
| `damage` | mark a region of a layer dirty |
| `present` | composite dirty layers and flush to the screen |
Text is intentionally *not* an operation — a client renders glyphs by blitting tiles
(the [PSF font](../system/kernel/font.psf) path the console already uses can move into a
client). Keeping the protocol to rectangles and tiles keeps the compositor small and the
policy in the client.
## `runtime.display`
Clients speak the protocol through a new [`library/runtime/display.zig`](../library/runtime/runtime.zig),
the [`runtime.block`](../library/runtime/block.zig) shape (a cached `.display` lookup
with a boot-race retry): `display.info()`, a `Layer` handle with `fill` / `blitTile` /
`damage`, and `present()`. Application code never issues the raw syscalls — it calls the
runtime, as with every other danos service.
## What v1 does not do (and why that's fine)
Two capabilities are deliberately out of the first cut. Neither reshapes anything above;
both are clean additions behind the interfaces v1 establishes.
- **Client-rendered surfaces (shared memory).** The fast path for a bitmap-heavy app is
to render into its *own* buffer and hand the compositor a *reference*, not a stream of
commands. That needs the missing cross-process shared-memory primitive — best built as
the natural generalization of the existing M13 [capability passing](driver-model.md)
from *endpoints* to *memory objects* (`shm_create(len) → {cap, vaddr}`, pass `cap` on
an `ipc_call`, receiver `shm_map(cap) → vaddr`). v1 avoids it because server-owned
surfaces already prove the whole pipeline.
- **Runtime mode-setting (a native backend).** Detecting the EDID mode list and changing
resolution / bpp at runtime needs the raw PCI device. The first native backend is
Bochs DISPI — the register interface QEMU's `-device VGA` exposes — behind the same
internal backend interface the dumb framebuffer sits behind. Refresh-rate and colour
management (a gamma LUT) are real-GPU-KMS territory, far beyond this.
## Verifying it
Three QEMU test cases ([tests.zig](../system/kernel/tests.zig), `python3
test/qemu_test.py <case>`), each layering on the last:
- **`display`** — the kernel handoff: the seeded `display` device is shaped correctly and
the claim → `mmio_map` leaf is genuinely **write-combining** (PAT entry 4), asserted at
the page-table level.
- **`display-service`** — the compositor comes up: it claims the framebuffer, allocates
the cacheable back buffer, presents a cleared frame through the double-buffer path
(`display: online … / presented frame 0`), and a startup **self-check** composites two
overlapping layers on the real framebuffer and reads them back — overlap = the top
layer — logging `display: compositor self-check ok`.
- **`display-demo`** — the full pipeline from a separate process: the hardware-free
[`display-demo`](../system/services/display-demo/) client (the
[`input-source`](../system/services/input-source/) analog) drives layers — a wallpaper, a
sliding rectangle, a cursor — through the layer client API and heartbeats
`display-demo: ok`, proving a frame travelled client → compositor → screen, exactly as
the [input test](input.md) proves an event travels source → service → subscriber. The
visible motion itself is a screenshot away via `zig build run-x86-64`.
The compositor's pixel math (rectangle clipping, fill, composite, tile blit) and colour
packing are additionally covered by pure host unit tests under `zig build test`.
## See also
- [framebuffer.md](framebuffer.md) — the linear framebuffer, pitch vs. width, `volatile`.
- [gop.md](gop.md) — GOP, and why only linear RGBX/BGRX modes are paintable.
- [input.md](input.md) — the sibling service; the async `ipc_send` fan-out.
- [driver-model.md](driver-model.md) — claim / `mmio_map`, capability passing, the trust model.
- [device-manager.md](device-manager.md) — matching and supervision (the native backend's route).
- [display-plan.md](display-plan.md) — the ordered build-out.
+6 -6
View File
@@ -17,12 +17,12 @@ was a mistake) without inheriting the mechanism, the API, or the names. The nami
rule is danos's own and it is strict: plain words that communicate intent
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
a concept that already has one. Literal POSIX arrives later and lives elsewhere: a
**musl-based C layer** (growing out of library/posix) that wires C programs to the
danos runtime — musl's syscall surface retargeted at danos system calls and IPC
protocols (files onto the VFS protocol, `sigaction`/`wait` onto this lifecycle,
sockets onto whatever networking becomes). Ported programs see POSIX; the system
underneath never does.
a concept that already has one. Literal POSIX arrives later and lives elsewhere: the
`std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C
layer** on the same native surface (see [zig-self-hosting.md](zig-self-hosting.md)) —
musl's syscall surface retargeted at danos system calls and IPC protocols (files onto
the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever
networking becomes). Ported programs see POSIX; the system underneath never does.
## Why a standard vocabulary
+9
View File
@@ -27,6 +27,15 @@ transcript. Serial is per-architecture (x86 uses port I/O; an ARM board uses a
memory-mapped UART), so it lives behind the [arch](arch.md) boundary — and adding
a new architecture's UART is what makes the same tests run there.
The serial log sink is **compiled in only under `-Dserial`** (off by default).
A real machine often has no live legacy COM1 — writing to a dead one is slow —
and the boot log is kept in a RAM buffer (`klog`) and flushed to disk instead,
so serial is now purely a QEMU/dev aid. The harness (`test/qemu_test.py`) builds
every case with `-Dserial=true`, and `zig build run-x86-64` boots a serial-enabled
image variant, so both get the transcript; a flashable `zig build` image leaves
serial out. (Even with `-Dserial`, a loopback probe disables a dead port at boot,
so a serial-enabled image is still safe on real hardware.)
## In-kernel test cases
Building with `-Dtest-case=<name>` makes the kernel, after normal bring-up, run one
+349
View File
@@ -0,0 +1,349 @@
# Running Zig on danos: the self-hosting roadmap
A design note (not built yet) on the path to making danos a **real Zig target** — a
target you can name (`-target x86_64-danos`) and, eventually, run the Zig compiler
itself on. It is forward-looking, like [vision.md](vision.md): it sets a direction
and the decisions that follow from it, so the code we write now bends toward it
instead of away.
This note deliberately does **not** cover a text editor or terminal. Those are
easier (single-process, I/O-bound) and fall out of the early phases here almost for
free; the hard, shaping problem is the standard-library surface, so that is what
this roadmap is about.
The analysis behind it was done against **Zig 0.16** (the pinned toolchain). Zig's
standard library moves between releases — especially the parts described here — so
treat upstream references as "the shape in 0.16.x," and expect to re-check them on a
toolchain bump.
## The win condition
danos runs the Zig compiler when a bare
```
zig build-exe hello.zig
```
completes **on danos** and produces a runnable danos binary. Note the milestone is
`build-exe`, not `zig build`: the `zig build` runner spawns child processes (the
build steps), which needs a whole process-control surface danos does not have yet.
A single `build-exe` needs none of that (see Phase 3). Reaching `build-exe` is
"self-hosting"; reaching `zig build` is a later, separate lift.
### Non-goals
- **No Linux syscall/ABI emulation.** danos will not implement the Linux `syscall`
interface so that stock `x86_64-linux` binaries run. That is a permanent
compatibility treadmill and it inverts the microkernel design — explicitly out.
- **No musl port yet.** A musl libc port is a reasonable *later* effort (it unlocks
the C ecosystem), but it is not on the critical path to Zig-on-danos, and it is
deferred. The roadmap below is arranged so the work still pays off if musl ever
happens (see "The same surface, twice").
- **Editor/terminal are out of scope for this note** (they are downstream of Phase 1).
**On FFI.** Foreign-function interop splits the same way as the doors below. Zig-level
and C-ABI-*exposing* FFI (`extern`, `callconv(.c)`, C-ABI structs) work on a real target
immediately — and the `std.os.danos` seam is C-ABI-shaped by construction, so it is
FFI-friendly from the start. *Consuming* C libraries (`@cImport`, linking archives) is
the part that needs a libc + headers, i.e. the deferred musl door. So an eventual FFI
need reinforces keeping that door open; it does not change the plan.
## The realization that shapes everything: 0.16 gives us *one* seam
The instinct "to target Zig we'd have to reimplement all the `std` namespaces" was
how older Zig worked. Zig 0.16 (post-"writergate") is far kinder:
- **`std.fs` is essentially gone.** It is now path helpers plus deprecated aliases;
there is no `std.fs.File`, `std.fs.Dir`, or `std.fs.cwd()`. File and directory
work goes through **`std.Io`** — a single runtime **vtable** (`Io.zig`) of
function pointers handed to `main` as `std.process.Init.io`. `std.Io.File` and
`std.Io.Dir` are thin forwarders to that vtable. `Io.zig` and the `fs` shim carry
**zero** per-OS branches.
- **`std.posix` is one generic body** parameterised over a single `system` module.
With no libc, `system` resolves **per target OS**: `.linux => std.os.linux`,
`.plan9 => std.os.plan9`, and so on. The generic `std.posix.read`/`write`/`open`
bodies are just `system.read(...)` plus an errno switch — *identical for every
OS*. The only variable is what `system` binds to.
- **`std.os.<tag>`** (e.g. `std/os/linux.zig`) is therefore the real porting seam: a
low-level, C-ABI-shaped module of `read/write/open/close/lseek/mmap/clock/exit/…`
plus an `errno` enum and the constant tables (`O_*`, `CLOCK_*`, `S_*`).
Put together: **to port danos we write `std.os.danos` once** — the ~30-operation
seam — and the whole `std.posix` / `std.fs` / `std.Io` tower above it lights up
generically, because none of it branches on the OS. That is a dramatically smaller
and more contained target than "reimplement the namespaces."
## Three doors, and why we take the first
| Door | What it is | Verdict |
|------|-----------|---------|
| **1. Implement the std seam** (`std.os.danos`) | Write the ~30-op `system` module over danos's native ABI + VFS; the generic std tower lights up. | **Take this.** The only door that touches neither C nor the Linux ABI. |
| **2. Port musl** | Port musl libc to danos, link Zig against it. | Defer. Good later for the *C* ecosystem; barely helps *Zig* (std only uses libc on the libc-linked path). |
| **3. Emulate the Linux ABI** | Implement Linux syscalls so stock linux binaries run. | Reject. Bottomless compatibility treadmill; against the design. |
### The same surface, twice
Doors 1 and 2 are the **same native surface at different layers**. `std.posix.read`
is `system.read(...)` + an errno switch *regardless of OS* — the only question is
whether `system` is **`std.os.danos` (Zig)** or **musl (C)**. Either way, the set of
danos-facing operations you must implement is the *same* ~30 ops, all bottoming out
in danos's native syscalls + the VFS/FAT server.
So the runtime work below is **not throwaway** if musl ever happens: you are building
the danos-native implementations of that surface either way. Door 1 just packages
them as Zig; a future musl re-uses the identical kernel/VFS operations underneath. The
two symmetries worth keeping in mind: doors 1 and 2 converge at the **top** (identical
POSIX surface); doors 2 and 3 converge at the **bottom** (unmodified musl needs the
Linux syscall ABI). Door 1 is the only one that avoids both C and Linux.
### A fork is table stakes — for any door
`std.Target.Os.Tag` is a **closed enum** baked into the compiler binary *and* into
the `std` linked with every program; `-target x86_64-danos` resolves through it. So
adding `danos` as a name requires patching and rebuilding the compiler — even the
musl door needs this. "Fork Zig" is therefore not an extra cost unique to door 1; it
is the price of admission for *any* real target. What door 1 adds on top is small and
localised (below).
## The architecture decision: `runtime.os` + `runtime.fs`, and retire `posix`
danos already has the right split ([the private-ABI boundary](../README.md)): the
kernel exposes a minimal syscall ABI ([syscall.md](syscall.md)); the **`runtime`**
library is the stable, danos-native application ABI. What this roadmap adds:
- **`runtime.os` — the seam.** A C-ABI-shaped module of the ~30 operations
(`read/write/open/close/lseek/mmap/munmap/clock/exit/…`) + an errno enum + the
constant tables, each backed by danos's native syscalls and the VFS. **Structure it
to mirror `std/os/linux.zig`.** This is the load-bearing, *non-throwaway* artifact:
when we fork Zig, `runtime.os` is copy-pasted (near-verbatim) into `std.os.danos`.
- **`runtime.fs` — the thin native file API** danos programs use *today*, layered
over `runtime.os`. It is also the concrete backing for the `std.Io` vtable's
file-write entry once we're a real target, which is why program stdout, diagnostics,
and file writes should all be *decided once at that seam* rather than as bespoke
per-call helpers (see "How this informs decisions now").
**Do not hand-mirror the high-level std namespaces.** `std.fs`/`std.Io`/`std.process`
are generic and OS-agnostic; once `std.os.danos` exists and we fork, upstream *gives*
them to danos for free. Hand-writing `runtime.std.fs` to imitate them would be
redundant the day the fork works, and it would chase a moving target (0.16's `std.Io`
is large and still shifting). Build the seam well; take the tower for free.
**Why not a library called `std`?** Because `@import("std")` resolves to the
compiler-provided standard library; a user module named `std` would *shadow* it for
anything that imports it that way. That is the real reason the seam lives *inside* a
forked std as `std/os/danos.zig`, not as a `runtime.std` library — and why danos's end
state (`@import("std")` just working, and knowing danos) is the most natively Zig it can
be. `runtime.os` is only the interim staging ground: developed against the stock
toolchain so Phase 1 need not wait on the fork, then promoted near-verbatim into the
fork's `std/os/danos.zig`.
### Retire `library/posix`
The `posix` compatibility layer (`unistd`, `stdio`) was the right instinct too early.
Its whole value is POSIX *spellings* for POSIX software — and danos has no POSIX
software; every current caller is danos-native code that could use `runtime.fs`
directly. The real POSIX story arrives later and from elsewhere (musl, or upstream
`std`'s own posix over `std.os.danos`), which supersedes a hand-rolled shim. So it is
premature abstraction that adds a "which layer do I use?" fork with no payoff yet.
Its footprint is tiny: **five** call sites, all `unistd` file operations —
`system/services/fat/fat.zig` (`mount`), the `vfs-test` and `fat-test` clients, and
(from the boot-log work) `init.zig` and `log-flush.zig`. `stdio.zig` is dead — nothing
imports it. The plan: build `runtime.fs`, migrate those five to it, delete
`library/posix/`, and drop the `posix` module from `build.zig`'s `addUserBinary`.
## Where danos stands: coverage vs. the gaps
What the seam needs, and what danos already provides:
| std need | danos today | Gap |
|----------|-------------|-----|
| open / read / write / close / lseek | VFS (via the current `unistd`, → `runtime.fs`) | none — repackage |
| directory read (`getdents`) | VFS `readdir` | none — repackage |
| mmap / munmap | native syscalls ([abi.zig](../system/abi.zig)) | none |
| page allocator | over `mmap`, via `root.os.heap.page_allocator` override | ~30-line hook |
| monotonic clock | `clock` syscall | none |
| args / argv | SysV entry stack ([sysv.md](sysv.md)), `runtime.process.Init` | none |
| stdout / stderr | `debug_write` today | wire fd 1/2 to a console **byte** stream |
| mkdir / unlink / rename / truncate | done — engine + VFS + `runtime.fs` (Phase 2) | — |
| stat fields | `{size, kind, mtime}` | **mode / inode** still missing (cache validity) |
| wall-clock / realtime | done — `wall_clock` syscall (CMOS RTC, Phase 2d) | — |
| **environment variables** | `Init` has no env field | missing (can start empty) |
| **cwd / chdir** | paths are absolute or bare | missing (no cwd anchor) |
| **entropy / random** | — | missing (needed behind `vtable.random`) |
| process spawn + exit status | `system_spawn` starts a *named ramdisk binary*; `ExitReason` is a *category* | no exec-of-path, no numeric `WEXITSTATUS` |
| threads | one thread per process | avoided via `-fsingle-threaded` (below) |
| symlinks | `NodeKind` has the tag; unimplemented | low priority |
The clustering is clear: reads and memory are basically done; the real work is
**filesystem mutation + richer stat + wall-clock**, and a few small seam pieces
(page-allocator hook, stdio bytes, entropy). Process spawning and threads are
side-stepped entirely for a single `build-exe`.
## The roadmap
### Phase 0 — Make `danos` a real target
**Host, target, self-host — keep the three roles straight.** The *host* is where the
compiler runs (your mac + linux dev machines); the *target* is what it emits (`danos`);
and eventually danos becomes a host too (self-hosting — the win condition). So the move
is: fork the compiler, build it **for** your dev hosts, and teach it to **cross-compile
to** danos. You already do this — danos is cross-compiled `freestanding` from your dev
host today; Phase 0 swaps that `freestanding` target for a real `x86_64-danos` one, which
is what unlocks the native `std`.
**Why a compiler fork, not just a `--zig-lib-dir` override.** `std.Target.Os.Tag` is a
*closed enum compiled into the compiler binary*, so `-target x86_64-danos` will not even
parse unless the compiler itself knows the tag. Overriding the std lib directory alone
cannot add a target — and there is no libc-only shortcut (a future musl needs the same
patch). The only alternative, staying on `freestanding` + hand-shims, is exactly the
non-native feel we are leaving: `@import("std")` there is stubbed, not real.
**The fork.** Clone `ziglang/zig` at the pinned 0.16 tag; build it with a stock
same-version `zig` (`zig build` in the tree — a standard, LLVM-pulling, roughly one-time
build); point danos's `build.zig`/CI at the resulting binary. Four localised patches:
- add `danos` to `std.Target.Os.Tag`, in the "no version range" group alongside
plan9/serenity;
- add `danos` to the freestanding/other **no-op `_start` list** in `std`'s `start.zig`,
so std does *not* emit its own System-V `_start` — danos keeps owning the entry shim
and `Init`/argv construction it already builds ([sysv.md](sysv.md));
- wire the `system` selector `.danos => std.os.danos` in `std.posix`;
- add `std/os/danos.zig` — **the seam itself**, promoted near-verbatim from the
`runtime.os` developed first in Phase 1 (against the stock toolchain, so the fork is
not a prerequisite for starting).
This is the fork treadmill we accept once. Keep the patch set tiny and `else`-friendly,
pin to one 0.16.x, and rebase on point releases.
### Phase 1 — `runtime.os` read-side + allocator + stdio + cwd; retire `posix`
Author `runtime.os` (→ `std.os.danos`): the `errno` enum, the constant tables, and
the C-convention `read / write / open / openat / close / lseek / mmap / munmap /
exit`, each returning result-or-`-errno`. Most backing already exists (VFS + native
mmap + clock).
- Provide `page_allocator` via `root.os.heap.page_allocator` (a thin override over
danos `mmap`). This sits **outside** the `std.Io` vtable, so it is wired separately.
- Wire fd 0/1/2 to a console **byte** stream (today output only reaches `debug_write`;
input is structured `InputEvent` IPC — a byte tty is a new, small thing in both
directions).
- Add a `getcwd`/`chdir` anchor so `std.fs.cwd()`-style resolution has something to
resolve against.
- Build `runtime.fs` over `runtime.os`; migrate the five `posix` callers to it; delete
`library/posix/` and drop its build module.
After Phase 1, the surface an editor or terminal needs (open/read/write/close/lseek/
readdir/isatty/args/exit) exists. Those are downstream and out of scope here.
### Phase 2 — Filesystem mutation + real stat (the compiler's cache tower)
danos's biggest genuine gap, and the correctness-critical one:
- Add **mkdir / unlink / rename / truncate** to *both* the VFS wire protocol
([protocol.zig](../system/services/vfs/protocol.zig)) and the FAT engine
([engine.zig](../system/services/fat/engine.zig)), then expose them via `runtime.os`.
- Extend `stat` beyond `{size, kind}` to carry **mtime + inode + mode** — `std`'s file
stat needs them for build-cache validity — which in turn needs **wall-clock** time
(danos is monotonic-only today; an RTC/time service is the dependency).
Because `std.fs`/`std.Io` have no per-OS branches, finishing this in `runtime.os`
lights up the whole file tower for the compiler at once. Environment can stay an empty
map until the kernel populates a non-empty `envp`.
**Status — Phase 2 complete.** `truncate` (O_TRUNC, closing the boot-log stale-tail
bug), `mkdir`, `unlink`, and `rename` are all wired through the FAT engine, the VFS
protocol + router, and `runtime.fs` (`makeDirectory` / `remove` / `rename`) —
host-tested and QEMU-tested (`fat-mutations` + `fat-rename` make a directory, write+read
a file in it, rename it, then remove it through the mount). `removeFile` and `rename`
are LFN-aware; `rename` is same-directory + 8.3 (cross-directory and long-name-
preserving rename are noted limitations). Wall-clock is now a kernel syscall
(`wall_clock`, a CMOS-RTC read anchored to the monotonic clock), and the FAT engine
stamps and reports **mtime** — `stat` / `runtime.fs.Attributes` carry a real
modification time (the `fat-mtime` case reads it back within seconds of the host clock).
The remaining `stat` fields, `mode`/`inode`, are deferred (not needed until the
compiler's cache layer wants them). **Everything past here is gated on Phase 0 (the
fork):** the `runtime.os` seam, `cwd`, stdio-as-fds, and the compiler bring-up.
### Phase 3 — Single-threaded, self-linked compiler bring-up
Build the compiler with **two load-bearing flags**:
- **`-fsingle-threaded`** removes `std.Thread` entirely — `Thread.spawn` is a hard
compile error under it, and `std.Io`'s threaded backend runs inline. danos being
one-thread-per-process is therefore **not** a blocker. Parallel codegen is a
throughput optimisation, not a correctness requirement.
- **`-fno-llvm -fno-lld`** keeps codegen and linking **in-process** (the self-hosted
x86-64 backend + self-linker), so a single `build-exe` **never forks a child**. That
is what lets us defer the entire spawn/exec/wait surface.
Then supply the few remaining seam pieces: `now` (wrap the danos clock), an entropy
source behind `vtable.random` (`randomSecure` can alias it initially — low volume, for
temp-file names and hashmap seeds), and the Phase-2 mkdir/rename/unlink for cache dir
trees and atomic temp-then-rename output.
**Explicitly deferred** (not on the `build-exe` path): child-process spawn/exec (only
`zig build` and external tools need it), `std.Thread`, `fsync` (FAT is write-through
today), symlinks, and musl.
## Risks and gotchas
- **The std-fork rebase treadmill is the main ongoing cost.** A new OS tag touches the
same broad file set plan9/serenity touch (hundreds of `native_os` sites, plus
"unsupported OS" `@compileError` dead-ends a new tag must be routed around), and the
entire `std.Io` layer is new in 0.16 and still moving. Stay pinned to one 0.16.x,
keep additions localised and `else`-friendly. Watch the closed-enum gotcha: adding
`danos` to `Os.Tag` can break existing *exhaustive* switches that lack an `else`, so
expect to touch switch sites beyond the ones you implement.
- **Single-threaded is load-bearing.** The "no `std.Thread`" simplification rests
entirely on `-fsingle-threaded`. If a dependency or flag flips threading back on, you
inherit an unescapable compile error (no root-hook exists) — the only outs are a full
thread-impl fork or linking libc for pthreads. Keep `single_threaded` asserted end to
end.
- **In-process linking is load-bearing.** Reaching the compiler without fork/exec
depends on `-fno-llvm -fno-lld`. The moment you shell out to LLD/`ld`, you need the
full `spawn`/`wait` surface — the hardest microkernel piece — and danos's
`system_spawn` only starts a *named ramdisk binary*, not exec of an arbitrary path.
Verify the self-hosted backend covers the target output before assuming child
processes are optional.
- **The shim cannot host the compiler.** danos's current `runtime`/`posix` is fine for
danos's *own* native programs, but the compiler `import`s *upstream* `std`, which on
a non-target hits the void `system` stub. So the compiler forces the real target
(Phase 0's fork). Do not over-invest in extending the hand-shim for compiler
purposes; put that effort into `runtime.os` + the VFS/FAT operations, which both the
fork *and* a future musl consume.
- **`"w"`/`O_CREAT` does not truncate — a silent-corruption bug on this road.** The FAT
engine's `writeFile` only *grows* `node.size`, so overwriting a shorter file leaves
trailing garbage. Harmless for the boot log today, but for a compiler it means
**corrupt `.o`/cache files that look like nondeterministic compiler bugs.** Land
`truncate` (Phase 2) before the compiler ever writes cache.
- **Exit status is categorical, not numeric.** `process_exit_reason` returns an
`ExitReason` *category*, not a numeric code (`WEXITSTATUS`). Fine while spawn is
stubbed; the day `zig build` or external tools arrive, plan a kernel exit-record
extension — do not let it surprise you.
## How this informs decisions now
Two current decisions fall out of this roadmap:
1. **The `runtime.fs` / `std.Io` question resolves at the vtable seam.** Because 0.16
routes *all* output through the `std.Io` vtable's file-write entry, and stdout/stderr
are just `File`s with well-known handles, build `runtime.fs` (and the console stdout)
as the concrete backing for that entry — not as a bespoke `std.Io.Writer`-only shim.
Decide it once, at the seam, and program stdout, diagnostics, and file writes all
flow through the same danos VFS/console path.
2. **The boot-log `truncate` caveat is now fixed** (Phase 2a). It was the same
`writeFile`-only-grows gap that on the self-hosting road would corrupt build output;
`engine.truncate` + an O_TRUNC open flag now free the old chain so a shorter rewrite
leaves no stale tail, and the boot-log flush opens with it.
## Related
- [vision.md](vision.md) — the north star this serves.
- [syscall.md](syscall.md) — the kernel↔runtime ABI `runtime.os` is built on.
- [sysv.md](sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs.
- [ipc.md](ipc.md) — the IPC the VFS/FAT operations travel over.
- [danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md) — the
filesystem layout the file surface serves.
- [coding-standards.md](coding-standards.md) — danos naming (why the compat spellings
are confined, and now retired).
-13
View File
@@ -1,13 +0,0 @@
//! DanOS's POSIX / C compatibility layer — `unistd`, `stdio`, and (later) the C
//! `errno` / `struct stat` / `extern "C"` surface. This is the *one* place POSIX and
//! C spellings are allowed to appear verbatim (see docs/coding-standards.md): a file
//! under library/posix/ *is* the foreign ABI, so it keeps the ABI's names. Everything
//! it touches on the danos side (the VFS protocol, the runtime) uses danos names,
//! which this layer translates to at the boundary.
//!
//! It is layered strictly *over* the runtime: it calls the runtime's IPC and heap,
//! never the kernel's system calls directly. danos-native applications use the
//! runtime; this exists so *POSIX* software can too.
pub const unistd = @import("unistd.zig");
pub const stdio = @import("stdio.zig");
-115
View File
@@ -1,115 +0,0 @@
//! A small C stdio layer over the POSIX-style file API (unistd.zig). Unbuffered
//! for now — each fread/fwrite is one VFS round trip; an internal buffer (fewer
//! IPC calls) is a later optimisation. Both a Zig-callable API and `extern "C"`
//! symbols are provided, so Zig and future C programs share it.
const std = @import("std");
const unistd = @import("unistd.zig");
const heap = @import("runtime").heap;
pub const SEEK_SET = unistd.SEEK_SET;
pub const SEEK_CURRENT = unistd.SEEK_CURRENT;
pub const SEEK_END = unistd.SEEK_END;
/// A C `FILE`: an fd plus sticky end-of-file / error flags. Allocated on the
/// heap; `fclose` frees it.
pub const FILE = extern struct {
fd: i32,
eof: c_int = 0,
err: c_int = 0,
};
fn flagsFor(mode: []const u8) u32 {
if (mode.len == 0) return 0;
return switch (mode[0]) {
'w', 'a' => unistd.O_CREAT,
else => 0,
};
}
/// Open `path` in `mode` ("r"/"w"/"a", '+' ignored for now). Returns null on error.
pub fn fopen(path: []const u8, mode: []const u8) ?*FILE {
const fd = unistd.open(path, flagsFor(mode));
if (fd < 0) return null;
const f = heap.allocator().create(FILE) catch {
unistd.close(fd);
return null;
};
f.* = .{ .fd = fd };
if (mode.len > 0 and mode[0] == 'a') _ = unistd.lseek(fd, 0, unistd.SEEK_END);
return f;
}
pub fn fclose(f: *FILE) c_int {
unistd.close(f.fd);
heap.allocator().destroy(f);
return 0;
}
/// Read `size*nmemb` bytes; returns the number of whole items read.
pub fn fread(buffer: []u8, size: usize, nmemb: usize, f: *FILE) usize {
const total = size * nmemb;
if (total == 0) return 0;
const n = unistd.read(f.fd, buffer[0..@min(buffer.len, total)]);
if (n <= 0) {
f.eof = 1;
return 0;
}
return @as(usize, @intCast(n)) / size;
}
/// Write `size*nmemb` bytes; returns the number of whole items written.
pub fn fwrite(data: []const u8, size: usize, nmemb: usize, f: *FILE) usize {
const total = @min(data.len, size * nmemb);
if (total == 0) return 0;
const n = unistd.write(f.fd, data[0..total]);
if (n <= 0) {
f.err = 1;
return 0;
}
return @as(usize, @intCast(n)) / size;
}
pub fn fseek(f: *FILE, off: i64, whence: u32) c_int {
f.eof = 0;
return if (unistd.lseek(f.fd, off, whence) < 0) -1 else 0;
}
pub fn ftell(f: *FILE) i64 {
return unistd.lseek(f.fd, 0, unistd.SEEK_CURRENT);
}
pub fn rewind(f: *FILE) void {
_ = fseek(f, 0, SEEK_SET);
}
pub fn feof(f: *FILE) c_int {
return f.eof;
}
pub fn ferror(f: *FILE) c_int {
return f.err;
}
pub fn fputs(s: []const u8, f: *FILE) c_int {
return if (unistd.write(f.fd, s) < 0) -1 else 0;
}
pub fn fputc(c: u8, f: *FILE) c_int {
const b = [_]u8{c};
return if (unistd.write(f.fd, &b) == 1) c else -1;
}
pub fn fgetc(f: *FILE) c_int {
var b: [1]u8 = undefined;
const n = unistd.read(f.fd, &b);
if (n <= 0) {
f.eof = 1;
return -1; // EOF
}
return b[0];
}
// Real `extern "C"` symbols (fopen/fread/fseek/...) — with a C-string signature
// distinct from the Zig slice API above — land with the first C program, wired
// via @export so they don't collide with these Zig names.
-148
View File
@@ -1,148 +0,0 @@
//! POSIX-style file API for user programs — the low level under C stdio. Files
//! are named objects served by the user-space VFS server (system/services/vfs/vfs.zig); each
//! call marshals a request, IPC_Calls the VFS, and unmarshals the reply. The
//! kernel knows nothing of files or fds — the fd table lives here, per process.
const std = @import("std");
const protocol = @import("vfs-protocol");
const ipc = @import("runtime").ipc;
pub const O_CREAT = protocol.create;
pub const SEEK_SET: u32 = 0;
pub const SEEK_CURRENT: u32 = 1;
pub const SEEK_END: u32 = 2;
// Resolve (and cache) the VFS server endpoint, looked up by well-known id.
var vfs_handle: usize = 0;
var vfs_resolved = false;
fn vfs() ?usize {
if (!vfs_resolved) {
vfs_handle = ipc.lookup(.vfs) orelse return null;
vfs_resolved = true;
}
return vfs_handle;
}
const maximum_fds = 32;
const Fd = struct { used: bool = false, node: u64 = 0, offset: u64 = 0 };
var fds = [_]Fd{.{}} ** maximum_fds;
fn allocFd() ?usize {
for (&fds, 0..) |*f, i| {
if (!f.used) {
f.* = .{ .used = true };
return i;
}
}
return null;
}
const Result = struct { reply: protocol.Reply, payload: []u8 };
/// One request/reply round trip: [Request header][send payload] -> VFS ->
/// [Reply header][receive payload]. The receive payload is written into `out`.
fn transact(request: protocol.Request, send: []const u8, out: []u8) ?Result {
const h = vfs() orelse return null;
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const slen = @min(send.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..slen], send[0..slen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(h, message[0 .. protocol.request_size + slen], &rbuf) catch return null;
if (n < protocol.reply_size) return null;
const reply = std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]);
const rpl = @min(n - protocol.reply_size, out.len);
@memcpy(out[0..rpl], rbuf[protocol.reply_size..][0..rpl]);
return .{ .reply = reply, .payload = out[0..rpl] };
}
/// Open (or create, with O_CREAT) `path`; returns an fd or -1.
pub fn open(path: []const u8, flags: u32) i32 {
const fd = allocFd() orelse return -1;
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = flags };
const r = transact(request, path, &.{}) orelse {
fds[fd].used = false;
return -1;
};
if (r.reply.status != 0) {
fds[fd].used = false;
return -1;
}
fds[fd] = .{ .used = true, .node = r.reply.node, .offset = 0 };
return @intCast(fd);
}
fn fdPtr(fd: i32) ?*Fd {
if (fd < 0 or fd >= maximum_fds) return null;
const f = &fds[@intCast(fd)];
return if (f.used) f else null;
}
/// Read up to `buffer.len` bytes at the current offset; returns the count or -1.
pub fn read(fd: i32, buffer: []u8) isize {
const f = fdPtr(fd) orelse return -1;
const want: u32 = @intCast(@min(buffer.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .read, .node = f.node, .offset = f.offset, .len = want, .flags = 0 };
const r = transact(request, &.{}, buffer) orelse return -1;
if (r.reply.status != 0) return -1;
f.offset += r.reply.len;
return @intCast(r.reply.len);
}
/// Write `data` at the current offset; returns the count or -1.
pub fn write(fd: i32, data: []const u8) isize {
const f = fdPtr(fd) orelse return -1;
const want: u32 = @intCast(@min(data.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .write, .node = f.node, .offset = f.offset, .len = want, .flags = 0 };
const r = transact(request, data[0..want], &.{}) orelse return -1;
if (r.reply.status != 0) return -1;
f.offset += r.reply.len;
return @intCast(r.reply.len);
}
/// Reposition the fd's offset. Returns the new offset or -1. (SEEK_END needs the
/// file size, which `stat` provides; handled by fetching it here.)
pub fn lseek(fd: i32, off: i64, whence: u32) i64 {
const f = fdPtr(fd) orelse return -1;
const base: i64 = switch (whence) {
SEEK_SET => 0,
SEEK_CURRENT => @intCast(f.offset),
SEEK_END => blk: {
const request = protocol.Request{ .operation = .status, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
var sbuf: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &sbuf) orelse return -1;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return -1;
const st = std.mem.bytesToValue(protocol.FileStatus, sbuf[0..@sizeOf(protocol.FileStatus)]);
break :blk @intCast(st.size);
},
else => return -1,
};
const pos = base + off;
if (pos < 0) return -1;
f.offset = @intCast(pos);
return pos;
}
/// Stat `path`. Returns 0 or -1.
pub fn stat(path: []const u8, out: *protocol.FileStatus) i32 {
// Open, stat by node, close — simple and enough for now.
const fd = open(path, 0);
if (fd < 0) return -1;
defer close(fd);
const f = fdPtr(fd).?;
const request = protocol.Request{ .operation = .status, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
var sbuf: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &sbuf) orelse return -1;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return -1;
out.* = std.mem.bytesToValue(protocol.FileStatus, sbuf[0..@sizeOf(protocol.FileStatus)]);
return 0;
}
/// Close an fd (best effort — tells the VFS to release the open file).
pub fn close(fd: i32) void {
const f = fdPtr(fd) orelse return;
const request = protocol.Request{ .operation = .close, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
_ = transact(request, &.{}, &.{});
f.used = false;
}
+69
View File
@@ -0,0 +1,69 @@
//! Block-device client: the helper a filesystem uses to read and write a block
//! device (a USB stick, via usb-storage) without hand-rolling the block-protocol
//! IPC. Layered over `ipc` and the shared `block-protocol` wire format, like
//! `runtime.usb` over the transfer protocol.
//!
//! Transfers name a caller-owned DMA buffer by physical address (from
//! `runtime.dma.alloc`), so whole sectors move without crossing the IPC size
//! limit — the same handoff usb-storage uses toward the controller.
const std = @import("std");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
const protocol = @import("block-protocol");
pub const Geometry = struct { block_size: u32, block_count: u64 };
pub const Device = struct {
endpoint: ipc.Handle,
/// The device's block size and total block count.
pub fn geometry(self: Device) ?Geometry {
var request = protocol.Request{ .operation = @intFromEnum(protocol.Operation.geometry), .lba = 0, .count = 0, .physical = 0 };
var reply: [protocol.reply_size]u8 = undefined;
const n = ipc.call(self.endpoint, std.mem.asBytes(&request), &reply) catch return null;
if (n < protocol.reply_size) return null;
const result = std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]);
if (result.status != 0) return null;
return .{ .block_size = result.block_size, .block_count = result.block_count };
}
/// Read `count` blocks starting at `lba` into the DMA buffer at `physical`.
pub fn read(self: Device, lba: u64, count: u32, physical: u64) bool {
return self.transfer(.read, lba, count, physical);
}
/// Write `count` blocks starting at `lba` from the DMA buffer at `physical`.
pub fn write(self: Device, lba: u64, count: u32, physical: u64) bool {
return self.transfer(.write, lba, count, physical);
}
/// Commit any device write cache to stable media (SCSI SYNCHRONIZE CACHE), so
/// prior writes survive a power-off. A filesystem calls this before the machine
/// goes down; no data transfer, so the buffer arguments are unused.
pub fn flush(self: Device) bool {
return self.transfer(.flush, 0, 0, 0);
}
fn transfer(self: Device, operation: protocol.Operation, lba: u64, count: u32, physical: u64) bool {
var request = protocol.Request{ .operation = @intFromEnum(operation), .lba = lba, .count = count, .physical = physical };
var reply: [protocol.reply_size]u8 = undefined;
const n = ipc.call(self.endpoint, std.mem.asBytes(&request), &reply) catch return false;
if (n < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]).status == 0;
}
};
/// Look up the block device, retrying generously while the USB storage chain
/// (controller reset, enumeration, mass-storage bring-up) comes up.
pub fn open() ?Device {
// Patient: the whole USB storage chain (firmware discovery, xHCI reset and
// enumeration, mass-storage bring-up) must complete first, which can take
// tens of seconds under emulation.
var attempts: usize = 0;
while (attempts < 1200) : (attempts += 1) {
if (ipc.lookup(.block)) |handle| return .{ .endpoint = handle };
system.sleep(50);
}
return null;
}
+171
View File
@@ -0,0 +1,171 @@
//! User-space display client: talk to the display service (query the mode, and — from D3
//! — create layers, draw, and present) without hand-rolling the IPC. The `runtime.block`
//! shape: a cached `.display` lookup with a boot-race retry, then extern-struct request/
//! reply marshalling. See system/services/display/ and docs/display.md.
const std = @import("std");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
const protocol = @import("display-protocol");
/// The display's current mode, as `info()` reports it.
pub const Info = struct {
width: u32,
height: u32,
pitch: u32, // bytes per row (may exceed width*4; see docs/framebuffer.md)
format: u32, // a device-abi DisplayFormat value (0 = rgbx, 1 = bgrx)
};
/// The service endpoint, looked up once and cached.
var handle: ?ipc.Handle = null;
/// Look up the display service, retrying while it comes up (a client races its
/// registration at boot). Returns the endpoint, or null if it never appears.
fn service() ?ipc.Handle {
if (handle) |h| return h;
var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.display)) |h| {
handle = h;
return h;
}
system.sleep(50);
}
return null;
}
/// Send one request, receive its reply; true on a zero status. `out` receives the reply
/// so callers can read `info`/`layer` fields on success.
fn transact(request: protocol.Request, out: *protocol.Reply) bool {
const h = service() orelse return false;
var req = request;
var reply: [protocol.reply_size]u8 = undefined;
const len = ipc.call(h, std.mem.asBytes(&req), &reply) catch return false;
if (len < protocol.reply_size) return false;
out.* = std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]);
return out.status == 0;
}
/// The display's current mode, or null if the service never came up.
pub fn info() ?Info {
var reply: protocol.Reply = undefined;
if (!transact(.{ .operation = @intFromEnum(protocol.Operation.info) }, &reply)) return null;
return .{ .width = reply.width, .height = reply.height, .pitch = reply.pitch, .format = reply.format };
}
/// Composite the dirty layers and flush the frame to the screen.
pub fn present() bool {
var reply: protocol.Reply = undefined;
return transact(.{ .operation = @intFromEnum(protocol.Operation.present) }, &reply);
}
/// The mode, cached after the first `info()` so `color()` doesn't round-trip per pixel.
var mode: ?Info = null;
fn cachedInfo() ?Info {
if (mode) |m| return m;
const i = info() orelse return null;
mode = i;
return i;
}
/// The native pixel value for an 8-bit-per-channel colour, in the display's format. A
/// client packs colours through this so it never has to know the byte order itself.
pub fn color(r: u8, g: u8, b: u8) u32 {
const format = if (cachedInfo()) |i| i.format else 0;
return protocol.pack(format, r, g, b);
}
/// A handle to a server-owned layer: a positioned, z-ordered surface the client draws
/// into by command. Create with `createLayer`; drawing and moves take effect on the next
/// `present`. Coordinates are signed (a layer may sit partly off-screen).
pub const Layer = struct {
id: u32,
/// Fill a rectangle of this layer (layer-local coordinates) with a native `colour`.
pub fn fill(self: Layer, x: i32, y: i32, w: u32, h: u32, colour: u32) bool {
var reply: protocol.Reply = undefined;
return transact(.{
.operation = @intFromEnum(protocol.Operation.fill_rect),
.layer = self.id,
.x = @bitCast(x),
.y = @bitCast(y),
.width = w,
.height = h,
.colour = colour,
}, &reply);
}
/// Copy a `w`×`h` tile of native pixels (row-major, little-endian bytes) into this
/// layer at (`x`, `y`). The tile rides inline in the request, so `w*h*4` must fit
/// `protocol.maximum_payload`.
pub fn blitTile(self: Layer, x: i32, y: i32, w: u32, h: u32, pixels: []const u8) bool {
var request = protocol.Request{
.operation = @intFromEnum(protocol.Operation.blit_tile),
.layer = self.id,
.x = @bitCast(x),
.y = @bitCast(y),
.width = w,
.height = h,
};
const header = std.mem.asBytes(&request);
if (header.len + pixels.len > protocol.message_maximum) return false;
var buffer: [protocol.message_maximum]u8 = undefined;
@memcpy(buffer[0..header.len], header);
@memcpy(buffer[header.len..][0..pixels.len], pixels);
const h_svc = service() orelse return false;
var reply: [protocol.reply_size]u8 = undefined;
const len = ipc.call(h_svc, buffer[0 .. header.len + pixels.len], &reply) catch return false;
if (len < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]).status == 0;
}
/// Move / restack / show or hide the layer.
pub fn configure(self: Layer, x: i32, y: i32, z: u32, visible: bool) bool {
var reply: protocol.Reply = undefined;
return transact(.{
.operation = @intFromEnum(protocol.Operation.configure_layer),
.layer = self.id,
.x = @bitCast(x),
.y = @bitCast(y),
.z = z,
.visible = if (visible) 1 else 0,
}, &reply);
}
/// Mark a rectangle of this layer (layer-local) dirty for the next present — for when
/// the layer's pixels changed without a drawing call the compositor already tracked.
pub fn damage(self: Layer, x: i32, y: i32, w: u32, h: u32) bool {
var reply: protocol.Reply = undefined;
return transact(.{
.operation = @intFromEnum(protocol.Operation.damage),
.layer = self.id,
.x = @bitCast(x),
.y = @bitCast(y),
.width = w,
.height = h,
}, &reply);
}
/// Release the layer and its surface.
pub fn destroy(self: Layer) bool {
var reply: protocol.Reply = undefined;
return transact(.{ .operation = @intFromEnum(protocol.Operation.destroy_layer), .layer = self.id }, &reply);
}
};
/// Create a server-owned layer of `w`×`h` pixels at screen (`x`, `y`) with stacking order
/// `z` (higher is nearer the front), initially visible. Returns a handle, or null.
pub fn createLayer(x: i32, y: i32, w: u32, h: u32, z: u32) ?Layer {
var reply: protocol.Reply = undefined;
if (!transact(.{
.operation = @intFromEnum(protocol.Operation.create_layer),
.x = @bitCast(x),
.y = @bitCast(y),
.width = w,
.height = h,
.z = z,
.visible = 1,
}, &reply)) return null;
return .{ .id = reply.layer };
}
+274
View File
@@ -0,0 +1,274 @@
//! runtime.fs — the danos-native file API. A program opens, reads, writes, and
//! lists files served by the user-space VFS (system/services/vfs), each call
//! marshalling a vfs-protocol request over IPC. This is the danos-native layer
//! danos programs use directly; it is also where the file operations that later
//! become `std.os.danos` are staged (see docs/zig-self-hosting.md). It replaces
//! the old POSIX `unistd` shim — a compatibility spelling danos does not need yet.
//!
//! Handles are *values*, not entries in a global descriptor table: a `File` /
//! `Directory` owns its VFS node id and (for files) a byte offset. So there is no
//! per-process fd limit and no shared table to synchronise — the danos-native
//! shape, unlike the POSIX fd model the old shim emulated.
const std = @import("std");
const ipc = @import("ipc.zig");
const protocol = @import("vfs-protocol");
/// The kind of a filesystem node — re-exported so a caller need not import the
/// wire protocol.
pub const Kind = protocol.NodeKind;
/// A node's metadata (the answer to a status request).
pub const Attributes = struct {
size: u64,
kind: Kind,
/// Modification time — Unix epoch seconds, UTC. 0 if the filesystem has none.
mtime: u64 = 0,
};
// Map a wire `NodeKind` value to the enum, defaulting anything unrecognised to
// `.regular` (the server is trusted, but a value outside the enum would be
// illegal to `@enumFromInt` directly).
fn kindFromWire(value: u32) Kind {
return switch (value) {
@intFromEnum(Kind.directory) => .directory,
@intFromEnum(Kind.character_device) => .character_device,
@intFromEnum(Kind.block_device) => .block_device,
@intFromEnum(Kind.symbolic_link) => .symbolic_link,
@intFromEnum(Kind.fifo) => .fifo,
@intFromEnum(Kind.socket) => .socket,
else => .regular,
};
}
/// How to open a path.
pub const OpenOptions = struct {
/// Create the file if it does not exist.
create: bool = false,
/// Open a directory node (for listing) rather than a file.
directory: bool = false,
/// Truncate an existing file to zero length on open (O_TRUNC) — replace its
/// contents rather than overwriting in place.
truncate: bool = false,
fn wireFlags(self: OpenOptions) u32 {
var f: u32 = 0;
if (self.create) f |= protocol.create;
if (self.directory) f |= protocol.directory;
if (self.truncate) f |= protocol.truncate;
return f;
}
};
// The VFS server endpoint, looked up once by well-known id and cached.
var vfs_handle: ipc.Handle = 0;
var vfs_resolved = false;
fn vfs() ?ipc.Handle {
if (!vfs_resolved) {
vfs_handle = ipc.lookup(.vfs) orelse return null;
vfs_resolved = true;
}
return vfs_handle;
}
const Result = struct { reply: protocol.Reply, payload: []u8 };
// One request/reply round trip: [Request header][send payload] -> VFS ->
// [Reply header][receive payload]. The receive payload lands in `out`.
fn transact(request: protocol.Request, send: []const u8, out: []u8) ?Result {
const h = vfs() orelse return null;
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const slen = @min(send.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..slen], send[0..slen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(h, message[0 .. protocol.request_size + slen], &rbuf) catch return null;
if (n < protocol.reply_size) return null;
const reply = std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]);
const rpl = @min(n - protocol.reply_size, out.len);
@memcpy(out[0..rpl], rbuf[protocol.reply_size..][0..rpl]);
return .{ .reply = reply, .payload = out[0..rpl] };
}
/// An open file: a VFS node plus a byte cursor. Read and write advance the cursor.
pub const File = struct {
node: u64,
offset: u64 = 0,
/// Read up to `buffer.len` bytes at the current offset; returns the count, or
/// null on error.
pub fn read(self: *File, buffer: []u8) ?usize {
const want: u32 = @intCast(@min(buffer.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .read, .node = self.node, .offset = self.offset, .len = want, .flags = 0 };
const r = transact(request, &.{}, buffer) orelse return null;
if (r.reply.status != 0) return null;
self.offset += r.reply.len;
return r.reply.len;
}
/// Write `data` at the current offset; returns the count written. A single
/// call is capped at the VFS payload size, so the return may be short — use
/// `writeAll` to write the whole slice. Null on error.
pub fn write(self: *File, data: []const u8) ?usize {
const want: u32 = @intCast(@min(data.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .write, .node = self.node, .offset = self.offset, .len = want, .flags = 0 };
const r = transact(request, data[0..want], &.{}) orelse return null;
if (r.reply.status != 0) return null;
self.offset += r.reply.len;
return r.reply.len;
}
/// Write all of `data`, looping past the per-call payload cap. Returns the
/// total written, or null if a write failed before any progress.
pub fn writeAll(self: *File, data: []const u8) ?usize {
var written: usize = 0;
while (written < data.len) {
const n = self.write(data[written..]) orelse return if (written == 0) null else written;
if (n == 0) return written; // no forward progress; stop rather than spin
written += n;
}
return written;
}
/// Move the read/write cursor to an absolute byte position.
pub fn seekTo(self: *File, position: u64) void {
self.offset = position;
}
/// This file's metadata.
pub fn attributes(self: *File) ?Attributes {
const request = protocol.Request{ .operation = .status, .node = self.node, .offset = 0, .len = 0, .flags = 0 };
var buffer: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return null;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return null;
const status = std.mem.bytesToValue(protocol.FileStatus, buffer[0..@sizeOf(protocol.FileStatus)]);
return .{ .size = status.size, .kind = kindFromWire(status.kind), .mtime = status.mtime };
}
/// Release the VFS's open handle for this file.
pub fn close(self: *File) void {
const request = protocol.Request{ .operation = .close, .node = self.node, .offset = 0, .len = 0, .flags = 0 };
_ = transact(request, &.{}, &.{});
}
};
/// Open (or create, with `.create`) `path`. Returns the open file, or null.
pub fn open(path: []const u8, options: OpenOptions) ?File {
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = options.wireFlags() };
const r = transact(request, path, &.{}) orelse return null;
if (r.reply.status != 0) return null;
return .{ .node = r.reply.node };
}
/// A path's metadata without keeping it open (open -> status -> close).
pub fn attributes(path: []const u8) ?Attributes {
var file = open(path, .{}) orelse return null;
defer file.close();
return file.attributes();
}
/// Whether `path` resolves — handy as a readiness check (e.g. waiting for a mount
/// to come up before writing to it).
pub fn exists(path: []const u8) bool {
return attributes(path) != null;
}
/// One entry returned by `Directory.next`.
pub const Entry = struct {
kind: Kind = .regular,
size: u64 = 0,
name_buffer: [64]u8 = undefined,
name_len: usize = 0,
pub fn name(self: *const Entry) []const u8 {
return self.name_buffer[0..self.name_len];
}
};
/// An open directory being listed, cursor-advanced by `next`.
pub const Directory = struct {
node: u64,
cursor: u64 = 0,
/// Fill `entry` with the next directory entry; false at end of directory or
/// on error.
pub fn next(self: *Directory, entry: *Entry) bool {
const request = protocol.Request{ .operation = .readdir, .node = self.node, .offset = self.cursor, .len = 0, .flags = 0 };
var buffer: [protocol.message_maximum]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return false;
if (r.reply.status != 0 or r.reply.len == 0) return false; // error or EOF
if (r.payload.len < protocol.directory_entry_size) return false;
const header = std.mem.bytesToValue(protocol.DirectoryEntry, r.payload[0..protocol.directory_entry_size]);
entry.kind = kindFromWire(header.kind);
entry.size = header.size;
const source = r.payload[protocol.directory_entry_size..];
const nlen = @min(@min(@as(usize, header.name_len), source.len), entry.name_buffer.len);
@memcpy(entry.name_buffer[0..nlen], source[0..nlen]);
entry.name_len = nlen;
self.cursor += 1;
return true;
}
/// Release the VFS's open handle for this directory.
pub fn close(self: *Directory) void {
var f = File{ .node = self.node };
f.close();
}
};
/// Open `path` as a directory for listing. Returns null if it isn't one / on error.
pub fn openDirectory(path: []const u8) ?Directory {
const file = open(path, .{ .directory = true }) orelse return null;
return .{ .node = file.node };
}
// A path-based request that returns only a status (mkdir, unlink).
fn pathOperation(operation: protocol.Operation, path: []const u8) bool {
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = 0 };
const r = transact(request, path, &.{}) orelse return false;
return r.reply.status == 0;
}
/// Create a directory at `path` (its parent must already exist). Returns true on
/// success. Only works under a mounted filesystem that supports directories.
pub fn makeDirectory(path: []const u8) bool {
return pathOperation(.mkdir, path);
}
/// Remove the file at `path`. Returns true on success. Directories are refused
/// (a separate directory-removal would have to check emptiness).
pub fn remove(path: []const u8) bool {
return pathOperation(.unlink, path);
}
/// Rename `old_path` to `new_path`. Both must be in the same directory (same-
/// directory, 8.3-name rename only for now). Returns true on success.
pub fn rename(old_path: []const u8, new_path: []const u8) bool {
const total = old_path.len + 1 + new_path.len;
if (total > protocol.maximum_payload) return false;
var payload: [protocol.maximum_payload]u8 = undefined;
@memcpy(payload[0..old_path.len], old_path);
payload[old_path.len] = 0;
@memcpy(payload[old_path.len + 1 ..][0..new_path.len], new_path);
const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 };
const r = transact(request, payload[0..total], &.{}) orelse return false;
return r.reply.status == 0;
}
/// Mount a filesystem backend (its server endpoint) at absolute path `target`;
/// the VFS then routes everything under `target` to that backend. This is the one
/// call that hands the VFS a capability (the backend endpoint). Returns true on
/// success.
pub fn mount(target: []const u8, backend: ipc.Handle) bool {
const h = vfs() orelse return false;
const request = protocol.Request{ .operation = .mount, .node = 0, .offset = 0, .len = @intCast(target.len), .flags = 0 };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const tlen = @min(target.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..tlen], target[0..tlen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const result = ipc.callCap(h, message[0 .. protocol.request_size + tlen], &rbuf, backend) catch return false;
if (result.len < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]).status == 0;
}
+19
View File
@@ -38,6 +38,25 @@ pub const device = @import("device.zig");
/// DMA-capable memory for drivers: contiguous, pinned, uncacheable buffers.
pub const dma = @import("dma.zig");
/// USB class-driver client: open a device on the xHCI bus and drive it
/// (control / interrupt / bulk transfers). See library/runtime/usb.zig.
pub const usb = @import("usb.zig");
/// Block-device client: read/write a block device (a USB stick, via
/// usb-storage). See library/runtime/block.zig.
pub const block = @import("block.zig");
/// Display-service client: query the mode, and (from D3) create layers, draw, and
/// present frames. See library/runtime/display.zig and system/services/display/.
pub const display = @import("display.zig");
/// The display wire protocol (shared with the display service and its clients).
pub const display_protocol = @import("display-protocol");
/// The danos-native file API (open/read/write/list over the user-space VFS) — the
/// layer danos programs use directly, and where the operations that later become
/// `std.os.danos` are staged. See docs/zig-self-hosting.md.
pub const fs = @import("fs.zig");
/// Re-exported so a user binary can `pub const panic = runtime.panic;`.
pub const panic = start.panic;
+18
View File
@@ -51,6 +51,24 @@ pub fn clock() u64 {
return @intCast(sc.systemCall0(.clock));
}
/// Wall-clock time in Unix epoch seconds (UTC) — the real date/time, from the RTC.
/// Unlike `clock` (monotonic since boot), this tracks calendar time, so it is what a
/// filesystem stamps as a file's modification time. Formatting it into a calendar
/// date/timezone is user-space policy layered on top.
pub fn wallClock() u64 {
return @intCast(sc.systemCall0(.wall_clock));
}
/// Copy bytes out of the kernel's in-memory diagnostic log — the accumulated
/// stream of everything `write` (and the kernel itself) has emitted — starting at
/// `offset`, into `out`. Returns the number of bytes copied (0 at end of buffer).
/// A program reads the whole log by looping from offset 0, advancing by the return
/// value, until it gets 0. This is how the boot log is persisted to disk on a
/// headless/real machine where serial output is otherwise lost.
pub fn klogRead(offset: usize, out: []u8) usize {
return sc.systemCall3(.klog_read, offset, @intFromPtr(out.ptr), out.len);
}
/// End the process. Never returns.
pub fn exit(code: usize) noreturn {
_ = sc.systemCall1(.exit, code);
+162
View File
@@ -0,0 +1,162 @@
//! USB class-driver client: the helper a keyboard, mouse, or mass-storage driver
//! uses to reach its device through the xHCI bus driver, so it never hand-rolls
//! the transfer-protocol IPC. Layered over `ipc` and the shared
//! `usb-transfer-protocol` wire format, the way `input.zig` layers over the input
//! service and `device.zig` over the raw device calls.
//!
//! A class driver, spawned with its interface's assigned device id as argv[1]:
//! if (!usb.helloManager(id)) return; // meet the spawn deadline
//! var device = usb.open(id) orelse return; // open + get its endpoints
//! _ = device.controlOut(usb_abi.setProtocol(...));// class requests, descriptors
//! _ = device.subscribeInterrupt(address, length); // reports arrive asynchronously
//! while (true) { ... ipc.replyWait(device.endpoint, ...) ... } // its own loop
//!
//! Reports are delivered to `device.endpoint` as asynchronous `InterruptReport`
//! messages (the class driver runs a bare `replyWait` loop to read them, because
//! the service harness drops buffered-message payloads — see service.zig).
const std = @import("std");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
const protocol = @import("usb-transfer-protocol");
const device_manager = @import("device-manager-protocol");
pub const Endpoint = protocol.Endpoint;
pub const InterruptReport = protocol.InterruptReport;
pub const max_report_data = protocol.max_report_data;
// Endpoint transfer types (EndpointDescriptor attributes), for `findEndpoint`.
pub const transfer_type_bulk: u8 = 2;
pub const transfer_type_interrupt: u8 = 3;
/// An opened USB device: the bus endpoint to send requests to, this driver's own
/// endpoint that reports arrive on, the device token, and the interface's
/// endpoints (so a driver need not re-read the configuration descriptor).
pub const Device = struct {
bus: ipc.Handle,
endpoint: ipc.Handle,
token: u64,
class: u8,
subclass: u8,
protocol_code: u8,
interface_number: u8,
endpoint_count: usize = 0,
endpoints: [protocol.max_reported_endpoints]Endpoint = undefined,
/// The interface's first endpoint of the given transfer type and direction
/// (`transfer_type_bulk` / `transfer_type_interrupt`), or null.
pub fn findEndpoint(self: *const Device, transfer_type: u8, direction_in: bool) ?Endpoint {
for (self.endpoints[0..self.endpoint_count]) |endpoint| {
if (endpoint.transfer_type == transfer_type and (endpoint.address & 0x80 != 0) == direction_in) return endpoint;
}
return null;
}
fn controlTransfer(self: *Device, setup: [8]u8, direction_in: bool, data: []u8) ?usize {
var request = protocol.ControlRequest{
.device_token = self.token,
.setup = setup,
.direction_in = @intFromBool(direction_in),
.data_length = @intCast(data.len),
};
if (!direction_in and data.len > 0) @memcpy(request.data[0..data.len], data);
var reply: [@sizeOf(protocol.ControlReply)]u8 = undefined;
const length = ipc.call(self.bus, std.mem.asBytes(&request), &reply) catch return null;
if (length < @sizeOf(protocol.ControlReply)) return null;
const control_reply = std.mem.bytesToValue(protocol.ControlReply, reply[0..@sizeOf(protocol.ControlReply)]);
if (control_reply.status != 0) return null;
const actual = @min(control_reply.actual_length, data.len);
if (direction_in and actual > 0) @memcpy(data[0..actual], control_reply.data[0..actual]);
return actual;
}
/// A control transfer with no data stage (SET_PROTOCOL, SET_IDLE, ...). The
/// `setup` is a bit-cast `usb_abi.Request`.
pub fn controlOut(self: *Device, setup: [8]u8) bool {
return self.controlTransfer(setup, false, &.{}) != null;
}
/// A device-to-host control transfer, returning the bytes read into `out`.
pub fn controlIn(self: *Device, setup: [8]u8, out: []u8) ?usize {
return self.controlTransfer(setup, true, out);
}
/// Begin periodic IN polling of an interrupt endpoint; reports flow back to
/// `self.endpoint` as asynchronous `InterruptReport` messages.
pub fn subscribeInterrupt(self: *Device, endpoint_address: u8, max_length: u16) bool {
var request = protocol.InterruptSubscribeRequest{
.device_token = self.token,
.endpoint_address = endpoint_address,
.max_length = max_length,
};
var reply: [@sizeOf(protocol.InterruptSubscribeReply)]u8 = undefined;
const length = ipc.call(self.bus, std.mem.asBytes(&request), &reply) catch return false;
if (length < @sizeOf(protocol.InterruptSubscribeReply)) return false;
return std.mem.bytesToValue(protocol.InterruptSubscribeReply, reply[0..@sizeOf(protocol.InterruptSubscribeReply)]).status == 0;
}
/// One bulk transfer (IN or OUT per `endpoint_address`'s direction bit) to or
/// from the caller's own DMA buffer at `physical`. Returns the bytes moved.
pub fn bulk(self: *Device, endpoint_address: u8, physical: u64, length: u32) ?u32 {
var request = protocol.BulkRequest{
.device_token = self.token,
.physical_address = physical,
.length = length,
.endpoint_address = endpoint_address,
};
var reply: [@sizeOf(protocol.BulkReply)]u8 = undefined;
const replied = ipc.call(self.bus, std.mem.asBytes(&request), &reply) catch return null;
if (replied < @sizeOf(protocol.BulkReply)) return null;
const bulk_reply = std.mem.bytesToValue(protocol.BulkReply, reply[0..@sizeOf(protocol.BulkReply)]);
if (bulk_reply.status != 0) return null;
return bulk_reply.actual_length;
}
};
/// Look up the USB bus and open the device with the assigned id, handing over a
/// freshly created endpoint for asynchronous interrupt reports. Retries while the
/// bus is still coming up (a class driver races the bus driver at boot).
pub fn open(device_id: u64) ?Device {
var attempts: usize = 0;
const bus = while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.usb_bus)) |handle| break handle;
system.sleep(20);
} else return null;
const endpoint = ipc.createIpcEndpoint() orelse return null;
var request = protocol.OpenRequest{ .device_id = device_id };
var reply: [@sizeOf(protocol.OpenReply)]u8 = undefined;
const result = ipc.callCap(bus, std.mem.asBytes(&request), &reply, endpoint) catch return null;
if (result.len < @sizeOf(protocol.OpenReply)) return null;
const open_reply = std.mem.bytesToValue(protocol.OpenReply, reply[0..@sizeOf(protocol.OpenReply)]);
if (open_reply.status != 0) return null;
var device = Device{
.bus = bus,
.endpoint = endpoint,
.token = open_reply.device_token,
.class = open_reply.interface_class,
.subclass = open_reply.interface_subclass,
.protocol_code = open_reply.interface_protocol,
.interface_number = open_reply.interface_number,
.endpoint_count = @min(open_reply.endpoint_count, protocol.max_reported_endpoints),
};
for (0..device.endpoint_count) |index| device.endpoints[index] = open_reply.endpoints[index];
return device;
}
/// Hello the device manager as a class driver (Role.device) so a supervised
/// spawn meets its hello deadline. Retries while the manager comes up.
pub fn helloManager(device_id: u64) bool {
var attempts: usize = 0;
const manager = while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.device_manager)) |handle| break handle;
system.sleep(20);
} else return false;
const hello = device_manager.Hello{ .role = @intFromEnum(device_manager.Role.device), .device_id = device_id };
var reply: [device_manager.message_maximum]u8 = undefined;
const length = ipc.call(manager, std.mem.asBytes(&hello), &reply) catch return false;
if (length < device_manager.reply_size) return false;
return std.mem.bytesToValue(device_manager.HelloReply, reply[0..device_manager.reply_size]).status == 0;
}
+6
View File
@@ -58,6 +58,8 @@ pub const SystemCall = enum(u64) {
signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on
process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself)
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
klog_read = 32, // klog_read(offset, ptr, len) -> bytes copied: copy the kernel RAM log buffer out to a user buffer (for persisting the boot log to disk)
wall_clock = 33, // wall_clock() -> Unix epoch seconds (UTC): the RTC wall-clock time, for filesystem timestamps (mtime). Monotonic time is `clock`.
_,
};
@@ -178,6 +180,10 @@ pub const ServiceId = enum(u32) {
ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes
device_manager = 4, // the tree, the matcher, the supervisor (docs/device-manager.md)
power = 5, // system power: events (button, lid, battery) + shutdown (docs/power.md; domain-named per docs/discovery.md — the acpi service registers it on x86, a PSCI service will on ARM)
usb_bus = 6, // the xHCI host-controller driver's transfer endpoint; USB class drivers look it up and `callCap`-open their device to get a private per-device transfer channel (docs/driver-model.md)
block = 7, // a block-device driver (USB mass storage today): read/write of fixed-size blocks, the storage a filesystem sits on
fat = 8, // the FAT filesystem server; the VFS mounts it and forwards paths under its mount point (/mnt/usb) to it
display = 9, // the display service: owns the framebuffer, composites a layer stack, presents frames (docs/display.md)
_,
};
+28 -62
View File
@@ -3,11 +3,13 @@
//! Walks the ACPI tables the firmware left in memory (starting from the RSDP the
//! bootloader handed us) and translates the static tables into the generic
//! `device` model, so the kernel enumerates hardware without knowing ACPI is the
//! source. This is deliberately the *static-table* path: MADT (CPUs / interrupt
//! controllers), MCFG (PCIe ECAM -> PCI enumeration), HPET (timer), and FADT
//! (power register map). The DSDT/SSDT bytecode is handed to the `aml` submodule
//! only to extract the sleep-state (`_Sx`) values for power management; full AML namespace
//! interpretation is a separate, larger subproject.
//! source. This is deliberately the *static-table* path, and **only** that: MADT
//! (CPUs / interrupt controllers), MCFG (PCIe ECAM -> PCI enumeration), HPET
//! (timer), and FADT (power register map). The DSDT/SSDT bytecode is *not*
//! interpreted here — the kernel collects the blobs and publishes them on the
//! acpi-tables node for the ring-3 acpi service to parse (device enumeration and
//! soft-off). Keeping the ~0.5 MB AML interpretation out of kernel init keeps it
//! off the single-core critical path (nothing else runs alongside it there).
//!
//! ACPI tables live in `.acpi_tables` / `.acpi_nvs` memory, which the kernel
//! identity-maps, so table addresses are dereferenced directly. PCIe ECAM is MMIO
@@ -19,7 +21,6 @@ const boot_handoff = @import("boot-handoff");
const abi = @import("abi");
const parameters = @import("parameters");
const device_model = @import("device-model.zig");
const aml = @import("aml/aml.zig");
const DeviceTree = device_model.DeviceTree;
const Hal = device_model.Hal;
@@ -37,8 +38,11 @@ pub const RegisterAccess = struct {
}
};
/// Everything the power subsystem needs, extracted from the FADT and the AML
/// sleep packages during discovery. Populated by `discover`, read by `power`.
/// The power register map, extracted from the FADT during discovery. Populated by
/// `discover`, read by `power` (kernel reboot). The **sleep-state (`_Sx`) values
/// live in AML**, which the kernel no longer parses — soft-off (S5) is owned by the
/// ring-3 acpi service (it re-parses the blobs on the published acpi-tables node and
/// writes the PM1 control register itself). So this holds only the FADT scalars.
pub const PowerInformation = struct {
/// The System Control Interrupt's GSI (FADT SCI_INT) — the line ACPI events
/// (power button, GPEs) arrive on. Published to the acpi service for M21.
@@ -54,10 +58,6 @@ pub const PowerInformation = struct {
reset: RegisterAccess = .{},
reset_value: u8 = 0,
reset_supported: bool = false,
/// SLP_TYP values for S5 (soft off) and S3 (suspend), from the AML sleep-state (`_Sx`)
/// packages.
s5: ?aml.SleepType = null,
s3: ?aml.SleepType = null,
};
/// Filled in by `discover`; the power service reads it to reboot/shutdown.
@@ -141,19 +141,6 @@ const maximum_cpus = parameters.maximum_cpus;
/// Filled in by `discover` (from the MADT); SMP bring-up reads it to wake the APs.
pub var cpu_information: CpuInformation = .{};
/// Integrity/diagnostics for the AML parse. `consumed == total` means the parser
/// walked every byte of the DSDT/SSDTs without desyncing.
pub const AmlStats = struct {
nodes: usize = 0,
consumed: usize = 0,
total: usize = 0,
};
pub var aml_stats: AmlStats = .{};
/// The ACPI namespace built from the DSDT/SSDTs, kept for sleep-state (`_Sx`) lookup now and
/// device enumeration later. Null until `discover` runs successfully.
pub var namespace: ?aml.Namespace = null;
/// Physical address of the DSDT the FADT points at, or 0.
pub var dsdt_physical: u64 = 0;
@@ -165,8 +152,9 @@ var fadt_physical: u64 = 0;
var fadt_length: u64 = 0;
// AML blocks (DSDT + any SSDTs) collected during the table walk, as physical
// address + length of each table's post-header bytecode. Scanned after the walk
// for the sleep-state (`_Sx`) packages.
// address + length of each table's post-header bytecode. The kernel does not
// interpret them — it publishes them on the acpi-tables node for the ring-3 acpi
// service to parse (device enumeration + soft-off). See publishAcpiTablesNode.
var aml_block_physical: [32]u64 = undefined;
var aml_block_len: [32]usize = undefined;
var aml_block_count: usize = 0;
@@ -373,7 +361,7 @@ const Hpet = extern struct {
/// Discover hardware from the ACPI tables rooted at `rsdp_physical` and populate
/// `device_tree`. `hal` provides MMIO mapping (for PCIe ECAM) and port I/O. Also parses the
/// FADT and the AML sleep-state (`_Sx`) packages into `power_information` for the power service.
/// FADT into `power_information`, and publishes the AML blobs for the ring-3 acpi service.
pub fn discover(rsdp_physical: u64, memory_regions: []const boot_handoff.MemoryRegion, device_tree: *DeviceTree, hal: Hal) !void {
if (rsdp_physical == 0) return error.NoRsdp;
boot_memory_regions = memory_regions;
@@ -383,8 +371,6 @@ pub fn discover(rsdp_physical: u64, memory_regions: []const boot_handoff.MemoryR
fadt_physical = 0;
fadt_length = 0;
platform_information = .{};
aml_stats = .{};
namespace = null;
dsdt_physical = 0;
aml_block_count = 0;
@@ -401,34 +387,20 @@ pub fn discover(rsdp_physical: u64, memory_regions: []const boot_handoff.MemoryR
try walkRoot(u32, rsdp.root_system_description_table_address, device_tree, hal);
}
// Now that the DSDT and any SSDTs are collected, build the AML namespace and
// read the sleep types from it.
var blocks: [aml_block_physical.len][]const u8 = undefined;
for (0..aml_block_count) |i| {
blocks[i] = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(aml_block_physical[i])))[0..aml_block_len[i]];
}
const active = blocks[0..aml_block_count];
if (aml.parse(device_tree.allocator, active)) |pr| {
namespace = pr.namespace;
aml_stats = .{ .nodes = namespace.?.nodeCount(), .consumed = pr.consumed, .total = pr.total };
power_information.s5 = aml.sleepState(&namespace.?, 5);
power_information.s3 = aml.sleepState(&namespace.?, 3);
// The namespace's Device objects are no longer folded into the kernel
// tree (M20.3): the ring-3 acpi service claims the acpi-tables node
// (published below), re-parses the same blobs, and registers + reports
// the _HID devices itself. The kernel keeps the namespace only for the
// \_S5 sleep type above.
} else |_| {
// AML parse failed (e.g. out of memory); power stays best-effort with
// whatever the FADT alone provided.
}
// The kernel does **not** interpret the DSDT/SSDTs. Static-table discovery
// above (MADT/HPET/FADT/MCFG) is all the kernel needs — CPUs, timers, PCIe,
// and the power register map. The AML bytecode (device enumeration and the
// sleep-state `_Sx` values for soft-off) is entirely the ring-3 acpi service's
// job: it claims the acpi-tables node published below, parses the same blobs,
// and both registers the `_HID` devices and owns S5. Not parsing ~0.5 MB of
// AML in the kernel keeps boot latency off the critical, single-core path.
// Publish the acpi-tables node (docs/discovery.md): the AML blobs as
// memory resources for the acpi service to map and parse in ring 3, a broad
// io_port grant for the OperationRegion access its interpreter needs, and
// the SCI for the events track (M21). Exactly one node, one trusted
// claimant. Kept even when the kernel-side device building (above) retires
// in M20.3 — the kernel still owns the *static* tables and \_S5.
// claimant — the sole path by which AML (devices + soft-off) reaches ring 3,
// now that the kernel keeps only the *static* tables for itself.
publishAcpiTablesNode(device_tree) catch {};
}
@@ -461,12 +433,6 @@ fn publishAcpiTablesNode(device_tree: *DeviceTree) !void {
if (fadt_physical != 0) _ = node.addResource(.memory, fadt_physical, fadt_length);
}
/// The number of Device objects in the namespace built during discovery, or 0.
pub fn amlDeviceCount() usize {
if (namespace) |*ns| return aml.deviceCount(ns);
return 0;
}
/// Walk the RSDT (Entry = u32) or XSDT (Entry = u64): validate it, then dispatch
/// each SDT it points at. A bad individual table is skipped, not fatal.
fn walkRoot(comptime Entry: type, root_physical: u64, device_tree: *DeviceTree, hal: Hal) !void {
@@ -502,7 +468,7 @@ fn handleTable(device_tree: *DeviceTree, hal: Hal, sdt_physical: u64) !void {
} else if (std.mem.eql(u8, &sig, &DMAR)) {
parseDmar(hal, header);
} else if (std.mem.eql(u8, &sig, &SSDT)) {
// Secondary namespace bytecode — collect for the sleep-state (`_Sx`) scan.
// Secondary namespace bytecode — collect it to publish for the ring-3 parse.
addAmlBlock(sdt_physical);
}
// Any other signature is recognised but left opaque for now.
@@ -740,8 +706,8 @@ const fadt_x_pm_tmr_blk = 208; // GAS
const flag_reset_register_supported = 1 << 10;
const flag_tmr_value_ext = 1 << 8; // PM timer counter is 32-bit (else 24-bit)
/// FADT -> the power register map (into `power_information`) and the DSDT address, which
/// is queued for the AML sleep-state (`_Sx`) scan. No AML interpretation happens here.
/// FADT -> the power register map (into `power_information`) and the DSDT address,
/// whose bytecode is collected for the ring-3 parse. No AML interpretation here.
fn parseFadt(header: *const SystemDescriptorTableHeader) void {
const base: [*]align(1) const u8 = @ptrCast(header);
const len: usize = header.length;
+46
View File
@@ -33,6 +33,18 @@ pub const DeviceClass = enum(u32) {
/// a broad io_port grant for OperationRegion access, and the SCI interrupt.
/// The one node whose claimant is trusted to run firmware bytecode.
acpi_tables,
/// One interface of a USB device, registered by the xHCI bus driver. It owns
/// no MMIO — it is reached through its controller — so it carries no
/// resources; the (class, subclass, protocol) triple that says what it is
/// travels in the bus report's identity, not here.
usb_device,
/// A scanout framebuffer: a linear region of pixel memory the display service
/// claims and maps. Unlike the other classes this one is not firmware-discovered
/// — the kernel seeds it from the loader's [[boot-handoff]] framebuffer
/// (`devices_broker.seedDisplay`). Its one `memory` resource is the framebuffer,
/// flagged write-combining; the geometry to interpret it travels in
/// `DeviceDescriptor.display`.
display,
unknown,
};
@@ -54,10 +66,39 @@ pub const ResourceDescriptor = extern struct {
kind: u64, // a ResourceKind value
start: u64,
len: u64,
/// A bitmask of `resource_flag_*` hints. Zero for a plain register/RAM window;
/// the kernel reads it when it maps the resource. Defaulted so every existing
/// literal (which never set flags) keeps compiling and lays out identically.
flags: u64 = 0,
};
/// `ResourceDescriptor.flags`: map this `memory` resource **write-combining** rather
/// than strong-uncacheable — for a framebuffer, where batched bursts to pixel memory
/// are the whole point (an uncacheable framebuffer blit is glacial). See
/// `mmio_map` (system/kernel/process.zig) and `setupPat` (…/x86_64/paging.zig).
pub const resource_flag_write_combining: u64 = 1 << 0;
pub const maximum_device_resources = 8;
/// The byte order of a display's pixels — mirrors the loader's `PixelFormat`
/// ([[boot-handoff]]) with the same numeric values, but lives here so user space
/// (which must never import the loader↔kernel handoff) can name it. Only the two
/// linear 32-bpp layouts a console can paint into exist; see docs/gop.md.
pub const DisplayFormat = enum(u32) {
rgbx = 0, // byte 0 = Red, 1 = Green, 2 = Blue, 3 = reserved
bgrx = 1, // byte 0 = Blue, 1 = Green, 2 = Red, 3 = reserved
};
/// The geometry of a `display` device's framebuffer, carried in its descriptor so a
/// claiming driver knows how to interpret the pixel bytes its `memory` resource maps.
/// `pitch` is bytes per row (may exceed `width * 4`; see docs/framebuffer.md).
pub const DisplayInfo = extern struct {
width: u32 = 0, // visible pixels per row
height: u32 = 0, // visible rows
pitch: u32 = 0, // bytes from one row's start to the next
format: u32 = 0, // a DisplayFormat value
};
/// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree.
pub const no_parent: u64 = ~@as(u64, 0);
@@ -87,4 +128,9 @@ pub const DeviceDescriptor = extern struct {
resource_count: u64,
hid: [8]u8,
resources: [maximum_device_resources]ResourceDescriptor,
// Framebuffer geometry, meaningful only when `class` is `DeviceClass.display`
// (zeroed otherwise). Kept here — a class-specific field on the shared descriptor —
// the same way `pci_class` is meaningful only for `pci_device` and `hid` only for
// `acpi_device`.
display: DisplayInfo = .{},
};
+5 -20
View File
@@ -22,13 +22,14 @@ pub const Resource = device_model.Resource;
pub const ResourceKind = device_model.ResourceKind;
pub const Hal = device_model.Hal;
pub const PowerInformation = acpi.PowerInformation;
pub const AmlStats = acpi.AmlStats;
pub const PlatformInformation = acpi.PlatformInformation;
pub const RegisterAccess = acpi.RegisterAccess;
pub const IsoEntry = acpi.IsoEntry;
pub const Cpu = acpi.Cpu;
/// The register map + sleep types discovery extracted, for logging/diagnostics.
/// The FADT power register map discovery extracted (PM1 control, reset register),
/// for kernel reboot and diagnostics. Sleep-state values are userspace's (S5 is
/// owned by the ring-3 acpi service), so they are not here.
pub fn powerInformation() PowerInformation {
return acpi.power_information;
}
@@ -39,18 +40,6 @@ pub fn platformInformation() PlatformInformation {
return acpi.platform_information;
}
/// AML parse integrity/diagnostics (namespace node count, bytes consumed).
/// The number of Device objects in the kernel's own AML namespace, or 0 if the
/// parse produced none — the `acpi-parse` test compares the ring-3 service's
/// count against this.
pub fn amlDeviceCount() usize {
return acpi.amlDeviceCount();
}
pub fn amlStats() AmlStats {
return acpi.aml_stats;
}
/// The usable logical processors discovered during enumeration — one entry per
/// core danos may schedule on, each carrying the Local APIC ID an SMP wake targets.
/// `len` is the hardware's degree of parallelism: how many tasks *could* run at the
@@ -92,12 +81,8 @@ pub fn discover(
}
/// Restart the machine. Never returns on success; returns only if no reset method
/// worked (extremely unlikely). Backend-agnostic entry the kernel calls.
/// worked (extremely unlikely). Backend-agnostic entry the kernel calls. Soft-off
/// (S5) is not a kernel operation — the ring-3 acpi service owns it (docs/power.md).
pub fn reboot(hal: Hal) void {
power.reboot(hal);
}
/// Power the machine off (ACPI S5). Never returns on success.
pub fn shutdown(hal: Hal) void {
power.shutdown(hal);
}
+10 -63
View File
@@ -1,34 +1,17 @@
//! Machine power control: enter ACPI mode, reboot, and power off (ACPI S5).
//! Machine reboot: restart via the FADT reset register, with legacy fallbacks.
//!
//! Built entirely on the register map `acpi` extracted from the FADT plus the
//! sleep-state (`_Sx`) types the AML submodule pulled from the DSDT, driven through the
//! injected `Hal` (port I/O and MMIO). Nothing here is x86-specific beyond the
//! well-known legacy reset fallbacks, which are guarded behind the ACPI methods.
//!
//! S3 (suspend-to-RAM) is stubbed: it needs a wake trampoline and device
//! re-initialisation, a milestone of its own.
//! Built on the register map `acpi` extracted from the FADT, driven through the
//! injected `Hal` (port I/O and MMIO). Soft-off (ACPI S5) and suspend (S3) are
//! **not** here: they need the AML sleep-state (`_Sx`) values, which the kernel no
//! longer parses — the ring-3 acpi service owns power management (it re-parses the
//! blobs and writes the PM1 control register itself). See docs/power.md. Reboot
//! stays in the kernel because it needs no AML — only the FADT reset register and
//! the well-known legacy fallbacks — so it survives as a last-resort restart.
const acpi = @import("acpi.zig");
const device_model = @import("device-model.zig");
const Hal = device_model.Hal;
const slp_en: u32 = 1 << 13; // SLP_EN: writing 1 triggers the sleep transition
const sci_en: u32 = 1 << 0; // SCI_EN in PM1 control: set once ACPI mode is active
/// Switch the platform into ACPI mode if it isn't already, so the PM1 control
/// register is live. A no-op when the firmware exposes no SMI command port (ACPI
/// already enabled, as under QEMU/OVMF) — we still verify SCI_EN first.
pub fn enable(hal: Hal) void {
const pi = acpi.power_information;
if (!pi.pm1a_cnt.present()) return;
if (readRegister(hal, pi.pm1a_cnt) & sci_en != 0) return; // already in ACPI mode
if (pi.smi_cmd == 0 or pi.acpi_enable == 0) return; // no way to switch; assume fine
hal.pioWrite(1, pi.smi_cmd, pi.acpi_enable);
var spins: usize = 0;
while (readRegister(hal, pi.pm1a_cnt) & sci_en == 0 and spins < 1_000_000) : (spins += 1) {}
}
/// Restart the machine. Tries the ACPI reset register first, then the two legacy
/// fallbacks. Returns only if every method failed (very unlikely).
pub fn reboot(hal: Hal) void {
@@ -48,42 +31,6 @@ pub fn reboot(hal: Hal) void {
delay();
}
/// Power the machine off via ACPI S5. Requires the soft-off (`_S5`) sleep type; if
/// it wasn't found in the AML, there is nothing safe to do and this returns.
pub fn shutdown(hal: Hal) void {
enable(hal);
const pi = acpi.power_information;
const s5 = pi.s5 orelse return;
if (pi.pm1a_cnt.present()) {
writeRegister(hal, pi.pm1a_cnt, sleepValue(s5.slp_typ_a));
}
if (pi.pm1b_cnt.present()) {
writeRegister(hal, pi.pm1b_cnt, sleepValue(s5.slp_typ_b));
}
delay();
}
/// S3 suspend-to-RAM — not implemented (needs a wake path + device re-init).
pub fn sleepS3(hal: Hal) error{Unsupported}!void {
_ = hal;
return error.Unsupported;
}
/// The PM1 control write that requests sleep type `slp_typ`: SLP_TYP in bits
/// [12:10], SLP_EN in bit 13.
fn sleepValue(slp_typ: u8) u32 {
return (@as(u32, slp_typ & 0x7) << 10) | slp_en;
}
fn readRegister(hal: Hal, register: acpi.RegisterAccess) u32 {
if (register.mmio) {
const p: *align(1) volatile u32 = @ptrFromInt(hal.mapMmio(register.address, 4, true));
return p.*;
}
return hal.pioRead(register.width, @intCast(register.address));
}
fn writeRegister(hal: Hal, register: acpi.RegisterAccess, value: u32) void {
if (register.mmio) {
const p: *align(1) volatile u32 = @ptrFromInt(hal.mapMmio(register.address, 4, true));
@@ -93,8 +40,8 @@ fn writeRegister(hal: Hal, register: acpi.RegisterAccess, value: u32) void {
}
}
/// A short busy-wait so a reset/power-off takes effect before we fall through to
/// the next method. The empty asm is an architecture-neutral barrier that keeps the loop
/// A short busy-wait so a reset takes effect before we fall through to the next
/// method. The empty asm is an architecture-neutral barrier that keeps the loop
/// from being optimised away.
fn delay() void {
var i: usize = 0;
+139 -51
View File
@@ -8,7 +8,7 @@
//! buffer at any offset, and bitmap bytes are packed structs so no caller ever needs a magic
//! mask. Class, subclass, and protocol code tables live in usb-ids.zig.
const DeviceState = enum(u8) {
pub const DeviceState = enum(u8) {
// Immediately after the USB device is attached to the USB system, it is in this state.
// The USB specifications do not define the state of a USB device that is detached from
// a USB system.
@@ -47,7 +47,7 @@ const DeviceState = enum(u8) {
suspended,
};
const RequestCode = enum(u8) {
pub const RequestCode = enum(u8) {
get_status = 0,
clear_feature = 1,
set_feature = 3,
@@ -59,10 +59,15 @@ const RequestCode = enum(u8) {
get_interface = 10,
set_interface = 11,
sync_frame = 12,
// Non-exhaustive: class-specific requests (HID, mass storage) reuse this byte
// field with codes from their own class's namespace — see the class-request
// constructors below. Some class codes numerically coincide with a standard
// one; the wire byte is what matters, and the constructors set it explicitly.
_,
};
// Direction of an endpoint, from the host's point of view
const EndpointDirection = enum(u1) {
pub const EndpointDirection = enum(u1) {
out = 0,
in = 1,
};
@@ -74,7 +79,7 @@ const EndpointDirection = enum(u1) {
// The bus address of a device, assigned by the host with SET_ADDRESS. Addresses are 7 bits
// wide.
const DeviceAddress = enum(u7) {
pub const DeviceAddress = enum(u7) {
// The default address every device answers at after a reset, until SET_ADDRESS
// completes
default = 0,
@@ -82,7 +87,7 @@ const DeviceAddress = enum(u7) {
};
// Identifies a configuration; from ConfigurationDescriptor.configuration_value.
const ConfigurationValue = enum(u8) {
pub const ConfigurationValue = enum(u8) {
// Not configured: returned by GET_CONFIGURATION while the device is in the address
// state, and passed to SET_CONFIGURATION to return a configured device to the address
// state
@@ -92,11 +97,11 @@ const ConfigurationValue = enum(u8) {
// Identifies an interface within a configuration; from
// InterfaceDescriptor.interface_number.
const InterfaceNumber = enum(u8) { _ };
pub const InterfaceNumber = enum(u8) { _ };
// Selects between the alternate settings of one interface; from
// InterfaceDescriptor.alternate_setting.
const AlternateSetting = enum(u8) {
pub const AlternateSetting = enum(u8) {
// The default setting of an interface
default = 0,
_,
@@ -104,7 +109,7 @@ const AlternateSetting = enum(u8) {
// The number of an endpoint within a device, 4 bits wide. The direction bit carried
// alongside it tells the two endpoints sharing a number apart.
const EndpointNumber = enum(u4) {
pub const EndpointNumber = enum(u4) {
// Endpoint zero: the default control pipe every device provides
default_control = 0,
_,
@@ -112,7 +117,7 @@ const EndpointNumber = enum(u4) {
// Index of a STRING descriptor, stored in descriptors that reference a string and passed to
// GET_DESCRIPTOR to read it.
const StringIndex = enum(u8) {
pub const StringIndex = enum(u8) {
// The device has no string descriptor for this field
none = 0,
_,
@@ -121,7 +126,7 @@ const StringIndex = enum(u8) {
// Characteristics of a device request (the bmRequestType field of a set-up packet). Fields are
// declared least-significant first: recipient occupies bits 4...0, kind bits 6...5, and
// direction bit 7.
const RequestType = packed struct(u8) {
pub const RequestType = packed struct(u8) {
// The recipient of the request (values 4...31 are reserved)
recipient: Recipient,
// The type of the request
@@ -129,27 +134,27 @@ const RequestType = packed struct(u8) {
// Data transfer direction. The value of this bit is ignored when length is zero.
direction: Direction,
const Recipient = enum(u5) {
pub const Recipient = enum(u5) {
device = 0,
interface = 1,
endpoint = 2,
other = 3,
};
const Kind = enum(u2) {
pub const Kind = enum(u2) {
standard = 0,
class = 1,
vendor = 2,
reserved = 3,
};
const Direction = enum(u1) {
pub const Direction = enum(u1) {
host_to_device = 0,
device_to_host = 1,
};
};
const Request = extern struct {
pub const Request = extern struct {
// Characteristics of the request
request_type: RequestType,
// Specific request
@@ -175,7 +180,7 @@ const Request = extern struct {
// The format of the index field when request_type specifies an endpoint as the
// recipient. The host should always set the direction bit to zero (but the device
// should accept either value) when the endpoint is part of a control pipe.
const EndpointIndex = packed struct(u16) {
pub const EndpointIndex = packed struct(u16) {
// Endpoint number
number: EndpointNumber,
// Reserved (reset to zero)
@@ -188,7 +193,7 @@ const Request = extern struct {
// The format of the index field when request_type specifies an interface as the
// recipient.
const InterfaceIndex = packed struct(u16) {
pub const InterfaceIndex = packed struct(u16) {
// Interface number
number: u8,
// Reserved (reset to zero)
@@ -199,7 +204,7 @@ const Request = extern struct {
// descriptor type in the high byte, and the descriptor index in the low byte. The index
// is used to select a specific descriptor (only for CONFIGURATION and STRING
// descriptors) when several descriptors of that type are implemented by a device.
const DescriptorValue = packed struct(u16) {
pub const DescriptorValue = packed struct(u16) {
// Descriptor index
index: u8 = 0,
// Descriptor type
@@ -209,7 +214,7 @@ const Request = extern struct {
// Feature selectors, used as the value field of CLEAR_FEATURE and SET_FEATURE requests. The
// comment on each value notes the recipient the selector applies to.
const FeatureSelector = enum(u16) {
pub const FeatureSelector = enum(u16) {
// Halts an endpoint (recipient: endpoint)
endpoint_halt = 0,
// Enables or disables the device's remote wakeup capability (recipient: device)
@@ -223,7 +228,7 @@ const FeatureSelector = enum(u16) {
// with the test_mode feature selector. Values 06h...3Fh are reserved for standard test
// selectors and C0h...FFh for vendor-specific test modes; all other unlisted values are
// reserved.
const TestMode = enum(u8) {
pub const TestMode = enum(u8) {
test_j = 0x01,
test_k = 0x02,
test_se0_nak = 0x03,
@@ -234,7 +239,7 @@ const TestMode = enum(u8) {
// The two bytes returned by a GET_STATUS request directed at a device. Fields are declared
// least-significant first.
const DeviceStatus = packed struct(u16) {
pub const DeviceStatus = packed struct(u16) {
// Whether the device is currently self-powered (as opposed to bus-powered). This bit
// cannot be changed with the SET_FEATURE or CLEAR_FEATURE requests.
self_powered: bool,
@@ -248,7 +253,7 @@ const DeviceStatus = packed struct(u16) {
// The two bytes returned by a GET_STATUS request directed at an endpoint. (A GET_STATUS
// request directed at an interface returns two bytes that are entirely reserved.)
const EndpointStatus = packed struct(u16) {
pub const EndpointStatus = packed struct(u16) {
// Whether the endpoint is currently halted. Set with the SET_FEATURE request using the
// endpoint_halt feature selector, and cleared with CLEAR_FEATURE.
halted: bool,
@@ -258,7 +263,7 @@ const EndpointStatus = packed struct(u16) {
// A target for the standard requests that may be directed at the device, an interface, or
// an endpoint.
const Target = union(enum) {
pub const Target = union(enum) {
device,
interface: InterfaceNumber,
endpoint: Request.EndpointIndex,
@@ -287,7 +292,7 @@ const Target = union(enum) {
// Reads the status of the given target: bit-cast the two bytes the device returns into a
// DeviceStatus or an EndpointStatus. (The two bytes returned for an interface are entirely
// reserved.)
fn getStatus(target: Target) Request {
pub fn getStatus(target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
@@ -303,7 +308,7 @@ fn getStatus(target: Target) Request {
// Clears or disables the given feature. A device cannot be taken out of a test mode with
// this request; test_mode is only cleared by cycling power.
fn clearFeature(feature: FeatureSelector, target: Target) Request {
pub fn clearFeature(feature: FeatureSelector, target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
@@ -319,7 +324,7 @@ fn clearFeature(feature: FeatureSelector, target: Target) Request {
// Sets or enables the given feature. For the test_mode feature selector, use setTestMode
// instead: the test selector rides in the high byte of the index field.
fn setFeature(feature: FeatureSelector, target: Target) Request {
pub fn setFeature(feature: FeatureSelector, target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
@@ -335,7 +340,7 @@ fn setFeature(feature: FeatureSelector, target: Target) Request {
// Puts a hi-speed device into the given test mode: a SET_FEATURE request with the test_mode
// feature selector and the test selector in the high byte of the index field.
fn setTestMode(mode: TestMode) Request {
pub fn setTestMode(mode: TestMode) Request {
return .{
.request_type = .{
.recipient = .device,
@@ -352,7 +357,7 @@ fn setTestMode(mode: TestMode) Request {
// Assigns the device its bus address, moving it from the default state to the address
// state. The device does not answer at the new address until the status stage of this
// request completes.
fn setAddress(address: DeviceAddress) Request {
pub fn setAddress(address: DeviceAddress) Request {
return .{
.request_type = .{
.recipient = .device,
@@ -372,7 +377,7 @@ fn setAddress(address: DeviceAddress) Request {
// - language_id selects the language of a string descriptor, and is zero otherwise.
// - length is the number of bytes to read; a device never returns more than length bytes,
// but may return less if the descriptor is shorter.
fn getDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
pub fn getDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
return .{
.request_type = .{
.recipient = .device,
@@ -389,7 +394,7 @@ fn getDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, l
// Updates an existing descriptor or adds a new one (optional; many devices do not support
// this request). The parameters mirror getDescriptor; the descriptor itself is sent in the
// DATA stage.
fn setDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
pub fn setDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
return .{
.request_type = .{
.recipient = .device,
@@ -405,7 +410,7 @@ fn setDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, l
// Reads the currently active configuration: @enumFromInt the byte the device returns into a
// ConfigurationValue, which is none while the device is not configured.
fn getConfiguration() Request {
pub fn getConfiguration() Request {
return .{
.request_type = .{
.recipient = .device,
@@ -422,7 +427,7 @@ fn getConfiguration() Request {
// Selects the configuration with the given configuration_value (from
// ConfigurationDescriptor.configuration_value), moving the device from the address state to
// the configured state. Selecting none returns the device to the address state.
fn setConfiguration(configuration_value: ConfigurationValue) Request {
pub fn setConfiguration(configuration_value: ConfigurationValue) Request {
return .{
.request_type = .{
.recipient = .device,
@@ -438,7 +443,7 @@ fn setConfiguration(configuration_value: ConfigurationValue) Request {
// Reads the alternate setting currently selected for the given interface: @enumFromInt the
// byte the device returns into an AlternateSetting.
fn getInterface(interface: InterfaceNumber) Request {
pub fn getInterface(interface: InterfaceNumber) Request {
return .{
.request_type = .{
.recipient = .interface,
@@ -454,7 +459,7 @@ fn getInterface(interface: InterfaceNumber) Request {
// Selects an alternate setting (from InterfaceDescriptor.alternate_setting) for the given
// interface.
fn setInterface(interface: InterfaceNumber, alternate_setting: AlternateSetting) Request {
pub fn setInterface(interface: InterfaceNumber, alternate_setting: AlternateSetting) Request {
return .{
.request_type = .{
.recipient = .interface,
@@ -470,7 +475,7 @@ fn setInterface(interface: InterfaceNumber, alternate_setting: AlternateSetting)
// Reads the two-byte number of the frame in which the given isochronous endpoint's
// repeating pattern of transfers begins.
fn syncFrame(endpoint: Request.EndpointIndex) Request {
pub fn syncFrame(endpoint: Request.EndpointIndex) Request {
return .{
.request_type = .{
.recipient = .endpoint,
@@ -484,7 +489,79 @@ fn syncFrame(endpoint: Request.EndpointIndex) Request {
};
}
const DescriptorType = enum(u8) {
// Class-specific requests. These carry a `kind = .class` request_type and a
// request_code from the interface's class namespace (not the standard
// RequestCode set above); the code is written into the same byte field, which
// is why RequestCode is non-exhaustive. Each is directed at an interface, whose
// number rides in the index field.
// The HID class request codes (USB HID 1.11 §7.2). Only the ones danos issues
// are named; the field on the wire is the raw byte.
pub const HidRequestCode = enum(u8) {
get_report = 0x01,
get_idle = 0x02,
get_protocol = 0x03,
set_report = 0x09,
set_idle = 0x0A,
set_protocol = 0x0B,
};
// The two protocols a boot-capable HID device can run (USB HID 1.11 §7.2.5).
// A driver selects `boot` for the simplified fixed-format boot report, usable
// before a full report-descriptor parser exists.
pub const HidProtocol = enum(u8) {
boot = 0,
report = 1,
};
// SET_PROTOCOL: choose the boot or report protocol on a HID interface.
pub fn setProtocol(interface: InterfaceNumber, protocol: HidProtocol) Request {
return .{
.request_type = .{ .recipient = .interface, .kind = .class, .direction = .host_to_device },
.request_code = @enumFromInt(@intFromEnum(HidRequestCode.set_protocol)),
.value = @intFromEnum(protocol),
.index = @intFromEnum(interface),
.length = 0,
};
}
// SET_IDLE: bound a HID interface's report rate. `duration` is in 4 ms units
// (0 means report only on change); `report_id` selects a report (0 = all).
pub fn setIdle(interface: InterfaceNumber, duration: u8, report_id: u8) Request {
return .{
.request_type = .{ .recipient = .interface, .kind = .class, .direction = .host_to_device },
.request_code = @enumFromInt(@intFromEnum(HidRequestCode.set_idle)),
.value = (@as(u16, duration) << 8) | report_id,
.index = @intFromEnum(interface),
.length = 0,
};
}
// Bulk-Only Mass Storage Reset (USB MSC BOT §3.1): ready a mass-storage
// interface for the next Command Block Wrapper after a protocol error.
pub fn bulkOnlyMassStorageReset(interface: InterfaceNumber) Request {
return .{
.request_type = .{ .recipient = .interface, .kind = .class, .direction = .host_to_device },
.request_code = @enumFromInt(0xFF),
.value = 0,
.index = @intFromEnum(interface),
.length = 0,
};
}
// Get Max LUN (USB MSC BOT §3.2): read the highest logical unit number the
// device supports (0 for a single-LUN flash drive). One byte is returned.
pub fn getMaxLun(interface: InterfaceNumber) Request {
return .{
.request_type = .{ .recipient = .interface, .kind = .class, .direction = .device_to_host },
.request_code = @enumFromInt(0xFE),
.value = 0,
.index = @intFromEnum(interface),
.length = 1,
};
}
pub const DescriptorType = enum(u8) {
device = 1,
configuration = 2,
string = 3,
@@ -496,7 +573,7 @@ const DescriptorType = enum(u8) {
_,
};
const DeviceDescriptor = extern struct {
pub const DeviceDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// DEVICE Descriptor Type
@@ -541,7 +618,7 @@ const DeviceDescriptor = extern struct {
configuration_count: u8,
};
const DeviceQualifierDescriptor = extern struct {
pub const DeviceQualifierDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// DEVICE_QUALIFIER Descriptor Type
@@ -564,7 +641,7 @@ const DeviceQualifierDescriptor = extern struct {
reserved: u8,
};
const ConfigurationDescriptor = extern struct {
pub const ConfigurationDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// CONFIGURATION Descriptor Type
@@ -593,7 +670,7 @@ const ConfigurationDescriptor = extern struct {
max_power: u8,
// Configuration characteristics. Fields are declared least-significant first.
const Attributes = packed struct(u8) {
pub const Attributes = packed struct(u8) {
// Reserved, reset to zero (D4...0)
reserved: u5,
// Whether Remote Wakeup is supported by this configuration (D5)
@@ -612,9 +689,9 @@ const ConfigurationDescriptor = extern struct {
// its alternative speed. The structure of the OTHER_SPEED_CONFIGURATION is identical to that
// of the CONFIGURATION descriptor; the only difference is that the descriptor_type field
// reflects that the descriptor is an OTHER_SPEED_CONFIGURATION descriptor.
const OtherSpeedConfigurationDescriptor = ConfigurationDescriptor;
pub const OtherSpeedConfigurationDescriptor = ConfigurationDescriptor;
const InterfaceDescriptor = extern struct {
pub const InterfaceDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// INTERFACE Descriptor Type
@@ -654,7 +731,7 @@ const InterfaceDescriptor = extern struct {
interface_index: StringIndex,
};
const EndpointDescriptor = extern struct {
pub const EndpointDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// ENDPOINT Descriptor Type
@@ -683,7 +760,7 @@ const EndpointDescriptor = extern struct {
interval: u8,
// The address of an endpoint. Fields are declared least-significant first.
const Address = packed struct(u8) {
pub const Address = packed struct(u8) {
// Endpoint Number (D3...0)
number: EndpointNumber,
// Reserved, reset to zero (D6...4)
@@ -693,7 +770,7 @@ const EndpointDescriptor = extern struct {
};
// An endpoint's attributes. Fields are declared least-significant first.
const Attributes = packed struct(u8) {
pub const Attributes = packed struct(u8) {
// Transfer Type (D1...0)
transfer_type: TransferType,
// Synchronization Type; isochronous endpoints only, reserved and reset to zero for
@@ -706,21 +783,21 @@ const EndpointDescriptor = extern struct {
reserved: u2,
};
const TransferType = enum(u2) {
pub const TransferType = enum(u2) {
control = 0,
isochronous = 1,
bulk = 2,
interrupt = 3,
};
const Synchronization = enum(u2) {
pub const Synchronization = enum(u2) {
none = 0,
asynchronous = 1,
adaptive = 2,
synchronous = 3,
};
const Usage = enum(u2) {
pub const Usage = enum(u2) {
data = 0,
feedback = 1,
implicit_feedback_data = 2,
@@ -728,7 +805,7 @@ const EndpointDescriptor = extern struct {
};
// The maximum packet size of an endpoint. Fields are declared least-significant first.
const MaxPacketSize = packed struct(u16) {
pub const MaxPacketSize = packed struct(u16) {
// Maximum packet size in bytes (bits 10...0)
size: u11,
// Number of additional transaction opportunities per microframe, for high-speed
@@ -739,7 +816,7 @@ const EndpointDescriptor = extern struct {
reserved: u3,
};
const AdditionalTransactions = enum(u2) {
pub const AdditionalTransactions = enum(u2) {
// None (1 transaction per microframe)
none = 0,
// 1 additional (2 transactions per microframe)
@@ -755,7 +832,7 @@ const EndpointDescriptor = extern struct {
// header, followed by the variable-length payload:
// - index 0: an array of two-byte LANGID codes (wLangID[0] through wLangID[x])
// - other indices: a Unicode string of N bytes
const StringDescriptor = extern struct {
pub const StringDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// STRING Descriptor Type
@@ -834,7 +911,7 @@ test "bitmap packings match the specification" {
try expect(hid_type != .device);
}
fn expectRequestBytes(request: Request, expected: [8]u8) !void {
pub fn expectRequestBytes(request: Request, expected: [8]u8) !void {
try std.testing.expectEqualSlices(u8, &expected, std.mem.asBytes(&request));
}
@@ -855,3 +932,14 @@ test "standard request constructors encode the specification's set-up packets" {
try expectRequestBytes(setInterface(@enumFromInt(2), @enumFromInt(1)), .{ 0x01, 11, 1, 0, 2, 0, 0, 0 });
try expectRequestBytes(syncFrame(.{ .number = @enumFromInt(3), .direction = .in }), .{ 0x82, 12, 0, 0, 0x83, 0, 2, 0 });
}
test "class request constructors encode the specification's set-up packets" {
// bmRequestType for a host-to-device class request to an interface = 0x21;
// device-to-host = 0xA1. The request_code byte is the class code, not a
// standard one — SET_PROTOCOL 0x0B, SET_IDLE 0x0A, BOT reset 0xFF, Max LUN 0xFE.
try expectRequestBytes(setProtocol(@enumFromInt(0), .boot), .{ 0x21, 0x0B, 0, 0, 0, 0, 0, 0 });
try expectRequestBytes(setProtocol(@enumFromInt(1), .report), .{ 0x21, 0x0B, 1, 0, 1, 0, 0, 0 });
try expectRequestBytes(setIdle(@enumFromInt(1), 0, 0), .{ 0x21, 0x0A, 0, 0, 1, 0, 0, 0 });
try expectRequestBytes(bulkOnlyMassStorageReset(@enumFromInt(0)), .{ 0x21, 0xFF, 0, 0, 0, 0, 0, 0 });
try expectRequestBytes(getMaxLun(@enumFromInt(0)), .{ 0xA1, 0xFE, 0, 0, 0, 0, 1, 0 });
}
+62 -19
View File
@@ -11,7 +11,7 @@
// Base class codes (assigned by the USB-IF). The comment on each value notes where the code
// may legally appear: in the device descriptor, in interface descriptors, or both.
const Class = enum(u8) {
pub const Class = enum(u8) {
// Use class information in the interface descriptors (device descriptor only). Each
// interface within a configuration specifies its own class information and the various
// interfaces operate independently.
@@ -72,8 +72,8 @@ const Class = enum(u8) {
// Subclass and protocol codes qualified by Class.hub. Hubs have no subclass codes; the
// protocol distinguishes the hub's transaction-translator arrangement.
const hub = struct {
const Protocol = enum(u8) {
pub const hub = struct {
pub const Protocol = enum(u8) {
// Full-speed hub
full_speed = 0x00,
// Hi-speed hub with a single transaction translator
@@ -87,8 +87,8 @@ const hub = struct {
};
// Subclass and protocol codes qualified by Class.hid.
const hid = struct {
const SubClass = enum(u8) {
pub const hid = struct {
pub const SubClass = enum(u8) {
// No subclass
none = 0x00,
// Boot interface: the device also supports the simplified boot protocol, usable by
@@ -98,7 +98,7 @@ const hid = struct {
};
// Only meaningful when the subclass is boot
const Protocol = enum(u8) {
pub const Protocol = enum(u8) {
none = 0x00,
keyboard = 0x01,
mouse = 0x02,
@@ -109,8 +109,8 @@ const hid = struct {
// Subclass and protocol codes qualified by Class.mass_storage. The subclass identifies the
// command set the device understands; the protocol identifies the transport used to carry
// commands, data, and status over the bus.
const mass_storage = struct {
const SubClass = enum(u8) {
pub const mass_storage = struct {
pub const SubClass = enum(u8) {
// SCSI command set not reported; de facto, treat as scsi
not_reported = 0x00,
// Reduced Block Commands: typically flash devices
@@ -134,7 +134,7 @@ const mass_storage = struct {
_,
};
const Protocol = enum(u8) {
pub const Protocol = enum(u8) {
// Control/Bulk/Interrupt with command completion interrupt
cbi_completion_interrupt = 0x00,
// Control/Bulk/Interrupt without command completion interrupt
@@ -152,8 +152,8 @@ const mass_storage = struct {
// Subclass and protocol codes qualified by Class.communications (CDC). The protocol codes
// are model-specific; the useful invariant is the subclass, which selects the control model
// the interface implements.
const communications = struct {
const SubClass = enum(u8) {
pub const communications = struct {
pub const SubClass = enum(u8) {
// Direct line control model
direct_line = 0x01,
// Abstract control model: USB modems and serial adapters
@@ -185,15 +185,15 @@ const communications = struct {
};
// Subclass and protocol codes qualified by Class.wireless_controller.
const wireless_controller = struct {
const SubClass = enum(u8) {
pub const wireless_controller = struct {
pub const SubClass = enum(u8) {
// Radio frequency controllers
radio_frequency = 0x01,
_,
};
// Only meaningful when the subclass is radio_frequency
const Protocol = enum(u8) {
pub const Protocol = enum(u8) {
// Bluetooth programming interface
bluetooth = 0x01,
// Ultra-wideband radio control
@@ -207,15 +207,15 @@ const wireless_controller = struct {
};
// Subclass and protocol codes qualified by Class.miscellaneous.
const miscellaneous = struct {
const SubClass = enum(u8) {
pub const miscellaneous = struct {
pub const SubClass = enum(u8) {
// Common class
common = 0x02,
_,
};
// Only meaningful when the subclass is common
const Protocol = enum(u8) {
pub const Protocol = enum(u8) {
// Interface association descriptor: at the device level, announces that the
// configuration groups interfaces into functions with IADs
interface_association = 0x01,
@@ -224,8 +224,8 @@ const miscellaneous = struct {
};
// Subclass and protocol codes qualified by Class.application_specific.
const application_specific = struct {
const SubClass = enum(u8) {
pub const application_specific = struct {
pub const SubClass = enum(u8) {
// Device firmware upgrade
firmware_upgrade = 0x01,
// IrDA bridge
@@ -236,6 +236,23 @@ const application_specific = struct {
};
};
/// Pack a (class, subclass, protocol) triple into one 0xCCSSPP value — the
/// bus-native identity a USB bus driver reports in `ChildAdded.identity` and the
/// device manager matches on (the USB analog of a packed PCI class code). Mirrors
/// `pci_class.ClassCode.pack`, so both sides build/decode the identical u64.
pub fn packTriple(class: u8, subclass: u8, protocol: u8) u64 {
return (@as(u64, class) << 16) | (@as(u64, subclass) << 8) | protocol;
}
/// The inverse of `packTriple`.
pub fn unpackTriple(triple: u64) struct { class: u8, subclass: u8, protocol: u8 } {
return .{
.class = @truncate(triple >> 16),
.subclass = @truncate(triple >> 8),
.protocol = @truncate(triple),
};
}
test "class codes match the USB-IF assignments" {
const std = @import("std");
const expectEqual = std.testing.expectEqual;
@@ -262,3 +279,29 @@ test "class codes match the USB-IF assignments" {
_ = miscellaneous.Protocol.interface_association;
_ = application_specific.SubClass.firmware_upgrade;
}
test "packTriple / unpackTriple round-trip the identity a bus driver reports" {
const std = @import("std");
const expectEqual = std.testing.expectEqual;
// A boot keyboard interface: HID / boot / keyboard.
const keyboard = packTriple(
@intFromEnum(Class.hid),
@intFromEnum(hid.SubClass.boot),
@intFromEnum(hid.Protocol.keyboard),
);
try expectEqual(@as(u64, 0x03_01_01), keyboard);
// A flash drive interface: mass storage / SCSI / bulk-only.
const storage = packTriple(
@intFromEnum(Class.mass_storage),
@intFromEnum(mass_storage.SubClass.scsi),
@intFromEnum(mass_storage.Protocol.bulk_only),
);
try expectEqual(@as(u64, 0x08_06_50), storage);
const parts = unpackTriple(storage);
try expectEqual(@as(u8, 0x08), parts.class);
try expectEqual(@as(u8, 0x06), parts.subclass);
try expectEqual(@as(u8, 0x50), parts.protocol);
}
+191
View File
@@ -0,0 +1,191 @@
//! Pure decoders for USB HID **boot-protocol** reports — the simplified,
//! fixed-format reports a boot keyboard and boot mouse send, the USB analog of
//! the PS/2 scancode and mouse-packet decoders. No I/O: these turn report bytes
//! into make/break transitions and motion, which the usb-hid drivers publish to
//! the input service. Host-testable in isolation (like mouse-packet.zig).
//!
//! "Boot protocol" is a USB HID term (USB HID 1.11 §B) — the device reports in
//! this fixed layout after SET_PROTOCOL(boot); it has nothing to do with system
//! boot.
const std = @import("std");
// --- keyboard ---------------------------------------------------------------
/// The 8-byte boot keyboard report: a modifier bitmap, a reserved byte, and up
/// to six concurrently-pressed key usages.
pub const KeyboardReport = extern struct {
modifiers: u8 = 0,
reserved: u8 = 0,
keys: [6]u8 = .{ 0, 0, 0, 0, 0, 0 },
};
// The modifier byte's bits (HID keyboard boot report).
pub const modifier_left_control: u8 = 1 << 0;
pub const modifier_left_shift: u8 = 1 << 1;
pub const modifier_left_alt: u8 = 1 << 2;
pub const modifier_left_gui: u8 = 1 << 3;
pub const modifier_right_control: u8 = 1 << 4;
pub const modifier_right_shift: u8 = 1 << 5;
pub const modifier_right_alt: u8 = 1 << 6;
pub const modifier_right_gui: u8 = 1 << 7;
pub const TransitionKind = enum { pressed, released };
/// One key going down or up. `usage` is a HID keyboard-page usage — modifier keys
/// map to usages 224..231 — which is exactly the input protocol's `Keycode`.
pub const Transition = struct { kind: TransitionKind, usage: u8 };
// A report can change at most all 8 modifiers and all 6 keys at once.
pub const max_transitions = 8 + 6;
pub const Transitions = struct {
items: [max_transitions]Transition = undefined,
count: usize = 0,
fn add(self: *Transitions, transition: Transition) void {
if (self.count < self.items.len) {
self.items[self.count] = transition;
self.count += 1;
}
}
pub fn slice(self: *const Transitions) []const Transition {
return self.items[0..self.count];
}
};
/// Turns a stream of boot keyboard reports into make/break transitions by diffing
/// each report against the last.
pub const KeyboardDecoder = struct {
previous: KeyboardReport = .{},
pub fn feed(self: *KeyboardDecoder, current: KeyboardReport) Transitions {
var out = Transitions{};
// Rollover: 0x01 (ErrorRollOver) means more keys are held than the report
// can carry, so the key array is invalid. Emit nothing and keep the prior
// state (so the eventual releases still resolve against real keys).
for (current.keys) |key| {
if (key == 0x01) return out;
}
// Modifiers: one make/break per changed bit; modifier usages are 224..231.
const changed = current.modifiers ^ self.previous.modifiers;
var bit: u3 = 0;
while (true) : (bit += 1) {
const mask = @as(u8, 1) << bit;
if (changed & mask != 0) {
out.add(.{
.kind = if (current.modifiers & mask != 0) .pressed else .released,
.usage = 224 + @as(u8, bit),
});
}
if (bit == 7) break;
}
// Keys made: present now, absent before.
for (current.keys) |key| {
if (key != 0 and !contains(&self.previous.keys, key)) out.add(.{ .kind = .pressed, .usage = key });
}
// Keys broken: present before, absent now.
for (self.previous.keys) |key| {
if (key != 0 and !contains(&current.keys, key)) out.add(.{ .kind = .released, .usage = key });
}
self.previous = current;
return out;
}
};
fn contains(keys: *const [6]u8, value: u8) bool {
for (keys) |key| {
if (key == value) return true;
}
return false;
}
// --- mouse ------------------------------------------------------------------
/// A decoded boot mouse report: the button bitmap and relative motion. The wheel
/// byte is present only on 4-byte reports (QEMU's usb-mouse sends one).
pub const MouseReport = struct {
buttons: u8 = 0,
dx: i8 = 0,
dy: i8 = 0,
wheel: i8 = 0,
has_wheel: bool = false,
};
pub const mouse_button_left: u8 = 1 << 0;
pub const mouse_button_right: u8 = 1 << 1;
pub const mouse_button_middle: u8 = 1 << 2;
/// Parse a 3- or 4-byte boot mouse report. Note HID reports Y in screen
/// convention (positive = down), so — unlike PS/2 — `dy` is NOT negated.
pub fn parseMouse(bytes: []const u8) ?MouseReport {
if (bytes.len < 3) return null;
return .{
.buttons = bytes[0],
.dx = @bitCast(bytes[1]),
.dy = @bitCast(bytes[2]),
.wheel = if (bytes.len >= 4) @bitCast(bytes[3]) else 0,
.has_wheel = bytes.len >= 4,
};
}
// --- tests ------------------------------------------------------------------
test "keyboard diff produces make and break transitions" {
var decoder = KeyboardDecoder{};
// Press 'a' (usage 4).
var t = decoder.feed(.{ .keys = .{ 4, 0, 0, 0, 0, 0 } });
try std.testing.expectEqual(@as(usize, 1), t.count);
try std.testing.expectEqual(TransitionKind.pressed, t.items[0].kind);
try std.testing.expectEqual(@as(u8, 4), t.items[0].usage);
// Hold 'a', press 'b' (usage 5): only 'b' is new.
t = decoder.feed(.{ .keys = .{ 4, 5, 0, 0, 0, 0 } });
try std.testing.expectEqual(@as(usize, 1), t.count);
try std.testing.expectEqual(@as(u8, 5), t.items[0].usage);
// Release everything: 'a' and 'b' both break.
t = decoder.feed(.{ .keys = .{ 0, 0, 0, 0, 0, 0 } });
try std.testing.expectEqual(@as(usize, 2), t.count);
try std.testing.expectEqual(TransitionKind.released, t.items[0].kind);
// Press Left Shift (modifier bit 1 -> usage 225).
t = decoder.feed(.{ .modifiers = modifier_left_shift });
try std.testing.expectEqual(@as(usize, 1), t.count);
try std.testing.expectEqual(@as(u8, 225), t.items[0].usage);
try std.testing.expectEqual(TransitionKind.pressed, t.items[0].kind);
}
test "rollover report is ignored but state is preserved" {
var decoder = KeyboardDecoder{};
_ = decoder.feed(.{ .keys = .{ 4, 0, 0, 0, 0, 0 } }); // press 'a'
const rollover = decoder.feed(.{ .keys = .{ 0x01, 0x01, 0x01, 0x01, 0x01, 0x01 } });
try std.testing.expectEqual(@as(usize, 0), rollover.count);
// 'a' is still considered down, so releasing all keys now breaks it.
const release = decoder.feed(.{ .keys = .{ 0, 0, 0, 0, 0, 0 } });
try std.testing.expectEqual(@as(usize, 1), release.count);
try std.testing.expectEqual(@as(u8, 4), release.items[0].usage);
try std.testing.expectEqual(TransitionKind.released, release.items[0].kind);
}
test "mouse report parses motion without inverting Y" {
const three = parseMouse(&.{ mouse_button_left, 5, 0xFB }).?; // dy = -5
try std.testing.expectEqual(mouse_button_left, three.buttons);
try std.testing.expectEqual(@as(i8, 5), three.dx);
try std.testing.expectEqual(@as(i8, -5), three.dy);
try std.testing.expect(!three.has_wheel);
const four = parseMouse(&.{ 0, 0, 0, 0xFF }).?; // wheel = -1
try std.testing.expect(four.has_wheel);
try std.testing.expectEqual(@as(i8, -1), four.wheel);
try std.testing.expect(parseMouse(&.{ 0, 0 }) == null); // too short
}
+174
View File
@@ -0,0 +1,174 @@
//! USB HID boot keyboard driver.
//!
//! Spawned by the device manager when the xHCI bus driver reports a HID / boot /
//! keyboard interface (class 3, subclass 1, protocol 1); its assigned device id
//! arrives as argv[1] and an optional layout name ("us", "gb", ...) as argv[2].
//! It owns no hardware: it opens its device through the USB transfer protocol
//! (`runtime.usb`), asks the device for the boot protocol, subscribes to its
//! interrupt-IN endpoint, and turns each 8-byte boot report into input-protocol
//! events, published to the input service — the USB analogue of ps2-bus/keyboard.
//!
//! interrupt report -> hid-report diff -> key_down / key_up
//! -> xkeyboard-config -> character -> key_press
//!
//! Because a USB keyboard's usages ARE the input protocol's keycodes (both are
//! HID keyboard page 0x07), the decode is nearly 1:1 — no scancode translation.
const std = @import("std");
const runtime = @import("runtime");
const usb_abi = @import("usb-abi");
const xkb = @import("xkeyboard-config");
const hid = @import("hid-report.zig");
const ipc = runtime.ipc;
const process = runtime.process;
const input_protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// The modifier state a character lookup needs — derived from the report's
// modifier byte, plus the driver-tracked caps-lock toggle.
const ModifierSnapshot = struct {
shift: bool,
control: bool,
right_alt: bool,
caps_lock: bool,
};
/// The character a key produces under `modifiers`, or 0 for none — the layout
/// lookup for printable keys, with ASCII control characters for the keys every
/// consumer expects (Enter, Tab, Backspace, Escape), exactly as ps2-bus/keyboard.
fn characterFor(layout: *const xkb.Layout, usage: u8, modifiers: ModifierSnapshot) u32 {
const mapping = xkb.map(layout, usage, .{
.shift = modifiers.shift,
.caps_lock = modifiers.caps_lock,
.level3 = modifiers.right_alt,
.control = modifiers.control,
});
if (mapping.character) |character| return character;
return switch (@as(input_protocol.Keycode, @enumFromInt(usage))) {
.enter, .keypad_enter => '\n',
.tab => '\t',
.backspace => 0x08,
.escape => 0x1B,
else => 0,
};
}
fn modifierWord(modifiers: u8) u32 {
var word: u32 = 0;
if (modifiers & (hid.modifier_left_shift | hid.modifier_right_shift) != 0) word |= input_protocol.modifier_shift;
if (modifiers & (hid.modifier_left_control | hid.modifier_right_control) != 0) word |= input_protocol.modifier_control;
if (modifiers & (hid.modifier_left_alt | hid.modifier_right_alt) != 0) word |= input_protocol.modifier_alt;
return word;
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("/system/drivers/usb-hid/keyboard: missing device id (argv[1])\n");
return;
};
const device_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/usb-hid/keyboard: malformed device id '{s}'\n", .{argument});
return;
};
const layout = xkb.byName(init.arguments.get(2) orelse "us") orelse xkb.us;
// Hello the manager first (meet the spawn deadline), then open the device.
if (!runtime.usb.helloManager(device_id)) {
_ = runtime.system.write("/system/drivers/usb-hid/keyboard: hello to device manager failed\n");
return;
}
var device = runtime.usb.open(device_id) orelse {
writeLine("/system/drivers/usb-hid/keyboard: could not open device {d}\n", .{device_id});
return;
};
const endpoint = device.findEndpoint(runtime.usb.transfer_type_interrupt, true) orelse {
_ = runtime.system.write("/system/drivers/usb-hid/keyboard: no interrupt-IN endpoint\n");
return;
};
// Ask for the boot protocol and an indefinite idle (report only on change).
_ = device.controlOut(@bitCast(usb_abi.setProtocol(@enumFromInt(device.interface_number), .boot)));
_ = device.controlOut(@bitCast(usb_abi.setIdle(@enumFromInt(device.interface_number), 0, 0)));
if (!device.subscribeInterrupt(endpoint.address, endpoint.max_packet_size)) {
_ = runtime.system.write("/system/drivers/usb-hid/keyboard: interrupt subscribe failed\n");
return;
}
var source = runtime.input.connectSource() orelse {
_ = runtime.system.write("/system/drivers/usb-hid/keyboard: input service unavailable\n");
return;
};
_ = process.bindSignals(device.endpoint);
writeLine("/system/drivers/usb-hid/keyboard: ok (device {d}, interface {d}, layout {s})\n", .{ device_id, device.interface_number, layout.name });
var decoder = hid.KeyboardDecoder{};
var caps_lock = false;
var receive: [64]u8 = undefined;
while (true) {
const got = ipc.replyWait(device.endpoint, &.{}, &receive, null);
if (!got.isNotification()) continue;
if (process.signalsFrom(got.badge)) |signals| {
if (signals.has(.terminate)) return;
continue;
}
if (!got.isMessage() or got.len < @sizeOf(runtime.usb.InterruptReport)) continue;
const message = std.mem.bytesToValue(runtime.usb.InterruptReport, receive[0..@sizeOf(runtime.usb.InterruptReport)]);
if (message.length < @sizeOf(hid.KeyboardReport)) continue;
const report = std.mem.bytesToValue(hid.KeyboardReport, message.data[0..@sizeOf(hid.KeyboardReport)]);
const transitions = decoder.feed(report);
// Caps Lock toggles on its own key-down (a stateful lock, not a modifier).
for (transitions.slice()) |transition| {
if (transition.kind == .pressed and @as(input_protocol.Keycode, @enumFromInt(transition.usage)) == .caps_lock) caps_lock = !caps_lock;
}
const modifiers = ModifierSnapshot{
.shift = report.modifiers & (hid.modifier_left_shift | hid.modifier_right_shift) != 0,
.control = report.modifiers & (hid.modifier_left_control | hid.modifier_right_control) != 0,
.right_alt = report.modifiers & hid.modifier_right_alt != 0,
.caps_lock = caps_lock,
};
const modifier_word = modifierWord(report.modifiers);
for (transitions.slice()) |transition| {
switch (transition.kind) {
.pressed => {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(input_protocol.EventKind.key_down),
.keycode = transition.usage,
.character = 0,
.modifiers = modifier_word,
});
const character = characterFor(layout, transition.usage, modifiers);
if (character != 0) {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(input_protocol.EventKind.key_press),
.keycode = transition.usage,
.character = character,
.modifiers = modifier_word,
});
}
},
.released => {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(input_protocol.EventKind.key_up),
.keycode = transition.usage,
.character = 0,
.modifiers = modifier_word,
});
},
}
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+140
View File
@@ -0,0 +1,140 @@
//! USB HID boot mouse driver.
//!
//! Spawned by the device manager when the xHCI bus driver reports a HID / boot /
//! mouse interface (class 3, subclass 1, protocol 2); its assigned device id
//! arrives as argv[1]. Like the keyboard driver it owns no hardware: it opens its
//! device through the USB transfer protocol (`runtime.usb`), asks for the boot
//! protocol, subscribes to its interrupt-IN endpoint, and turns each 3- or 4-byte
//! boot report into input-protocol mouse events published to the input service.
//!
//! Unlike PS/2, HID reports Y in screen convention (positive = down), so motion
//! is passed straight through (the decode in hid-report.zig does not negate it).
const std = @import("std");
const runtime = @import("runtime");
const usb_abi = @import("usb-abi");
const hid = @import("hid-report.zig");
const ipc = runtime.ipc;
const process = runtime.process;
const input_protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// The current pressed-button bitmask in input-protocol terms.
fn buttonMask(buttons: u8) u32 {
var mask: u32 = 0;
if (buttons & hid.mouse_button_left != 0) mask |= input_protocol.mouse_button_left;
if (buttons & hid.mouse_button_right != 0) mask |= input_protocol.mouse_button_right;
if (buttons & hid.mouse_button_middle != 0) mask |= input_protocol.mouse_button_middle;
return mask;
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("/system/drivers/usb-hid/mouse: missing device id (argv[1])\n");
return;
};
const device_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/usb-hid/mouse: malformed device id '{s}'\n", .{argument});
return;
};
if (!runtime.usb.helloManager(device_id)) {
_ = runtime.system.write("/system/drivers/usb-hid/mouse: hello to device manager failed\n");
return;
}
var device = runtime.usb.open(device_id) orelse {
writeLine("/system/drivers/usb-hid/mouse: could not open device {d}\n", .{device_id});
return;
};
const endpoint = device.findEndpoint(runtime.usb.transfer_type_interrupt, true) orelse {
_ = runtime.system.write("/system/drivers/usb-hid/mouse: no interrupt-IN endpoint\n");
return;
};
_ = device.controlOut(@bitCast(usb_abi.setProtocol(@enumFromInt(device.interface_number), .boot)));
if (!device.subscribeInterrupt(endpoint.address, endpoint.max_packet_size)) {
_ = runtime.system.write("/system/drivers/usb-hid/mouse: interrupt subscribe failed\n");
return;
}
var source = runtime.input.connectSource() orelse {
_ = runtime.system.write("/system/drivers/usb-hid/mouse: input service unavailable\n");
return;
};
_ = process.bindSignals(device.endpoint);
writeLine("/system/drivers/usb-hid/mouse: ok (device {d}, interface {d})\n", .{ device_id, device.interface_number });
var previous_buttons: u8 = 0;
var receive: [64]u8 = undefined;
while (true) {
const got = ipc.replyWait(device.endpoint, &.{}, &receive, null);
if (!got.isNotification()) continue;
if (process.signalsFrom(got.badge)) |signals| {
if (signals.has(.terminate)) return;
continue;
}
if (!got.isMessage() or got.len < @sizeOf(runtime.usb.InterruptReport)) continue;
const message = std.mem.bytesToValue(runtime.usb.InterruptReport, receive[0..@sizeOf(runtime.usb.InterruptReport)]);
const length = @min(message.length, message.data.len);
const report = hid.parseMouse(message.data[0..length]) orelse continue;
const mask = buttonMask(report.buttons);
// Button transitions: one event per changed button bit.
const changed = report.buttons ^ previous_buttons;
inline for (.{
.{ hid.mouse_button_left, input_protocol.mouse_button_left },
.{ hid.mouse_button_right, input_protocol.mouse_button_right },
.{ hid.mouse_button_middle, input_protocol.mouse_button_middle },
}) |pair| {
if (changed & pair[0] != 0) {
_ = source.publishMouseEvent(.{
.kind = @intFromEnum(if (report.buttons & pair[0] != 0) input_protocol.MouseEventKind.button_down else input_protocol.MouseEventKind.button_up),
.button = pair[1],
.dx = 0,
.dy = 0,
.scroll_x = 0,
.scroll_y = 0,
.buttons = mask,
});
}
}
previous_buttons = report.buttons;
// Relative motion (dy straight through — HID Y is already screen convention).
if (report.dx != 0 or report.dy != 0) {
_ = source.publishMouseEvent(.{
.kind = @intFromEnum(input_protocol.MouseEventKind.motion),
.button = 0,
.dx = report.dx,
.dy = report.dy,
.scroll_x = 0,
.scroll_y = 0,
.buttons = mask,
});
}
// Wheel (4-byte reports only): positive = scroll up.
if (report.has_wheel and report.wheel != 0) {
_ = source.publishMouseEvent(.{
.kind = @intFromEnum(input_protocol.MouseEventKind.scroll),
.button = 0,
.dx = 0,
.dy = 0,
.scroll_x = 0,
.scroll_y = report.wheel,
.buttons = mask,
});
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
@@ -0,0 +1,73 @@
//! USB Mass Storage Bulk-Only Transport (BOT) wire structures — the Command and
//! Command Status Wrappers that bracket every command (USB MSC BOT §5). Pure data
//! definitions, host-testable in isolation. The command inside the CBW is a SCSI
//! CDB (see scsi.zig); the transport here just carries it and reports status.
//!
//! One command is three bulk transfers: CBW out, an optional data stage, CSW in.
const std = @import("std");
/// "USBC" — the signature at the head of every Command Block Wrapper.
pub const cbw_signature: u32 = 0x43425355;
/// "USBS" — the signature at the head of every Command Status Wrapper.
pub const csw_signature: u32 = 0x53425355;
/// CBW `flags`: set for a device-to-host (IN) data stage, clear for OUT.
pub const flag_data_in: u8 = 0x80;
/// The 31-byte Command Block Wrapper, sent on the bulk-OUT endpoint.
pub const CommandBlockWrapper = extern struct {
signature: u32 align(1) = cbw_signature,
tag: u32 align(1),
data_transfer_length: u32 align(1),
flags: u8,
lun: u8,
cdb_length: u8,
cdb: [16]u8 = [_]u8{0} ** 16,
};
/// A device's answer to a command (the CSW `status` byte).
pub const CommandStatus = enum(u8) {
passed = 0,
failed = 1,
phase_error = 2,
_,
};
/// The 13-byte Command Status Wrapper, read from the bulk-IN endpoint.
pub const CommandStatusWrapper = extern struct {
signature: u32 align(1) = csw_signature,
tag: u32 align(1),
data_residue: u32 align(1),
status: u8,
};
comptime {
std.debug.assert(@sizeOf(CommandBlockWrapper) == 31);
std.debug.assert(@sizeOf(CommandStatusWrapper) == 13);
}
test "wrapper sizes and signatures match the specification" {
const cbw = CommandBlockWrapper{
.tag = 0x11223344,
.data_transfer_length = 512,
.flags = flag_data_in,
.lun = 0,
.cdb_length = 10,
};
const bytes = std.mem.asBytes(&cbw);
try std.testing.expectEqual(@as(usize, 31), bytes.len);
// "USBC" little-endian.
try std.testing.expectEqualSlices(u8, "USBC", bytes[0..4]);
try std.testing.expectEqual(flag_data_in, bytes[12]);
const csw = std.mem.bytesToValue(CommandStatusWrapper, &[_]u8{
0x55, 0x53, 0x42, 0x53, // "USBS"
0x44, 0x33, 0x22, 0x11, // tag
0x00, 0x00, 0x00, 0x00, // residue
0x00, // passed
});
try std.testing.expectEqual(csw_signature, csw.signature);
try std.testing.expectEqual(@as(u32, 0x11223344), csw.tag);
try std.testing.expectEqual(@as(u8, @intFromEnum(CommandStatus.passed)), csw.status);
}
+97
View File
@@ -0,0 +1,97 @@
//! The SCSI command descriptor blocks a transparent-SCSI (subclass 0x06) mass
//! storage device understands, and the parsers for what they return. Pure data —
//! host-testable. These CDBs go inside a Bulk-Only-Transport CBW (see
//! bulk-only-transport.zig).
//!
//! Every multi-byte SCSI field is **big-endian** — the opposite of the USB wire
//! ABI — so the LBA and transfer-length encodings are the load-bearing detail.
const std = @import("std");
// SCSI operation codes.
const op_test_unit_ready: u8 = 0x00;
const op_request_sense: u8 = 0x03;
const op_inquiry: u8 = 0x12;
const op_read_capacity_10: u8 = 0x25;
const op_read_10: u8 = 0x28;
const op_write_10: u8 = 0x2A;
const op_synchronize_cache_10: u8 = 0x35;
/// INQUIRY: standard device data (36 bytes: peripheral type, removable, vendor
/// and product strings).
pub fn inquiry(allocation_length: u8) [6]u8 {
return .{ op_inquiry, 0, 0, 0, allocation_length, 0 };
}
/// TEST UNIT READY: no data; success (CSW passed) means the unit is ready.
pub fn testUnitReady() [6]u8 {
return .{ op_test_unit_ready, 0, 0, 0, 0, 0 };
}
/// REQUEST SENSE: 18 bytes of sense data (sense key + ASC/ASCQ) explaining the
/// previous failure.
pub fn requestSense(allocation_length: u8) [6]u8 {
return .{ op_request_sense, 0, 0, 0, allocation_length, 0 };
}
/// READ CAPACITY(10): 8 bytes back — the last LBA and the block size, both u32
/// big-endian. Block count is last_lba + 1.
pub fn readCapacity10() [10]u8 {
return .{ op_read_capacity_10, 0, 0, 0, 0, 0, 0, 0, 0, 0 };
}
/// READ(10): read `blocks` logical blocks starting at `lba` into the data stage.
pub fn read10(lba: u32, blocks: u16) [10]u8 {
var cdb = [_]u8{0} ** 10;
cdb[0] = op_read_10;
std.mem.writeInt(u32, cdb[2..6], lba, .big);
std.mem.writeInt(u16, cdb[7..9], blocks, .big);
return cdb;
}
/// WRITE(10): write `blocks` logical blocks starting at `lba` from the data stage.
pub fn write10(lba: u32, blocks: u16) [10]u8 {
var cdb = [_]u8{0} ** 10;
cdb[0] = op_write_10;
std.mem.writeInt(u32, cdb[2..6], lba, .big);
std.mem.writeInt(u16, cdb[7..9], blocks, .big);
return cdb;
}
/// SYNCHRONIZE CACHE(10): commit the device's write cache to stable media. LBA 0
/// and block count 0 mean "the whole medium". No data stage. Without this a write
/// can sit in the USB flash controller's cache and be lost if power is cut right
/// after — which is exactly what a shutdown-time log flush hits on real hardware.
pub fn synchronizeCache10() [10]u8 {
var cdb = [_]u8{0} ** 10;
cdb[0] = op_synchronize_cache_10;
return cdb;
}
/// Decode an 8-byte READ CAPACITY(10) reply.
pub fn parseCapacity(bytes: [8]u8) struct { last_lba: u32, block_size: u32 } {
return .{
.last_lba = std.mem.readInt(u32, bytes[0..4], .big),
.block_size = std.mem.readInt(u32, bytes[4..8], .big),
};
}
test "read/write CDBs encode the LBA and length big-endian" {
const read = read10(0x01020304, 8);
try std.testing.expectEqualSlices(u8, &.{ 0x28, 0x00, 0x01, 0x02, 0x03, 0x04, 0x00, 0x00, 0x08, 0x00 }, &read);
const write = write10(0xAABBCCDD, 1);
try std.testing.expectEqualSlices(u8, &.{ 0x2A, 0x00, 0xAA, 0xBB, 0xCC, 0xDD, 0x00, 0x00, 0x01, 0x00 }, &write);
try std.testing.expectEqual(@as(u8, 0x25), readCapacity10()[0]);
try std.testing.expectEqual(@as(u8, 0x12), inquiry(36)[0]);
try std.testing.expectEqual(@as(u8, 36), inquiry(36)[4]);
try std.testing.expectEqual(@as(u8, 0x00), testUnitReady()[0]);
}
test "read capacity parses last LBA and block size" {
// last_lba = 0x0003FFFF (262144 blocks), block_size = 512.
const capacity = parseCapacity(.{ 0x00, 0x03, 0xFF, 0xFF, 0x00, 0x00, 0x02, 0x00 });
try std.testing.expectEqual(@as(u32, 0x0003FFFF), capacity.last_lba);
try std.testing.expectEqual(@as(u32, 512), capacity.block_size);
}
+187
View File
@@ -0,0 +1,187 @@
//! USB mass-storage class driver (Bulk-Only Transport + transparent SCSI).
//!
//! Spawned by the device manager when the xHCI bus driver reports a mass-storage
//! / SCSI / bulk-only interface (class 8, subclass 6, protocol 0x50); its device
//! id arrives as argv[1]. It owns no hardware: it opens its device through the
//! USB transfer protocol (`runtime.usb`), then drives it with the BOT command
//! cycle — CBW out, an optional data stage, CSW in — carrying SCSI commands
//! (READ CAPACITY, READ(10), WRITE(10)). Upward it is a block device: it serves
//! the block protocol under `.block`, the storage a FAT filesystem sits on.
//!
//! Block data never crosses IPC: read/write name a caller-owned DMA buffer by
//! physical address, which the data stage DMAs straight to/from.
const std = @import("std");
const runtime = @import("runtime");
const scsi = @import("scsi.zig");
const bot = @import("bulk-only-transport.zig");
const block_protocol = @import("block-protocol");
const dma = runtime.dma;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
var device_id: u64 = 0;
var device: runtime.usb.Device = undefined;
var bulk_in: runtime.usb.Endpoint = undefined;
var bulk_out: runtime.usb.Endpoint = undefined;
// DMA buffers for the transport: the 31-byte CBW, the 13-byte CSW, and a page
// for the small command data (INQUIRY / READ CAPACITY / the self-check sector).
var command_wrapper: dma.Region = undefined;
var status_wrapper: dma.Region = undefined;
var command_data: dma.Region = undefined;
var next_tag: u32 = 1;
var block_size: u32 = 512;
var block_count: u64 = 0;
/// One Bulk-Only-Transport command: send the CBW, run the data stage (to/from
/// `data_physical`), read and validate the CSW. Returns true on a passed status.
fn transact(cdb: []const u8, direction_in: bool, data_physical: u64, data_length: u32) bool {
const tag = next_tag;
next_tag +%= 1;
const wrapper: *bot.CommandBlockWrapper = @ptrFromInt(command_wrapper.virtual);
wrapper.* = .{
.tag = tag,
.data_transfer_length = data_length,
.flags = if (direction_in) bot.flag_data_in else 0,
.lun = 0,
.cdb_length = @intCast(cdb.len),
};
@memcpy(wrapper.cdb[0..cdb.len], cdb);
if (device.bulk(bulk_out.address, command_wrapper.physical, @sizeOf(bot.CommandBlockWrapper)) == null) return false;
if (data_length > 0) {
const endpoint = if (direction_in) bulk_in.address else bulk_out.address;
if (device.bulk(endpoint, data_physical, data_length) == null) return false;
}
if (device.bulk(bulk_in.address, status_wrapper.physical, @sizeOf(bot.CommandStatusWrapper)) == null) return false;
const status: *const bot.CommandStatusWrapper = @ptrFromInt(status_wrapper.virtual);
if (status.signature != bot.csw_signature or status.tag != tag) return false;
return status.status == @intFromEnum(bot.CommandStatus.passed);
}
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint;
if (!runtime.usb.helloManager(device_id)) {
_ = runtime.system.write("/system/drivers/usb-storage: hello to device manager failed\n");
return false;
}
device = runtime.usb.open(device_id) orelse {
writeLine("/system/drivers/usb-storage: could not open device {d}\n", .{device_id});
return false;
};
bulk_in = device.findEndpoint(runtime.usb.transfer_type_bulk, true) orelse {
_ = runtime.system.write("/system/drivers/usb-storage: no bulk-IN endpoint\n");
return false;
};
bulk_out = device.findEndpoint(runtime.usb.transfer_type_bulk, false) orelse {
_ = runtime.system.write("/system/drivers/usb-storage: no bulk-OUT endpoint\n");
return false;
};
command_wrapper = dma.alloc(4096, dma.coherent) orelse return false;
status_wrapper = dma.alloc(4096, dma.coherent) orelse return false;
command_data = dma.alloc(4096, dma.coherent) orelse return false;
// Bring the LUN up: wait for it to be ready (clearing the initial unit-attention
// with REQUEST SENSE), identify it, and read its capacity.
var tries: u32 = 0;
while (tries < 10) : (tries += 1) {
const ready = scsi.testUnitReady();
if (transact(&ready, false, 0, 0)) break;
const sense = scsi.requestSense(18);
_ = transact(&sense, true, command_data.physical, 18);
runtime.system.sleep(50);
}
const inquiry = scsi.inquiry(36);
_ = transact(&inquiry, true, command_data.physical, 36);
const capacity_command = scsi.readCapacity10();
if (!transact(&capacity_command, true, command_data.physical, 8)) {
_ = runtime.system.write("/system/drivers/usb-storage: READ CAPACITY failed\n");
return false;
}
var capacity_bytes: [8]u8 = undefined;
const capacity_source: [*]const u8 = @ptrFromInt(command_data.virtual);
@memcpy(&capacity_bytes, capacity_source[0..8]);
const capacity = scsi.parseCapacity(capacity_bytes);
block_size = capacity.block_size;
block_count = @as(u64, capacity.last_lba) + 1;
writeLine("/system/drivers/usb-storage: ready ({d} blocks x {d} bytes)\n", .{ block_count, block_size });
// Self-check: read block 0 and log its trailing signature (0x55AA for a boot
// sector) — proof READ(10) works end to end over the bulk path.
const read0 = scsi.read10(0, 1);
if (block_size <= 4096 and transact(&read0, true, command_data.physical, block_size)) {
const sector: [*]const u8 = @ptrFromInt(command_data.virtual);
writeLine("/system/drivers/usb-storage: block 0 signature 0x{x:0>2}{x:0>2}\n", .{ sector[510], sector[511] });
}
return true;
}
/// Serve the block protocol: geometry, and whole-block read/write to/from the
/// caller's DMA buffer (named by physical address).
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = sender;
_ = capability;
if (message.len < block_protocol.request_size) return 0;
const request = std.mem.bytesToValue(block_protocol.Request, message[0..block_protocol.request_size]);
switch (request.operation) {
@intFromEnum(block_protocol.Operation.geometry) => {
return writeReply(reply, .{ .status = 0, .block_size = block_size, .block_count = block_count });
},
@intFromEnum(block_protocol.Operation.read) => {
const count: u16 = @intCast(request.count);
const cdb = scsi.read10(@intCast(request.lba), count);
const ok = transact(&cdb, true, request.physical, request.count * block_size);
return writeReply(reply, .{ .status = if (ok) 0 else -1, .block_size = block_size, .block_count = if (ok) request.count else 0 });
},
@intFromEnum(block_protocol.Operation.write) => {
const count: u16 = @intCast(request.count);
const cdb = scsi.write10(@intCast(request.lba), count);
const ok = transact(&cdb, false, request.physical, request.count * block_size);
return writeReply(reply, .{ .status = if (ok) 0 else -1, .block_size = block_size, .block_count = if (ok) request.count else 0 });
},
@intFromEnum(block_protocol.Operation.flush) => {
// SYNCHRONIZE CACHE: commit the device's write cache to flash. No data
// stage. Makes prior writes durable before a caller (init at shutdown)
// cuts power. A device without a volatile cache reports success anyway.
const cdb = scsi.synchronizeCache10();
const ok = transact(&cdb, false, 0, 0);
return writeReply(reply, .{ .status = if (ok) 0 else -1, .block_size = block_size, .block_count = 0 });
},
else => return 0,
}
}
fn writeReply(reply: []u8, value: block_protocol.Reply) usize {
const bytes = std.mem.asBytes(&value);
@memcpy(reply[0..bytes.len], bytes);
return bytes.len;
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("/system/drivers/usb-storage: missing device id (argv[1])\n");
return;
};
device_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/usb-storage: malformed device id '{s}'\n", .{argument});
return;
};
runtime.service.run(block_protocol.message_maximum, .{
.service = .block,
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
@@ -0,0 +1,159 @@
//! The USB transfer protocol: what a USB class driver (a keyboard, mouse, or
//! mass-storage driver) says to the xHCI bus driver over its well-known
//! `.usb_bus` endpoint to drive its device. The class driver owns no hardware —
//! it reaches its device entirely through these messages, the way a PS/2 keyboard
//! driver reaches the 8042 through the ps2-bus. Extern-struct messages tagged by
//! `Operation`, the vfs-protocol / device-manager-protocol pattern.
//!
//! The shape:
//! - **open** (a capability-passing `ipc.callCap`): the class driver hands over
//! its own endpoint (for asynchronous interrupt reports) and its assigned
//! device id, and receives a `device_token` plus its interface's endpoints.
//! - **control / bulk** (synchronous `ipc.call`): one transfer, answered when
//! it completes. Control data travels inline (descriptors, HID/MSC class
//! requests are all small); bulk data travels by **physical address** — the
//! class driver's own `dma_alloc`'d buffer — so a 512-byte sector never has
//! to cross the 256-byte IPC boundary.
//! - **interrupt_subscribe** (synchronous): arm periodic IN polling of an
//! interrupt endpoint; each report the device produces is then pushed to the
//! class driver's endpoint as an asynchronous `InterruptReport` (`ipc.send`),
//! exactly how the input service delivers events.
//!
//! Single controller assumption: one `.usb_bus` singleton serves QEMU's one xHCI.
//! A multi-controller machine would need a per-controller endpoint (the device
//! manager handing each class driver the right one); noted, not built.
/// Fits one synchronous IPC message (kernel MESSAGE_MAXIMUM).
pub const message_maximum: usize = 256;
/// The largest inline control-transfer payload. Sized so a whole message
/// (header + data) stays under `message_maximum`: descriptors and HID/MSC class
/// requests are all far smaller.
pub const max_inline_data: usize = 200;
/// The largest interrupt report pushed asynchronously. Sized so `InterruptReport`
/// fits an `ipc_send` payload slot (POST_MAXIMUM = 64): boot keyboard reports are
/// 8 bytes, boot mouse reports 3–4.
pub const max_report_data: usize = 48;
/// Endpoints per interface reported back in an open reply (a boot HID interface
/// has one interrupt endpoint, a mass-storage interface two bulk endpoints).
pub const max_reported_endpoints: usize = 4;
pub const Operation = enum(u32) {
open = 0,
control = 1,
interrupt_subscribe = 2,
bulk = 3,
};
/// The endpoint facts a class driver needs, lifted from the endpoint descriptor
/// the bus driver already parsed during enumeration.
pub const Endpoint = extern struct {
/// EndpointDescriptor address: direction in bit 7, number in bits 3:0.
address: u8,
/// 0 control, 1 isochronous, 2 bulk, 3 interrupt.
transfer_type: u8,
max_packet_size: u16,
interval: u8,
reserved: [3]u8 = .{ 0, 0, 0 },
};
/// open: the class driver's receive endpoint rides as the call's capability, and
/// `device_id` is the interface's assigned id (its argv[1]).
pub const OpenRequest = extern struct {
operation: u32 = @intFromEnum(Operation.open),
reserved: u32 = 0,
device_id: u64,
};
/// The answer to open: a token scoping every later request to this device, the
/// interface's class triple (a sanity check), and its endpoints.
pub const OpenReply = extern struct {
status: i32,
endpoint_count: u32,
device_token: u64,
interface_class: u8,
interface_subclass: u8,
interface_protocol: u8,
interface_number: u8,
reserved2: u32 = 0,
endpoints: [max_reported_endpoints]Endpoint = [_]Endpoint{.{ .address = 0, .transfer_type = 0, .max_packet_size = 0, .interval = 0 }} ** max_reported_endpoints,
};
/// control: one EP0 control transfer. `setup` is a bit-cast `usb_abi.Request`.
/// For an OUT transfer `data[0..data_length]` is sent; for an IN transfer the
/// reply carries up to `data_length` bytes back.
pub const ControlRequest = extern struct {
operation: u32 = @intFromEnum(Operation.control),
reserved: u32 = 0,
device_token: u64,
setup: [8]u8,
direction_in: u8, // 1 = device-to-host (IN), 0 = host-to-device (OUT)
reserved2: u8 = 0,
data_length: u16,
reserved3: u32 = 0,
data: [max_inline_data]u8 = [_]u8{0} ** max_inline_data,
};
pub const ControlReply = extern struct {
status: i32, // 0 success, negative on failure/stall
actual_length: u32,
data: [max_inline_data]u8 = [_]u8{0} ** max_inline_data,
};
/// interrupt_subscribe: begin periodic IN polling of an interrupt endpoint. Each
/// report the device returns is pushed to the caller's endpoint (handed over at
/// open) as an asynchronous `InterruptReport`.
pub const InterruptSubscribeRequest = extern struct {
operation: u32 = @intFromEnum(Operation.interrupt_subscribe),
reserved: u32 = 0,
device_token: u64,
endpoint_address: u8,
reserved2: u8 = 0,
max_length: u16, // bytes to request per poll (the endpoint's max packet size)
};
pub const InterruptSubscribeReply = extern struct {
status: i32,
reserved: u32 = 0,
};
/// bulk: one bulk IN or OUT transfer. `physical_address` is the class driver's own
/// `dma_alloc`'d buffer — the controller DMAs straight to/from it, so the bulk
/// data never crosses IPC. `endpoint_address`'s bit 7 selects IN vs OUT.
pub const BulkRequest = extern struct {
operation: u32 = @intFromEnum(Operation.bulk),
reserved: u32 = 0,
device_token: u64,
physical_address: u64,
length: u32,
endpoint_address: u8,
reserved2: u8 = 0,
reserved3: u16 = 0,
};
pub const BulkReply = extern struct {
status: i32,
actual_length: u32,
};
/// An asynchronous interrupt report, pushed with `ipc.send` to a subscriber's
/// endpoint. `Received.isMessage()` is set; there is no reply owed.
pub const InterruptReport = extern struct {
device_token: u64,
endpoint_address: u8,
length: u8,
reserved: u16 = 0,
data: [max_report_data]u8 = [_]u8{0} ** max_report_data,
};
comptime {
const std = @import("std");
// Every synchronous message must fit one IPC message; the async report must
// fit an ipc_send payload slot.
std.debug.assert(@sizeOf(ControlRequest) <= message_maximum);
std.debug.assert(@sizeOf(ControlReply) <= message_maximum);
std.debug.assert(@sizeOf(OpenReply) <= message_maximum);
std.debug.assert(@sizeOf(InterruptReport) <= 64);
}
+268 -35
View File
@@ -17,6 +17,52 @@ const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
const device = runtime.device;
const usb_ids = @import("usb-ids");
const usb_abi = @import("usb-abi");
const transfer = @import("usb-transfer-protocol");
const library = @import("usb-xhci-library.zig");
/// The controller engine (reset, rings, transfers), stood up in `initialise`.
var controller: ?library.Controller = null;
/// This driver's service endpoint (registered as `.usb_bus`), where class-driver
/// requests, signals, and the interrupt-poll timer all arrive.
var service_endpoint: runtime.ipc.Handle = 0;
/// How often the driver drains the event ring for interrupt reports (~125 Hz),
/// re-armed each tick. Frequent enough for responsive input.
const poll_interval_ms: u64 = 8;
/// The class driver endpoints that opened each device, so interrupt reports can
/// be pushed back to them. Keyed by the device token (the interface's device id).
const Open = struct {
used: bool = false,
device_token: u64 = 0,
report_endpoint: usize = 0,
};
var opens = [_]Open{.{}} ** 16;
fn recordOpen(device_token: u64, report_endpoint: usize) void {
for (&opens) |*open| {
if (open.used and open.device_token == device_token) {
open.report_endpoint = report_endpoint;
return;
}
}
for (&opens) |*open| {
if (!open.used) {
open.* = .{ .used = true, .device_token = device_token, .report_endpoint = report_endpoint };
return;
}
}
}
fn reportEndpointFor(device_token: u64) ?usize {
for (&opens) |*open| {
if (open.used and open.device_token == device_token) return open.report_endpoint;
}
return null;
}
/// Format one whole log line and emit it in a single `debug_write`, so
/// concurrent instances (one per controller) can never interleave mid-line.
@@ -31,7 +77,7 @@ var controller_id: u64 = protocol.no_device;
/// manager. Any failure returns false: the process exits cleanly, which the
/// manager reads as "meant to stop" — a missing assignment is not a crash loop.
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint;
service_endpoint = endpoint;
if (!device.claim(controller_id)) {
writeLine("/system/drivers/usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id});
return false;
@@ -72,6 +118,26 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
return false;
};
// Bring the controller up: reset it, stand up the command and event rings,
// and start it running (the hardware half lives in usb-xhci-library.zig).
controller = library.Controller.init(register_base) orelse {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: controller reset/bring-up failed\n");
return false;
};
writeLine("/system/drivers/usb-xhci-bus: controller running ({d} slots, {d}-byte contexts)\n", .{
controller.?.max_slots,
controller.?.context_size,
});
// The proof of life: a No-Op command round-trips the command ring, the event
// ring, the doorbell, and the cycle-bit bookkeeping. If this completes, the
// engine is sound; transfers build on exactly this machinery.
if (controller.?.noOpCommand()) {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: command ring running (no-op ok)\n");
} else {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: no-op command did not complete\n");
return false;
}
// The handshake: role, protocol version, assignment — inside the manager's
// deadline (the lookup retries cover the manager still registering).
var manager: ?runtime.ipc.Handle = null;
@@ -97,17 +163,15 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: hello acknowledged\n");
scanPorts(h);
// Arm the poll timer that drains interrupt reports from the event ring. It is
// re-armed on each tick in onNotification; class drivers subscribe later.
_ = runtime.system.timerOnce(service_endpoint, poll_interval_ms);
return true;
}
var register_base: usize = 0;
/// One 32-bit volatile register read at `offset` from the mapped window.
fn readRegister(offset: usize) u32 {
const register: *volatile u32 = @ptrFromInt(register_base + offset);
return register.*;
}
/// The xHCI default Protocol Speed IDs (the PORTSC port-speed field, bits 13:10)
/// decoded to human names — the boot-log breadcrumb for what actually enumerated on
/// a port, the USB analog of the pci-bus class-code line. A controller may redefine
@@ -124,50 +188,217 @@ fn speedName(speed: u32) []const u8 {
};
}
/// The root-hub port scan: read the capability registers for the port count
/// and the operational-register offset, then one PORTSC per port. The connect
/// bit (CCS) and the speed field reflect hardware state directly — no
/// controller reset or run needed to *see* the devices; driving them needs the
/// rings (the USB track).
/// The root-hub scan and enumeration: for each connected port, bring the device
/// up (reset → enable slot → address), read its descriptors, and register +
/// report one child per interface — carrying the interface's (class, subclass,
/// protocol) triple as identity, which is what the device manager matches a
/// class driver against.
fn scanPorts(manager: runtime.ipc.Handle) void {
// Capability registers: CAPLENGTH is byte 0 of the first dword; HCSPARAMS1
// carries MaxPorts in bits 31:24.
const capability_length = readRegister(0) & 0xFF;
const structural = readRegister(0x04);
const maximum_ports: u32 = structural >> 24;
writeLine("/system/drivers/usb-xhci-bus: {d} root-hub ports\n", .{maximum_ports});
const engine = if (controller) |*c| c else {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: controller not initialised\n");
return;
};
writeLine("/system/drivers/usb-xhci-bus: {d} root-hub ports\n", .{engine.max_ports});
// PORTSC registers: operational base + 0x400 + 0x10 per port (1-based).
var port: u32 = 1;
var connected: u32 = 0;
while (port <= maximum_ports) : (port += 1) {
const port_status = readRegister(capability_length + 0x400 + 0x10 * (port - 1));
while (port <= engine.max_ports) : (port += 1) {
const port_status = engine.portStatus(port);
if (port_status & 1 == 0) continue; // CCS: nothing connected
connected += 1;
const speed = (port_status >> 10) & 0xF; // the PORTSC port-speed class
writeLine("/system/drivers/usb-xhci-bus: port {d} connected — {s} (speed class {d})\n", .{ port, speedName(speed), speed });
const report = protocol.ChildAdded{
.parent = controller_id,
.bus_address = port,
.identity = speed,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&report), &reply) catch {
writeLine("/system/drivers/usb-xhci-bus: child report for port {d} failed\n", .{port});
const usb_device = engine.setupDevice(port, speed) orelse {
writeLine("/system/drivers/usb-xhci-bus: port {d} device setup failed\n", .{port});
continue;
};
if (!engine.enumerate(usb_device)) {
writeLine("/system/drivers/usb-xhci-bus: port {d} enumeration failed\n", .{port});
continue;
}
writeLine("/system/drivers/usb-xhci-bus: port {d} device vendor 0x{x:0>4} product 0x{x:0>4}, {d} interface(s)\n", .{
port,
usb_device.device_descriptor.vendor_id,
usb_device.device_descriptor.product_id,
usb_device.interface_count,
});
for (usb_device.interfaces[0..usb_device.interface_count]) |*interface| {
// Record the id each interface was registered as, so a class driver
// opening the interface (by that id) resolves to it.
if (reportInterface(manager, port, interface.*)) |registered| {
interface.registered_device_id = registered;
}
}
}
if (connected == 0) _ = runtime.system.write("/system/drivers/usb-xhci-bus: no devices connected\n");
}
/// No bus protocol to serve yet — transfer requests arrive with the USB track.
/// Register one interface as a resource-less child of the controller and report
/// it to the device manager. The identity is the packed USB class triple, so the
/// manager can match a class driver (HID keyboard, mouse, mass storage); the
/// registered device id becomes that driver's argv[1] assignment. Returns the
/// registered device id, or null if registration or the report failed.
fn reportInterface(manager: runtime.ipc.Handle, port: u32, interface: library.InterfaceInfo) ?u64 {
const identity = usb_ids.packTriple(interface.class, interface.subclass, interface.protocol);
// A USB device is reached through its controller, not by MMIO, so the child
// carries no resources; register() allows that. Its bus-local identity — the
// (port, interface) address, written as a short "P<port>I<interface>" tag in
// the hid field — makes each interface a distinct kernel node (the register
// dedup keys on class/pci_class/hid/resources, all otherwise identical here)
// and keeps re-registration idempotent across a bus restart: the same port
// and interface always map back to the same device id.
var descriptor = std.mem.zeroes(device.DeviceDescriptor);
descriptor.class = @intFromEnum(device.DeviceClass.usb_device);
descriptor.pci_class = device.no_pci_class;
descriptor.resource_count = 0;
var hid_buffer: [8]u8 = undefined;
const hid_text = std.fmt.bufPrint(&hid_buffer, "P{d}I{d}", .{ port, interface.number }) catch "";
descriptor.hid_len = hid_text.len;
@memcpy(descriptor.hid[0..hid_text.len], hid_text);
const registered = device.register(controller_id, &descriptor) orelse {
writeLine("/system/drivers/usb-xhci-bus: register refused for port {d} interface {d}\n", .{ port, interface.number });
return null;
};
const report = protocol.ChildAdded{
.parent = controller_id,
.bus_address = (@as(u64, port) << 8) | interface.number,
.identity = identity,
.device_id = registered,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&report), &reply) catch {
writeLine("/system/drivers/usb-xhci-bus: child report for port {d} interface {d} failed\n", .{ port, interface.number });
return null;
};
writeLine("/system/drivers/usb-xhci-bus: port {d} interface {d} class {d}/{d}/{d} registered as device {d}\n", .{
port,
interface.number,
interface.class,
interface.subclass,
interface.protocol,
registered,
});
return registered;
}
/// Serve the USB transfer protocol: a class driver opens its device, then issues
/// control / interrupt-subscribe / bulk requests against it.
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
if (message.len < 4) return 0;
const operation = std.mem.readInt(u32, message[0..4], .little);
return switch (operation) {
@intFromEnum(transfer.Operation.open) => handleOpen(message, reply, capability),
@intFromEnum(transfer.Operation.control) => handleControl(message, reply),
@intFromEnum(transfer.Operation.interrupt_subscribe) => handleSubscribe(message, reply),
@intFromEnum(transfer.Operation.bulk) => handleBulk(message, reply),
else => 0,
};
}
fn writeReply(reply: []u8, value: anytype) usize {
const bytes = std.mem.asBytes(&value);
@memcpy(reply[0..bytes.len], bytes);
return bytes.len;
}
/// open: resolve the assigned device id to an interface, remember the caller's
/// endpoint (for interrupt reports), and answer with a device token + the
/// interface's endpoints so the class driver need not re-read the config.
fn handleOpen(message: []const u8, reply: []u8, capability: ?runtime.ipc.Handle) usize {
if (message.len < @sizeOf(transfer.OpenRequest)) return writeReply(reply, transfer.OpenReply{ .status = -1, .endpoint_count = 0, .device_token = 0, .interface_class = 0, .interface_subclass = 0, .interface_protocol = 0, .interface_number = 0 });
const request = std.mem.bytesToValue(transfer.OpenRequest, message[0..@sizeOf(transfer.OpenRequest)]);
const engine = if (controller) |*c| c else return writeReply(reply, transfer.OpenReply{ .status = -1, .endpoint_count = 0, .device_token = 0, .interface_class = 0, .interface_subclass = 0, .interface_protocol = 0, .interface_number = 0 });
const found = engine.findInterface(request.device_id) orelse return writeReply(reply, transfer.OpenReply{ .status = -1, .endpoint_count = 0, .device_token = 0, .interface_class = 0, .interface_subclass = 0, .interface_protocol = 0, .interface_number = 0 });
if (capability) |endpoint| recordOpen(request.device_id, endpoint);
var open_reply = transfer.OpenReply{
.status = 0,
.endpoint_count = found.interface.endpoint_count,
.device_token = request.device_id,
.interface_class = found.interface.class,
.interface_subclass = found.interface.subclass,
.interface_protocol = found.interface.protocol,
.interface_number = found.interface.number,
};
const count = @min(found.interface.endpoint_count, transfer.max_reported_endpoints);
for (found.interface.endpoints[0..count], 0..) |endpoint, index| {
open_reply.endpoints[index] = .{
.address = endpoint.address,
.transfer_type = endpoint.transfer_type,
.max_packet_size = endpoint.max_packet_size,
.interval = endpoint.interval,
};
}
return writeReply(reply, open_reply);
}
/// control: one EP0 control transfer, small data inline both ways.
fn handleControl(message: []const u8, reply: []u8) usize {
if (message.len < @sizeOf(transfer.ControlRequest)) return writeReply(reply, transfer.ControlReply{ .status = -1, .actual_length = 0 });
const request = std.mem.bytesToValue(transfer.ControlRequest, message[0..@sizeOf(transfer.ControlRequest)]);
const engine = if (controller) |*c| c else return writeReply(reply, transfer.ControlReply{ .status = -1, .actual_length = 0 });
const found = engine.findInterface(request.device_token) orelse return writeReply(reply, transfer.ControlReply{ .status = -1, .actual_length = 0 });
const setup = std.mem.bytesToValue(usb_abi.Request, &request.setup);
const direction_in = request.direction_in != 0;
const data_length = @min(request.data_length, transfer.max_inline_data);
var data: [transfer.max_inline_data]u8 = undefined;
if (!direction_in) @memcpy(data[0..data_length], request.data[0..data_length]);
const ok = engine.controlTransfer(found.device, setup, data[0..data_length], direction_in);
var control_reply = transfer.ControlReply{ .status = if (ok) 0 else -1, .actual_length = if (ok) data_length else 0 };
if (ok and direction_in) @memcpy(control_reply.data[0..data_length], data[0..data_length]);
return writeReply(reply, control_reply);
}
/// interrupt_subscribe: arm periodic IN polling; reports flow back asynchronously.
fn handleSubscribe(message: []const u8, reply: []u8) usize {
if (message.len < @sizeOf(transfer.InterruptSubscribeRequest)) return writeReply(reply, transfer.InterruptSubscribeReply{ .status = -1 });
const request = std.mem.bytesToValue(transfer.InterruptSubscribeRequest, message[0..@sizeOf(transfer.InterruptSubscribeRequest)]);
const engine = if (controller) |*c| c else return writeReply(reply, transfer.InterruptSubscribeReply{ .status = -1 });
const found = engine.findInterface(request.device_token) orelse return writeReply(reply, transfer.InterruptSubscribeReply{ .status = -1 });
const endpoint = library.Controller.endpointForAddress(found.interface, request.endpoint_address) orelse return writeReply(reply, transfer.InterruptSubscribeReply{ .status = -1 });
const report_endpoint = reportEndpointFor(request.device_token) orelse return writeReply(reply, transfer.InterruptSubscribeReply{ .status = -1 });
const ok = engine.subscribeInterrupt(found.device, endpoint, request.device_token, report_endpoint);
return writeReply(reply, transfer.InterruptSubscribeReply{ .status = if (ok) 0 else -1 });
}
/// bulk: one bulk transfer to/from the class driver's own DMA buffer (by physical
/// address), so sector-sized data never crosses IPC.
fn handleBulk(message: []const u8, reply: []u8) usize {
if (message.len < @sizeOf(transfer.BulkRequest)) return writeReply(reply, transfer.BulkReply{ .status = -1, .actual_length = 0 });
const request = std.mem.bytesToValue(transfer.BulkRequest, message[0..@sizeOf(transfer.BulkRequest)]);
const engine = if (controller) |*c| c else return writeReply(reply, transfer.BulkReply{ .status = -1, .actual_length = 0 });
const found = engine.findInterface(request.device_token) orelse return writeReply(reply, transfer.BulkReply{ .status = -1, .actual_length = 0 });
const endpoint = library.Controller.endpointForAddress(found.interface, request.endpoint_address) orelse return writeReply(reply, transfer.BulkReply{ .status = -1, .actual_length = 0 });
const transferred = engine.bulkTransfer(found.device, endpoint, request.physical_address, request.length);
return writeReply(reply, transfer.BulkReply{ .status = if (transferred != null) 0 else -1, .actual_length = transferred orelse 0 });
}
/// The poll timer landed: drain any interrupt reports off the event ring and push
/// each to the class driver that subscribed, then re-arm the timer.
fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_timer_bit == 0) return;
if (controller) |*engine| {
engine.pump();
while (engine.takeReport()) |report| {
var message = transfer.InterruptReport{
.device_token = report.device_token,
.endpoint_address = report.endpoint_address,
.length = @intCast(@min(report.length, transfer.max_report_data)),
};
const n = @min(report.length, transfer.max_report_data);
@memcpy(message.data[0..n], report.data[0..n]);
_ = runtime.ipc.send(report.report_endpoint, std.mem.asBytes(&message));
}
}
_ = runtime.system.timerOnce(service_endpoint, poll_interval_ms);
}
pub fn main(init: runtime.process.Init) void {
@@ -179,9 +410,11 @@ pub fn main(init: runtime.process.Init) void {
writeLine("/system/drivers/usb-xhci-bus: malformed controller device id '{s}'\n", .{argument});
return;
};
runtime.service.run(protocol.message_maximum, .{
runtime.service.run(transfer.message_maximum, .{
.service = .usb_bus,
.init = initialise,
.on_message = onMessage,
.on_notification = onNotification,
});
}
File diff suppressed because it is too large Load Diff
+107 -4
View File
@@ -106,6 +106,12 @@ pub fn serialWrite(bytes: []const u8) void {
serial.write(bytes);
}
/// Whether a working UART was detected (loopback probe). When false the serial
/// sink is silently inert — a dead legacy COM1 costs nothing per byte.
pub fn serialPresent() bool {
return serial.present();
}
/// Emit a one-byte progress checkpoint to whatever hardware debug sink the
/// platform has — here the POST diagnostic port (0x80), which a POST card or BMC
/// displays. The last-resort progress signal when there's no text output at all.
@@ -162,10 +168,18 @@ pub fn mapUserPageInto(root: u64, virtual: u64, physical: u64, writable: bool, e
paging.mapUserInto(root, virtual, physical, writable, executable);
}
/// Map a device MMIO window into address space `root`: strong-uncacheable, RW+NX,
/// and marked so teardown won't free the MMIO frames as RAM. For IO passthrough.
pub fn mapUserDeviceInto(root: u64, virtual: u64, physical: u64, len: u64) void {
paging.mapUserDeviceInto(root, virtual, physical, len);
/// Map a device MMIO window into address space `root`: RW+NX, and marked so teardown
/// won't free the MMIO frames as RAM. `write_combining` picks the cache type —
/// false = strong-uncacheable (registers), true = write-combining (a framebuffer).
/// For IO passthrough.
pub fn mapUserDeviceInto(root: u64, virtual: u64, physical: u64, len: u64, write_combining: bool) void {
paging.mapUserDeviceInto(root, virtual, physical, len, write_combining);
}
/// Is the user leaf mapping `virtual` in address space `root` write-combining? Null if
/// unmapped. For tests verifying the framebuffer map's cache type.
pub fn userLeafIsWriteCombining(root: u64, virtual: u64) ?bool {
return paging.leafIsWriteCombining(root, virtual);
}
/// Map coherent DMA RAM into address space `root`: strong-uncacheable, RW+NX, but
@@ -469,6 +483,95 @@ pub fn clockHz() u64 {
return apic.tscHz();
}
// --- real-time clock (CMOS) --------------------------------------------------
//
// The battery-backed CMOS clock, read once at boot and thereafter anchored to the
// monotonic clock (see kernel/wall-clock.zig) — so this is never on a hot path and
// needs no lock. Wall-clock *seconds* are mechanism the kernel owns (the hardware's
// value), like the monotonic clock; calendars/timezones are policy layered on top.
fn cmosRead(register: u8) u8 {
io.outb(0x70, register);
return io.inb(0x71);
}
const RtcFields = struct { second: u8, minute: u8, hour: u8, day: u8, month: u8, year: u8 };
fn rtcRaw() RtcFields {
while (cmosRead(0x0A) & 0x80 != 0) {} // wait out any update in progress (status A bit 7)
return .{
.second = cmosRead(0x00),
.minute = cmosRead(0x02),
.hour = cmosRead(0x04),
.day = cmosRead(0x07),
.month = cmosRead(0x08),
.year = cmosRead(0x09),
};
}
fn bcdToBinary(v: u8) u8 {
return (v & 0x0F) + ((v >> 4) * 10);
}
fn isLeapYear(y: u32) bool {
return (y % 4 == 0 and y % 100 != 0) or (y % 400 == 0);
}
/// Read the CMOS real-time clock and convert it to Unix epoch seconds (UTC).
pub fn readRtcUnixSeconds() u64 {
// Read until two consecutive reads agree, so we never latch a half-updated time.
var a = rtcRaw();
while (true) {
const b = rtcRaw();
if (a.second == b.second and a.minute == b.minute and a.hour == b.hour and
a.day == b.day and a.month == b.month and a.year == b.year) break;
a = b;
}
const status_b = cmosRead(0x0B);
const binary_mode = status_b & 0x04 != 0; // else BCD
const hour_24 = status_b & 0x02 != 0; // else 12-hour with a PM bit
var second = a.second;
var minute = a.minute;
var hour_field = a.hour;
var day = a.day;
var month = a.month;
var year = a.year;
if (!binary_mode) {
second = bcdToBinary(second);
minute = bcdToBinary(minute);
hour_field = bcdToBinary(hour_field & 0x7F) | (hour_field & 0x80); // preserve the PM bit
day = bcdToBinary(day);
month = bcdToBinary(month);
year = bcdToBinary(year);
}
var hour: u32 = hour_field & 0x7F;
if (!hour_24) {
const pm = hour_field & 0x80 != 0;
hour %= 12; // 12 AM/PM -> 0
if (pm) hour += 12;
}
// The CMOS year is 0..99; QEMU and modern hardware mean 20xx (there is no
// reliable century register on QEMU). Treat < 70 as 20xx, else 19xx.
const full_year: u32 = if (year < 70) 2000 + @as(u32, year) else 1900 + @as(u32, year);
var days: u64 = 0;
var y: u32 = 1970;
while (y < full_year) : (y += 1) days += if (isLeapYear(y)) 366 else 365;
const month_lengths = [_]u8{ 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31 };
var m: u8 = 1;
while (m < month) : (m += 1) {
days += month_lengths[m - 1];
if (m == 2 and isLeapYear(full_year)) days += 1;
}
days += @as(u64, day) - 1;
return ((days * 24 + hour) * 60 + minute) * 60 + second;
}
/// Whether the CPU guarantees an **invariant** TSC (CPUID 0x80000007 EDX[8] on
/// x86; the analogous architectural guarantee elsewhere). When false the TSC is not
/// used as the clocksource.
+118 -18
View File
@@ -7,8 +7,13 @@
//! unmapped as a null guard. It also exposes map/unmap for on-demand mapping,
//! which the kernel heap will build on.
//!
//! Everything is 4 KiB pages — precise and simple; the extra table memory is
//! negligible against available RAM.
//! The physmap (the permanent window onto all physical RAM) is built with 2 MiB
//! huge pages wherever the range is 2 MiB-aligned, falling back to 4 KiB for the
//! unaligned edges. On a big machine that is the difference between ~16.7M page-
//! table entries (128 MiB of tables) and ~32K — it makes both the build and the
//! footprint scale sanely with RAM. Everything else (kernel segments, heap, user
//! space, on-demand MMIO) stays 4 KiB: precise, and the table memory is
//! negligible there.
const boot_handoff = @import("boot-handoff");
const abi = @import("abi");
@@ -22,10 +27,22 @@ const writable: u64 = 1 << 1;
const user: u64 = 1 << 2; // U/S: accessible from ring 3 (must be set at every level)
const pwt: u64 = 1 << 3; // page write-through
const pcd: u64 = 1 << 4; // page cache disable (with PWT: strong-uncacheable under the default PAT)
const page_size_bit: u64 = 1 << 7; // PS: this PDPT/PD entry is a 1 GiB/2 MiB leaf, not a pointer to the next table
const device_grant: u64 = 1 << 9; // available bit: this leaf maps device MMIO, not RAM — do not reclaim
const no_execute: u64 = 1 << 63;
const address_mask: u64 = 0x000F_FFFF_FFFF_F000;
// The PAT-index bit. In a 4 KiB PTE it is bit 7; in a huge leaf (2 MiB PDE / 1 GiB
// PDPTE) bit 7 is PS, so the PAT bit moves to bit 12. With PCD=PWT=0 this selects
// PAT entry 4, which `setupPat` programs to write-combining (see mapRangePhysmap).
const pte_pat: u64 = 1 << 7;
const huge_pat: u64 = 1 << 12;
const ia32_pat: u32 = 0x277;
/// The physmap's page size for 2 MiB-aligned RAM: one PD leaf covers this instead
/// of 512 PT entries. 4 KiB pages fill the unaligned edges (see mapRangePhysmap).
const huge_page_size: u64 = 2 << 20; // 2 MiB
// ELF segment flags (p_flags).
const pf_x: u32 = 1;
const pf_w: u32 = 2;
@@ -74,7 +91,15 @@ fn allocTable() u64 {
/// entries are writable and executable so the leaf's bits govern (a page is
/// writable only if every level is; non-executable if any level is).
fn descend(entry: *u64) u64 {
if (entry.* & present != 0) return entry.* & address_mask;
if (entry.* & present != 0) {
// A present-but-huge entry is a leaf, not a table: descending would read
// its 2 MiB/1 GiB data frame as a page table and corrupt RAM. This only
// fires on a bug — a 4 KiB map landing inside a physmap huge page — and a
// loud panic beats silent corruption. (The physmap and the 4 KiB regions
// live in disjoint PML4 slots, so it should never happen.)
if (entry.* & page_size_bit != 0) @panic("paging: descend through a huge-page leaf");
return entry.* & address_mask;
}
const frame = allocTable();
entry.* = frame | present | writable;
return frame;
@@ -96,15 +121,39 @@ fn mapPage(pml4: u64, virtual: u64, physical: u64, flags: u64) void {
tableAt(pt)[(virtual >> 12) & 0x1FF] = (physical & address_mask) | flags | present;
}
/// Map one 2 MiB huge page `virtual` -> `physical` with `flags` — a leaf at the PD
/// level (PS bit set), with no PT beneath it. Both addresses must be 2 MiB-aligned.
/// One of these replaces 512 `mapPage`s (and the PT frame they'd need).
fn mapHugePage(pml4: u64, virtual: u64, physical: u64, flags: u64) void {
const pml4e = &tableAt(pml4)[(virtual >> 39) & 0x1FF];
if (init_done and (virtual >> 63) == 1 and pml4e.* & present == 0)
@panic("paging: new higher-half PML4 entry after init");
const pdpt = descend(pml4e);
const pdpte = &tableAt(pdpt)[(virtual >> 30) & 0x1FF];
const pd = descend(pdpte);
tableAt(pd)[(virtual >> 21) & 0x1FF] = (physical & address_mask) | flags | present | page_size_bit;
}
/// Map [physical_base, physical_base+len) into the physmap (at physicalToVirtual(physical)) with
/// `flags`, rounded out to whole pages. This is how the kernel keeps a permanent
/// window onto physical memory once the low identity map goes away.
fn mapRangePhysmap(pml4: u64, physical_base: u64, len: u64, flags: u64) void {
/// window onto physical memory once the low identity map goes away. The 2 MiB-
/// aligned interior is mapped with huge pages; the unaligned head/tail with 4 KiB.
/// `write_combining` selects the WC memory type (setupPat's PAT entry 4) via the
/// PAT bit — bit 7 in a 4 KiB PTE, bit 12 in a huge leaf — for the framebuffer.
fn mapRangePhysmap(pml4: u64, physical_base: u64, len: u64, flags: u64, write_combining: bool) void {
const pte_flags = if (write_combining) flags | pte_pat else flags;
const huge_flags = if (write_combining) flags | huge_pat else flags;
var address = physical_base & ~@as(u64, page_size - 1);
const end = physical_base + len;
while (address < end) : (address += page_size) {
mapPage(pml4, boot_handoff.physicalToVirtual(address), address, flags);
}
// Head: 4 KiB pages up to the next 2 MiB boundary.
while (address < end and address & (huge_page_size - 1) != 0) : (address += page_size)
mapPage(pml4, boot_handoff.physicalToVirtual(address), address, pte_flags);
// Interior: 2 MiB huge pages while a whole one still fits.
while (address + huge_page_size <= end) : (address += huge_page_size)
mapHugePage(pml4, boot_handoff.physicalToVirtual(address), address, huge_flags);
// Tail: 4 KiB pages for whatever is left.
while (address < end) : (address += page_size)
mapPage(pml4, boot_handoff.physicalToVirtual(address), address, pte_flags);
}
fn regions(mm: boot_handoff.MemoryMap) []const boot_handoff.MemoryRegion {
@@ -118,11 +167,26 @@ fn enableNx() void {
io.wrmsr(efer_msr, io.rdmsr(efer_msr) | (1 << 11));
}
/// Program this core's PAT so entry 4 (selected by the PAT bit with PCD=PWT=0) is
/// **write-combining**, leaving the other seven at their reset types. Nothing else
/// in danos sets the PAT bit, so this changes no existing mapping — it only gives
/// the framebuffer a write-combining type, which turns its full-screen clear from
/// glacial (uncached writes to a GPU BAR, the real-hardware default via MTRRs) into
/// a batched burst. Must run on **every** core (PAT is per-logical-processor) — the
/// framebuffer mapping lives in the shared kernel half, so a core with the reset
/// PAT would see it as write-back and alias. Called from `init` (BSP) and each AP.
pub fn setupPat() void {
// Reset PAT is PA0=WB PA1=WT PA2=UC- PA3=UC PA4=WB PA5=WT PA6=UC- PA7=UC; flip
// PA4 from WB (0x06) to WC (0x01). Type codes: UC=0 WC=1 WT=4 WP=5 WB=6 UC-=7.
io.wrmsr(ia32_pat, 0x0007_0401_0007_0406);
}
/// Build the address space and switch onto it.
pub fn init(allocFrame: *const fn () ?u64, freeFrame: *const fn (u64) void, boot_information: *const boot_handoff.BootInformation) void {
alloc_frame = allocFrame;
free_frame = freeFrame;
enableNx();
setupPat(); // BSP: PAT entry 4 = write-combining, for the framebuffer window
const pml4 = allocTable();
// 1. All RAM in the physmap (physicalToVirtual(physical)) RW + NX. No identity/low-half
@@ -130,13 +194,15 @@ pub fn init(allocFrame: *const fn () ?u64, freeFrame: *const fn (u64) void, boot
// mapped on demand (mapMmio) or explicitly below.
for (regions(boot_information.memory_map)) |r| {
if (r.kind == .mmio) continue;
mapRangePhysmap(pml4, r.base, r.pages * page_size, present | writable | no_execute);
mapRangePhysmap(pml4, r.base, r.pages * page_size, present | writable | no_execute, false);
}
// 2. Physmap windows for the framebuffer and the Local APIC (device memory
// the kernel touches directly), RW + NX.
// the kernel touches directly), RW + NX. The framebuffer is **write-
// combining** (see setupPat) so the console's full-screen clear is a burst,
// not millions of uncached single-word writes.
const fb = boot_information.framebuffer;
mapRangePhysmap(pml4, fb.base, @as(u64, fb.height) * fb.pitch, present | writable | no_execute);
mapRangePhysmap(pml4, fb.base, @as(u64, fb.height) * fb.pitch, present | writable | no_execute, true);
mapPage(pml4, boot_handoff.physicalToVirtual(0xFEE00000), 0xFEE00000, present | writable | no_execute);
// 3. The kernel's own segments at their higher-half link addresses, mapped
@@ -254,8 +320,12 @@ pub fn mapUserInto(pml4: u64, virtual: u64, physical: u64, writable_page: bool,
/// RAM allocator (`freeSubtree`). RW + NX; the caller places `virtual` in a
/// user-exclusive range (PML4[225]). Both `virtual` and `physical` are page-aligned by
/// the caller; a sub-page `physical` offset is the caller's to re-apply.
pub fn mapUserDeviceInto(pml4: u64, virtual: u64, physical: u64, len: u64) void {
const flags: u64 = present | user | writable | no_execute | pcd | pwt | device_grant;
pub fn mapUserDeviceInto(pml4: u64, virtual: u64, physical: u64, len: u64, write_combining: bool) void {
// Registers are strong-uncacheable (PCD|PWT). A framebuffer instead wants
// write-combining — the PAT bit (bit 7 in a 4 KiB PTE) with PCD=PWT=0 selects PAT
// entry 4, which `setupPat` programs to WC — so pixel writes batch into bursts.
const cache: u64 = if (write_combining) pte_pat else (pcd | pwt);
const flags: u64 = present | user | writable | no_execute | device_grant | cache;
const first = physical & ~@as(u64, page_size - 1);
const last = (physical + (if (len == 0) 1 else len) - 1) & ~@as(u64, page_size - 1);
var off: u64 = 0;
@@ -296,6 +366,33 @@ pub fn mapUserDmaInto(pml4: u64, virtual: u64, physical: u64, len: u64) void {
}
}
/// The raw leaf entry mapping `virtual` in the address space rooted at `pml4`, or null
/// if any level of the walk is absent. **Read-only** — never allocates or descends into
/// a missing table (unlike the `map*` paths' `descendUser`). Stops at the first huge
/// leaf. For tests and introspection that need a page's actual flag bits.
pub fn leafEntryOf(pml4: u64, virtual: u64) ?u64 {
const l4 = tableAt(pml4)[(virtual >> 39) & 0x1FF];
if (l4 & present == 0) return null;
const l3 = tableAt(l4 & address_mask)[(virtual >> 30) & 0x1FF];
if (l3 & present == 0) return null;
if (l3 & page_size_bit != 0) return l3; // 1 GiB leaf
const l2 = tableAt(l3 & address_mask)[(virtual >> 21) & 0x1FF];
if (l2 & present == 0) return null;
if (l2 & page_size_bit != 0) return l2; // 2 MiB leaf
const l1 = tableAt(l2 & address_mask)[(virtual >> 12) & 0x1FF];
if (l1 & present == 0) return null;
return l1;
}
/// Is the 4 KiB leaf mapping `virtual` write-combining — the PAT bit set with PCD and
/// PWT clear, which `setupPat` makes PAT entry 4 (WC)? Null if unmapped. The device
/// mapping path (`mapUserDeviceInto`) always uses 4 KiB leaves, so bit 7 (`pte_pat`)
/// is the PAT selector in play.
pub fn leafIsWriteCombining(pml4: u64, virtual: u64) ?bool {
const e = leafEntryOf(pml4, virtual) orelse return null;
return (e & pte_pat != 0) and (e & pcd == 0) and (e & pwt == 0);
}
/// Create a new address space: a fresh PML4 with an empty user half and the
/// kernel's higher half shared in (copying PML4[256..512), whose entries point
/// at the kernel's PDPTs — pre-created at init and never restaled, so growth in
@@ -338,8 +435,8 @@ fn freeSubtree(physical: u64, level: u32) void {
}
/// Whether `virtual` is currently mapped **executable** — present with the NX bit
/// clear. Walks the 4-level tables (all danos mappings are 4 KiB, so no huge-page
/// case). Returns false if unmapped. Used for W^X checks in tests.
/// clear. Walks the 4-level tables, stopping at a 2 MiB huge-page leaf (the physmap
/// uses them). Returns false if unmapped. Used for W^X checks in tests.
pub fn isExecutable(virtual: u64) bool {
const pml4e = tableAt(kernel_pml4)[(virtual >> 39) & 0x1FF];
if (pml4e & present == 0) return false;
@@ -347,6 +444,7 @@ pub fn isExecutable(virtual: u64) bool {
if (pdpte & present == 0) return false;
const pde = tableAt(pdpte & address_mask)[(virtual >> 21) & 0x1FF];
if (pde & present == 0) return false;
if (pde & page_size_bit != 0) return pde & no_execute == 0; // 2 MiB huge leaf
const pte = tableAt(pde & address_mask)[(virtual >> 12) & 0x1FF];
if (pte & present == 0) return false;
return pte & no_execute == 0;
@@ -385,9 +483,9 @@ pub fn unmapInto(pml4: u64, virtual: u64) void {
/// Resolve a virtual address to a physical one in the address space rooted at
/// `pml4`, walking the tables through the physmap (CR3-independent — works for
/// any address space, not just the live one). Returns null if `virtual` is not
/// mapped at any level. All danos mappings are 4 KiB, so there is no huge-page
/// case. The foundation for cross-address-space copies and for munmap (which
/// needs the frame behind a user vaddr to free it).
/// mapped at any level. Stops at a 2 MiB huge-page leaf (the physmap uses them),
/// resolving the offset within it. The foundation for cross-address-space copies
/// and for munmap (which needs the frame behind a user vaddr to free it).
pub fn translateIn(pml4: u64, virtual: u64) ?u64 {
const pml4e = tableAt(pml4)[(virtual >> 39) & 0x1FF];
if (pml4e & present == 0) return null;
@@ -395,6 +493,8 @@ pub fn translateIn(pml4: u64, virtual: u64) ?u64 {
if (pdpte & present == 0) return null;
const pde = tableAt(pdpte & address_mask)[(virtual >> 21) & 0x1FF];
if (pde & present == 0) return null;
if (pde & page_size_bit != 0) // 2 MiB huge leaf: frame base is bits 51:21
return (pde & address_mask & ~@as(u64, huge_page_size - 1)) | (virtual & (huge_page_size - 1));
const pte = tableAt(pde & address_mask)[(virtual >> 12) & 0x1FF];
if (pte & present == 0) return null;
return (pte & address_mask) | (virtual & (page_size - 1));
+42 -4
View File
@@ -17,6 +17,12 @@ const Access = enum { port, mmio };
var access: Access = .port;
var base: u64 = 0x3F8; // COM1
/// Whether `init`/`reconfigure` found a *working* UART at `base`. False on a
/// legacy-free machine whose COM1 is decoded but dead: writing to it is then a
/// no-op, so `write` never spins waiting for a transmit register that will never
/// drain. Cleared until proven by the loopback probe.
var uart_present: bool = false;
fn portOut(p: u16, value: u8) void {
asm volatile ("outb %[value], %[p]"
:
@@ -57,6 +63,34 @@ pub fn init() void {
setRegister(3, 0x03); // 8 bits, no parity, one stop bit; DLAB off
setRegister(2, 0xC7); // enable + clear FIFO, 14-byte threshold
setRegister(4, 0x0B); // RTS/DSR set
uart_present = probe();
}
/// Detect a *working* UART by internal loopback: route the transmitter back to
/// the receiver (MCR bit 4), send a byte, and check it comes back. A port that is
/// merely decoded but has nothing behind it (the common case on a legacy-free
/// board that still answers I/O at 0x3F8) never echoes, so this returns false.
///
/// This matters for speed, not just correctness: a dead UART's line-status
/// register reads back 0x00, so its transmit-holding-empty bit never sets, and
/// `writeByte` would otherwise spin its full guard — tens of milliseconds — on
/// *every* logged byte. On real hardware that alone can add ~a minute to boot.
fn probe() bool {
const saved_mcr = register(4);
setRegister(4, 0x1E); // MCR: LOOP | OUT2 | OUT1 | RTS — internal loopback
setRegister(0, 0xAE); // push a distinctive byte into the loopback path
var guard: u32 = 0;
while (register(5) & 0x01 == 0 and guard < 10_000) : (guard += 1) {} // await Data Ready
const echo = register(0);
setRegister(4, saved_mcr); // restore the modem-control lines
return echo == 0xAE;
}
/// Whether a working UART was detected (see `probe`). The log sink stays
/// registered regardless — it simply does nothing until this is true — so a UART
/// that only `reconfigure` discovers (via SPCR) still starts logging.
pub fn present() bool {
return uart_present;
}
/// Point the console at the UART ACPI's SPCR table names (MMIO or I/O port) and
@@ -70,15 +104,19 @@ pub fn reconfigure(is_mmio: bool, address: u64) void {
}
fn writeByte(c: u8) void {
// Wait for the transmit-holding register to empty — but bounded, so an absent
// UART (whose line-status register reads back as 0x00) can't hang the kernel.
// Wait for the transmit-holding register to empty. `write` only reaches here
// for a UART the loopback probe proved live, so this bounds a momentary stall
// (e.g. deasserted flow control), not an absent port: ~5000 legacy-port reads
// is a few ms — comfortably longer than one 38400-baud byte-time (~260 µs).
var guard: u32 = 0;
while (register(5) & 0x20 == 0 and guard < 100_000) : (guard += 1) {}
while (register(5) & 0x20 == 0 and guard < 5_000) : (guard += 1) {}
setRegister(0, c);
}
/// Write bytes, translating LF to CRLF so terminals and logs line up.
/// Write bytes, translating LF to CRLF so terminals and logs line up. A no-op
/// when no working UART was detected, so a dead COM1 costs nothing per byte.
pub fn write(bytes: []const u8) void {
if (!uart_present) return;
for (bytes) |c| {
if (c == '\n') writeByte('\r');
writeByte(c);
@@ -173,6 +173,7 @@ fn delayMicros(us: u64) void {
/// signals the BSP, then jumps to the generic scheduler entry. Never returns.
fn apEntry(percpu: usize) callconv(.c) noreturn {
const cpu = boot_index;
paging.setupPat(); // this core's PAT: entry 4 = write-combining, to match the BSP
gdt.loadOnThisCpu(cpu); // this core's GDT (with its own TSS slot)
tss.setupThisCpu(cpu); // this core's TSS + IST stack, loaded into TR
idt.loadOnThisCpu(); // the shared IDT
+16 -3
View File
@@ -20,6 +20,12 @@ const boot_handoff = @import("boot-handoff");
var con: Console = undefined;
var con_present: bool = false;
/// Set while a user-space display service owns the framebuffer: `write` falls silent so
/// the kernel doesn't paint over the compositor. Driven by the display device's
/// claim/release (system/kernel/process.zig). The terminal panic/exception paths clear
/// it first (`setSuppressed(false)`) — a dying machine's message wins over any display.
var suppressed: bool = false;
/// Set up the console over `fb`, or mark it absent if there's no usable
/// framebuffer. Clears the screen when present.
pub fn init(fb: boot_handoff.Framebuffer) void {
@@ -41,13 +47,20 @@ pub fn present() bool {
return con_present;
}
/// Output sink: draw `bytes` on screen. A no-op when no framebuffer is present,
/// so it's always safe to call.
/// Output sink: draw `bytes` on screen. A no-op when no framebuffer is present, or
/// while a display service owns the screen (`suppressed`), so it's always safe to call.
pub fn write(bytes: []const u8) void {
if (!con_present) return;
if (!con_present or suppressed) return;
for (bytes) |c| con.putChar(c);
}
/// Quiesce (or resume) the bootstrap console. Set true when a display service claims the
/// framebuffer; set false when that claim is released, or by the panic path to force a
/// last message onto a screen a (now-irrelevant) service was holding.
pub fn setSuppressed(value: bool) void {
suppressed = value;
}
/// The console font, embedded at compile time. cp850-8x16, PSF2 format:
/// a 32-byte header, then 256 glyphs of 16 bytes each (one byte per 8-pixel
/// row). We index glyphs straight by byte value, so ASCII maps 1:1.
+50
View File
@@ -36,6 +36,11 @@ var devices: [maximum_devices]device_abi.DeviceDescriptor = undefined;
var claimed: [maximum_devices]?u32 = .{null} ** maximum_devices; // owner task id, or null
var count: usize = 0;
/// The id of the seeded framebuffer node (`seedDisplay`), or null when the machine
/// handed over no framebuffer. Lets the process layer recognise the display claim
/// (to quiesce the bootstrap console) without threading the id through every caller.
var display_device: ?u64 = null;
/// Devices discovery found but the table had no room for. Non-zero means the machine
/// is bigger than `maximum_devices` and some hardware is simply invisible to drivers —
/// which would otherwise be an entirely silent failure. Logged at boot.
@@ -45,10 +50,55 @@ pub var dropped: usize = 0;
pub fn init(device_tree: *const platform.DeviceTree) void {
count = 0;
dropped = 0;
display_device = null;
for (&claimed) |*c| c.* = null;
walk(device_tree.root, device_abi.no_parent);
}
/// Publish the loader's framebuffer as a `display` device — a root-level node with one
/// write-combining `memory` resource over the linear framebuffer and its geometry in
/// `.display`. The framebuffer is *not* firmware-discovered (it rides the
/// [[boot-handoff]], not the device tree), so it is seeded explicitly, after `init`.
/// Returns the new device id, or null when there is no framebuffer (headless) or the
/// table is full. Idempotent-ish: only ever call once per boot.
pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32) ?u64 {
if (base == 0 or width == 0 or height == 0) return null; // headless
if (count >= maximum_devices) {
dropped += 1;
return null;
}
var d = std.mem.zeroes(device_abi.DeviceDescriptor);
d.id = count;
d.parent = device_abi.no_parent;
d.class = @intFromEnum(device_abi.DeviceClass.display);
d.pci_class = device_abi.no_pci_class;
d.resource_count = 1;
d.resources[0] = .{
.kind = @intFromEnum(device_abi.ResourceKind.memory),
.start = base,
.len = @as(u64, height) * pitch,
.flags = device_abi.resource_flag_write_combining,
};
d.display = .{ .width = width, .height = height, .pitch = pitch, .format = format };
devices[count] = d;
display_device = d.id;
count += 1;
return d.id;
}
/// The id of the seeded framebuffer device, or null when none was seeded.
pub fn displayDevice() ?u64 {
return display_device;
}
/// Whether the framebuffer device is currently claimed by some process. The bootstrap
/// console uses this (via the process layer) to fall silent while a display service
/// owns the screen, and to resume if that service dies and its claim is released.
pub fn displayClaimed() bool {
const id = display_device orelse return false;
return ownerOf(id) != null;
}
/// Record `node` (unless it's the synthetic root) and recurse, threading the id we
/// assigned it down to its children as their parent.
fn walk(node: *platform.Device, parent_id: u64) void {
+3 -1
View File
@@ -37,7 +37,9 @@ const Task = scheduler.Task;
pub const MESSAGE_MAXIMUM: usize = 256;
pub const maximum_handles = scheduler.ipc_maximum_handles;
pub const maximum_services = 8;
// The name registry is indexed directly by ServiceId, so this must exceed the
// largest id (currently fat = 8). Sized with headroom for new services.
pub const maximum_services = 16;
/// Errno-style failures, returned as `-value` in the system_call result register.
pub const EBADF: i64 = 1; // bad handle
+51 -16
View File
@@ -5,6 +5,7 @@ const parameters = @import("parameters");
const architecture = @import("architecture");
const console = @import("console.zig");
const log = @import("log.zig");
const wall_clock = @import("wall-clock.zig");
const pmm = @import("pmm.zig");
const heap = @import("heap.zig");
const scheduler = @import("scheduler.zig");
@@ -59,16 +60,34 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// file on a ramdisk/USB/SSD), so a message survives as long as any is present.
// A headless, serial-less machine still boots correctly — it just goes quiet,
// with port-0x80 checkpoints as the only progress signal.
architecture.serialInit();
log.addSink(architecture.serialWrite);
//
// Serial is compiled in only under -Dserial (build.zig): a real machine often
// has no live legacy COM1, and the log survives in the RAM buffer (below) and
// is flushed to disk — so serial is now a QEMU/dev convenience the flashable
// image leaves out. When it *is* built in, `serialInit`'s loopback probe still
// guards against a dead port (so a -Dserial image is safe on real hardware).
if (build_options.serial) {
architecture.serialInit();
log.addSink(architecture.serialWrite);
}
if (architecture.debugconPresent()) log.addSink(architecture.debugconWrite);
// Retain the whole stream in a RAM buffer too, so a user program can later
// read it back (klog_read) and persist the boot log to disk — the only way to
// see it on a headless/real machine with no host capturing serial.
log.addSink(log.ramSink);
// The **framebuffer** is deliberately *not* a log sink. It's a separate output
// surface — a bootstrap text console today, a graphics device driver later — so
// we never assume the OS is text-based. Only a few user-facing status lines
// (via `status`) and panics are mirrored to it; the verbose log stays out.
//
// The console is brought up *after* paging (below), not here: its one-time
// full-screen clear then runs on the kernel's **write-combining** mapping of the
// framebuffer instead of the loader's uncached one — a fast burst rather than
// millions of uncached writes on real hardware. Until then, on-screen output is
// absent (an early panic still lands in the serial/RAM log); the trade is worth
// a near-instant boot. `console.write` is a safe no-op while the console is down.
const fb = boot_information.framebuffer;
console.init(fb);
log.checkpoint(cp_entry);
@@ -78,10 +97,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
architecture.init();
status("/system/kernel: initialising kernel...\n");
log.write(if (console.present())
"/system/kernel: framebuffer console online (bootstrap; graphics driver later)\n"
if (build_options.serial) log.write(if (architecture.serialPresent())
"/system/kernel: serial console online (COM1)\n"
else
"/system/kernel: no framebuffer (headless) -> logging to serial/debugcon only\n");
"/system/kernel: no serial UART (COM1 absent) -> log kept in RAM/debugcon\n");
log.write("/system/kernel: cpu tables online (GDT, IDT, TSS)\n");
log.print(" resolution : {d}x{d}\n", .{ fb.width, fb.height });
log.print(" pitch : {d} bytes\n", .{fb.pitch});
@@ -137,6 +156,15 @@ fn kmain(boot_information: *const BootInformation) noreturn {
log.print(" page tables: root = 0x{x:0>16}\n", .{architecture.activePageTable()});
log.print(" kernel segs: {d} (mapped with W^X permissions)\n", .{boot_information.kernel_segment_count});
// Now on our own tables, the framebuffer window is write-combining: bring up
// the on-screen console and clear it (a fast burst here, not the loader's
// uncached crawl). From here `status` reaches the screen as well as the log.
console.init(fb);
log.write(if (console.present())
"/system/kernel: framebuffer console online (bootstrap; graphics driver later)\n"
else
"/system/kernel: no framebuffer (headless) -> logging to serial/debugcon only\n");
// Bring up the kernel heap (dynamic allocation), built on the VMM.
heap.init();
log.checkpoint(cp_heap);
@@ -167,25 +195,24 @@ fn kmain(boot_information: *const BootInformation) noreturn {
log.print("/system/kernel: WARNING {d} device(s) dropped — table full\n", .{devices_broker.dropped});
}
// Publish the loader's framebuffer as a claimable `display` device, so a
// user-space display service can take it over the same claim + mmio_map path as
// any other hardware (it is not firmware-discovered; it rides the boot handoff).
if (devices_broker.seedDisplay(fb.base, fb.width, fb.height, fb.pitch, @intFromEnum(fb.format))) |display_id| {
log.print("/system/kernel: framebuffer device {d} seeded ({d}x{d}, pitch {d}, write-combining)\n", .{ display_id, fb.width, fb.height, fb.pitch });
}
// Install the device-IRQ trampolines, so a driver's irq_bind has vectors to
// land on. Every line stays masked until something binds it (ioapic.init).
irq.init();
// Power register map extracted from the FADT + AML, for confidence it parsed.
// Power register map, from the FADT (the SLP_TYP sleep values live in AML,
// which the kernel doesn't parse — the ring-3 acpi service owns soft-off).
const pw = platform.powerInformation();
log.write("/system/kernel: power\n");
log.print(" pm1a_cnt : {s} 0x{x} (width {d})\n", .{ if (pw.pm1a_cnt.mmio) "mmio" else "io", pw.pm1a_cnt.address, pw.pm1a_cnt.width });
if (pw.s5) |s| {
log.print(" S5 slp_typ : a={d} b={d}\n", .{ s.slp_typ_a, s.slp_typ_b });
} else {
log.write(" S5 slp_typ : (not found)\n");
}
log.print(" reset : supported={} {s} 0x{x} val 0x{x}\n", .{ pw.reset_supported, if (pw.reset.mmio) "mmio" else "io", pw.reset.address, pw.reset_value });
// AML namespace parse integrity: consumed should equal total.
const am = platform.amlStats();
log.print(" aml : {d} namespace nodes, parsed {d}/{d} bytes\n", .{ am.nodes, am.consumed, am.total });
// Feed the architecture layer the discovered addresses/facts so it makes no legacy
// assumptions — the point of all this on UEFI Class 3 firmware. MMIO bases
// (HPET, I/O APIC) come from the device tree; scalar facts from ACPI.
@@ -281,6 +308,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
if (!architecture.clockSynchronized())
log.write("/system/kernel: WARNING: per-core TSCs are not synchronized; monotonic clock moved off the TSC\n");
// Anchor wall-clock time: read the RTC once, now the monotonic clock is final.
wall_clock.init();
log.print("/system/kernel: wall clock {d} (Unix epoch seconds, UTC, from the RTC)\n", .{wall_clock.nowSeconds()});
// In a test build (`zig build -Dtest-case=<name>`), run that case and stop.
// Normal builds fall through to the idle halt.
if (build_options.test_case) |case| {
@@ -455,6 +486,9 @@ fn onException(state: *const architecture.CpuState) noreturn {
}
log.checkpoint(cp_exception);
// The machine is going down: force the console back on even if a display service
// was holding the framebuffer, so the exception actually reaches the screen.
console.setSuppressed(false);
const core = scheduler.currentCpuIndex();
// A fault is user-facing enough to paint on screen too (via statusPrint), on
// top of the diagnostic log.
@@ -477,6 +511,7 @@ pub const panic = std.debug.FullPanic(struct {
_ = first_trace_address;
log.checkpoint(cp_panic);
log.recordPanic(message);
console.setSuppressed(false); // a panic outranks any display service holding the screen
status("\nKERNEL PANIC: ");
status(message);
status("\n");
+33
View File
@@ -41,6 +41,39 @@ pub fn write(bytes: []const u8) void {
for (sinks[0..sink_count]) |sink| sink(bytes);
}
// --- the RAM sink: a retained copy of the whole diagnostic stream ------------
//
// A fixed in-image buffer that accumulates every logged byte, so a user program
// (`log-flush`, and init at shutdown) can read it back through `klog_read` and
// persist it to a file — the boot log survives on a headless/real machine that
// has no host capturing serial. It is a *sink like any other*: register it with
// `addSink(ramSink)` at boot. No allocation (works pre-heap and in a panic).
//
// It fills linearly and stops when full: the earliest output — the most valuable
// for diagnosing a boot — is kept, and the tail is still on the live serial sink.
// 256 KiB comfortably holds a full boot plus a long run (a boot is ~15 KiB).
const ram_capacity = 256 * 1024;
var ram_buffer: [ram_capacity]u8 = undefined;
var ram_len: usize = 0;
/// The RAM sink. Best-effort and self-guarding like every sink: appends what fits
/// and silently drops the rest once full. (Concurrency matches the other sinks —
/// the dominant writer, debug_write, already holds the kernel lock; a rare torn
/// append on a kernel-internal line is an accepted diagnostic imperfection.)
pub fn ramSink(bytes: []const u8) void {
const n = @min(ram_buffer.len - ram_len, bytes.len);
if (n != 0) {
@memcpy(ram_buffer[ram_len..][0..n], bytes[0..n]);
ram_len += n;
}
}
/// The accumulated log so far — what `klog_read` copies out.
pub fn ramSnapshot() []const u8 {
return ram_buffer[0..ram_len];
}
/// A formatted log line. Truncates past 256 bytes; the buffer is on the stack, so
/// this is safe to call from interrupt context and from a panic.
pub fn print(comptime fmt: []const u8, args: anytype) void {
+81 -19
View File
@@ -28,12 +28,14 @@ const parameters = @import("parameters");
const architecture = @import("architecture");
const pmm = @import("pmm.zig");
const scheduler = @import("scheduler.zig");
const console = @import("console.zig");
const sync = @import("sync.zig");
const ipc = @import("ipc-synchronous.zig");
const devices_broker = @import("devices-broker.zig");
const irq = @import("irq.zig");
const initial_ramdisk = @import("initial-ramdisk");
const log = @import("log.zig");
const wall_clock = @import("wall-clock.zig");
const page_size = abi.page_size;
const SystemCall = abi.SystemCall;
@@ -79,9 +81,11 @@ pub const device_arena_end: u64 = device_arena_base + (4 << 30);
pub const dma_arena_base: u64 = 0x0000_7200_0000_0000;
pub const dma_arena_end: u64 = dma_arena_base + (256 << 20); // 256 MiB per process
/// Largest single `mmap` grant, in pages (1 MiB). The user heap grows in small
/// chunks, so this bound is generous; it also caps the frame scratch array below.
const maximum_mmap_pages = 256;
/// Largest single `mmap` grant, in pages (32 MiB). Big enough for a display service's
/// back buffer at up to 4K (3840x2160x4 ≈ 8100 pages); the user heap otherwise grows in
/// small chunks. `systemMmap` maps page by page with rollback, so this is only a sanity
/// bound (and an overflow guard on the page count), not the size of any scratch array.
const maximum_mmap_pages = 8192;
/// Ceiling on a process's argv entries, including argv[0]. Arguments are spawn
/// parameters ("you are the driver for device 12"), not bulk data — IPC carries
@@ -206,6 +210,8 @@ fn system_call(state: *architecture.CpuState) void {
.signal_bind => systemSignalBind(state),
.process_signal => systemProcessSignal(state),
.timer_bind => systemTimerBind(state),
.klog_read => systemKlogRead(state),
.wall_clock => systemWallClock(state),
_ => fail(state),
}
}
@@ -298,12 +304,19 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
/// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process.
fn systemDeviceClaim(state: *architecture.CpuState) void {
const device_id = architecture.systemCallArg(state, 0);
const claim_flags = sync.enter();
defer sync.leave(claim_flags);
if (devices_broker.claim(architecture.systemCallArg(state, 0), scheduler.current().id))
architecture.setSystemCallResult(state, 0)
else
fail(state);
if (devices_broker.claim(device_id, scheduler.current().id)) {
// A display service just took the framebuffer — quiesce the bootstrap console
// so the kernel and the service don't scribble over each other's pixels. The
// claim releases (and the console resumes) automatically if the service dies;
// see releaseTaskResourcesLocked.
if (devices_broker.displayDevice()) |display_id| {
if (device_id == display_id) console.setSuppressed(true);
}
architecture.setSystemCallResult(state, 0);
} else fail(state);
}
/// mmio_map(device_id, resource_index) -> vaddr: map a claimed device's MMIO window into
@@ -338,7 +351,10 @@ fn systemMmioMap(state: *architecture.CpuState) void {
const base_v = t.device_map_next;
if (base_v + pages * page_size > device_arena_end) return fail(state);
architecture.mapUserDeviceInto(t.aspace, base_v, r.start, r.len);
// A framebuffer resource asks (via its flag) to be mapped write-combining rather
// than the strong-uncacheable default that register MMIO needs.
const write_combining = (r.flags & device_abi.resource_flag_write_combining) != 0;
architecture.mapUserDeviceInto(t.aspace, base_v, r.start, r.len, write_combining);
t.device_map_next = base_v + pages * page_size;
architecture.setSystemCallResult(state, base_v + (r.start & (page_size - 1))); // register base
}
@@ -598,6 +614,9 @@ fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
recordExitLocked(t);
irq.releaseOwner(t.id);
devices_broker.releaseAllOwnedBy(t.id);
// If that dropped the framebuffer claim (this task was the display service), let the
// bootstrap console draw again — the screen is nobody's now, so panics/status land.
if (!devices_broker.displayClaimed()) console.setSuppressed(false);
// The dying task's signal endpoint and one-shot timers go with it.
if (t.signal_endpoint) |raw| {
ipc.dropRef(@ptrCast(@alignCast(raw)));
@@ -974,6 +993,37 @@ fn systemDebugWrite(state: *architecture.CpuState) void {
}
}
/// klog_read(offset, ptr, len) -> bytes copied: copy the kernel's in-memory
/// diagnostic log (the RAM sink in log.zig) out to the user buffer at `ptr`,
/// starting at `offset`. Returns the count copied — 0 once `offset` reaches the
/// end — so a program reads the whole log by looping from 0 until it gets 0.
///
/// The mirror of `debug_write`: the same overflow-safe user-half bounds check,
/// but the copy runs kernel -> user. Written under the kernel lock so the source
/// snapshot can't grow underneath the copy. A read-only diagnostic — it exposes
/// only the log the kernel already broadcasts to serial, nothing else.
fn systemKlogRead(state: *architecture.CpuState) void {
const offset = architecture.systemCallArg(state, 0);
const ptr = architecture.systemCallArg(state, 1);
const len = architecture.systemCallArg(state, 2);
// Confine the whole destination span to the user (low) half. `len <=
// user_half_end - ptr` bounds the length without an overflowing add.
if (ptr < user_half_end and len <= user_half_end - ptr) {
const flags = sync.enter();
defer sync.leave(flags);
const snapshot = log.ramSnapshot();
var n: usize = 0;
if (offset < snapshot.len) {
n = @min(len, snapshot.len - offset);
const dest: [*]u8 = @ptrFromInt(ptr);
@memcpy(dest[0..n], snapshot[offset..][0..n]);
}
architecture.setSystemCallResult(state, n);
} else {
fail(state);
}
}
/// mmap(len, prot) -> base: grant `len` bytes (rounded up to whole pages) of
/// fresh, zeroed, writable+NX memory in the caller's mmap arena, and return the
/// base virtual address. `prot` is accepted but not yet honoured (grants are
@@ -990,21 +1040,26 @@ fn systemMmap(state: *architecture.CpuState) void {
const base = t.heap_next;
if (base + pages * page_size > heap_arena_end) return fail(state); // arena exhausted
// Reserve all frames up front so a mid-way exhaustion rolls back cleanly
// (no partially-mapped grant leaks into the address space).
var frames: [maximum_mmap_pages]u64 = undefined;
var got: usize = 0;
while (got < pages) : (got += 1) {
frames[got] = pmm.alloc() orelse {
for (frames[0..got]) |f| pmm.free(f);
// Map page by page. On mid-way frame exhaustion, roll back the pages already mapped
// (unmap + free) so no partial grant leaks into the address space — the same
// all-or-nothing guarantee as before, but without a fixed scratch array, so the
// per-call size can be a multi-MiB framebuffer.
var mapped: usize = 0;
while (mapped < pages) : (mapped += 1) {
const frame = pmm.alloc() orelse {
var i: usize = 0;
while (i < mapped) : (i += 1) {
const va = base + i * page_size;
if (architecture.translate(t.aspace, va)) |physical| {
architecture.unmapUserPageInto(t.aspace, va);
pmm.free(physical);
}
}
return fail(state);
};
}
for (frames[0..pages], 0..) |frame, i| {
const destination: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(frame));
@memset(destination[0..page_size], 0); // hand out zeroed memory
architecture.mapUserPageInto(t.aspace, base + i * page_size, frame, true, false); // RW + NX
architecture.mapUserPageInto(t.aspace, base + mapped * page_size, frame, true, false); // RW + NX
}
t.heap_next = base + pages * page_size;
architecture.setSystemCallResult(state, base);
@@ -1304,3 +1359,10 @@ pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []c
fn systemClock(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, architecture.nanos());
}
/// wall_clock() -> Unix epoch seconds (UTC). The RTC value, read at boot and offset
/// by the monotonic clock (wall-clock.zig) — mechanism, not policy: calendars and
/// timezones layer on top in user space. Needed for filesystem timestamps (mtime).
fn systemWallClock(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, wall_clock.nowSeconds());
}
+246 -36
View File
@@ -14,6 +14,7 @@ const boot_handoff = @import("boot-handoff");
const abi = @import("abi");
const device_abi = @import("device-abi");
const architecture = @import("architecture");
const wall_clock = @import("wall-clock.zig");
const devices_broker = @import("devices-broker.zig");
const platform = @import("platform");
const pmm = @import("pmm.zig");
@@ -66,6 +67,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
timer();
} else if (eql(case, "clock")) {
clock();
} else if (eql(case, "wall-clock")) {
wallClock();
} else if (eql(case, "vmm")) {
vmm();
} else if (eql(case, "heap")) {
@@ -92,6 +95,12 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
iommuTest();
} else if (eql(case, "ioport")) {
ioPortTest();
} else if (eql(case, "display")) {
displayTest(boot_information);
} else if (eql(case, "display-service")) {
displayServiceTest(boot_information);
} else if (eql(case, "display-demo")) {
displayDemoTest(boot_information);
} else if (eql(case, "clock")) {
clockTest();
} else if (eql(case, "smp")) {
@@ -144,6 +153,12 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
driverRestartTest(boot_information);
} else if (eql(case, "usb-report")) {
usbReportTest(boot_information);
} else if (eql(case, "usb-hid")) {
usbHidTest(boot_information);
} else if (eql(case, "usb-storage")) {
usbStorageTest(boot_information);
} else if (eql(case, "fat-mount")) {
fatMountTest(boot_information);
} else if (eql(case, "device-list")) {
deviceListTest(boot_information);
} else if (eql(case, "pci-scan")) {
@@ -172,10 +187,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
containmentTest();
} else if (eql(case, "device-manager")) {
deviceManagerTest(boot_information);
} else if (eql(case, "poweroff")) {
powerTest(.off);
} else if (eql(case, "reboot")) {
powerTest(.reboot);
rebootTest();
} else {
log("DANOS-TEST-RESULT: FAIL (unknown case '{s}')\n", .{case});
}
@@ -192,15 +205,14 @@ fn platformHal() platform.Hal {
/// Drive an ACPI power transition. On success the machine powers off or resets,
/// so QEMU exits — the harness observes the process exit. If control returns, the
/// transition failed and we emit a FAIL result.
fn powerTest(comptime action: enum { off, reboot }) void {
const name = if (action == .off) "poweroff" else "reboot";
log("DANOS-TEST-BEGIN: {s}\n", .{name});
// Soft-off (S5) is no longer a kernel operation — the ring-3 acpi service owns it
// (exercised end-to-end by `orderly-shutdown`). Reboot stays in the kernel (FADT
// reset register, no AML), so it keeps its own case.
fn rebootTest() void {
log("DANOS-TEST-BEGIN: reboot\n", .{});
const hal = platformHal();
log("DANOS-POWER: attempting {s}\n", .{name});
switch (action) {
.off => platform.shutdown(hal),
.reboot => platform.reboot(hal),
}
log("DANOS-POWER: attempting reboot\n", .{});
platform.reboot(hal);
check("power transition took effect", false);
result();
}
@@ -267,22 +279,19 @@ fn timer() void {
}
/// Verify device discovery populated the platform facts the rest of the kernel
/// depends on — the results ACPI parsing stashed in globals at boot. These are
/// stable for the QEMU q35 + OVMF machine the harness runs, and span the tables:
/// MADT (LAPIC base, CPU count), FADT (PM/reset registers), and the AML parse
/// (the sleep type, plus the integrity check that every byte was consumed).
/// depends on — the results the static ACPI tables stashed in globals at boot.
/// These are stable for the QEMU q35 + OVMF machine the harness runs, and span the
/// tables: MADT (LAPIC base, CPU count) and FADT (PM/reset registers). The kernel
/// no longer interprets AML — sleep types are the ring-3 acpi service's concern.
fn discoveryTest() void {
log("DANOS-TEST-BEGIN: discovery\n", .{});
const pinfo = platform.platformInformation();
const pw = platform.powerInformation();
const am = platform.amlStats();
check("LAPIC base discovered (MADT)", pinfo.lapic_base == 0xFEE00000);
check("ACPI PM timer found (FADT)", pinfo.pm_timer.present());
check("PM1a control register found (FADT)", pw.pm1a_cnt.present());
check("reset register supported (FADT)", pw.reset_supported);
check("S5 sleep type found (AML)", pw.s5 != null);
check("AML parsed completely (consumed == total)", am.total > 0 and am.consumed == am.total);
check("at least one CPU enumerated (MADT)", platform.cpus().len >= 1);
// M15: every PCI function now carries its own 4 KiB ECAM configuration space as
@@ -452,6 +461,17 @@ fn heapTest() void {
/// Verify the calibrated clocks: sane measured frequencies, monotonic uptime that
/// advances with real ticks, and — the point of the TSC clock — nanosecond
/// resolution far finer than the 1 ms tick, with the unit functions consistent.
fn wallClock() void {
log("DANOS-TEST-BEGIN: wall-clock\n", .{});
// The RTC was read and anchored at boot (kmain -> wall_clock.init()).
const seconds = wall_clock.nowSeconds();
log(" epoch: {d}\n", .{seconds});
// A plausible current wall-clock: after 2020-01-01 (1577836800) and before 2050
// (2524608000) — catches a broken CMOS read or a wrong epoch conversion.
check("wall clock reads a plausible current epoch", seconds > 1_577_836_800 and seconds < 2_524_608_000);
result();
}
fn clock() void {
log("DANOS-TEST-BEGIN: clock\n", .{});
@@ -1968,6 +1988,66 @@ fn pciScanTest(boot_information: *const BootInformation) void {
/// the acpi service publishes it; init runs the stop sequence over its children
/// and asks the power service for S5; the machine powers off (QEMU exits). The
/// kernel test only spawns init — the ordered chain is the harness assertion.
/// The USB HID chain, end to end: boot the full service tree (init spawns vfs,
/// input, device-manager), and let discovery run — the manager matches the PCI
/// host bridge to pci-bus, pci-bus reports the xHCI controller, usb-xhci-bus
/// enumerates the HID interfaces, and the manager spawns the class drivers. The
/// harness's expect regex requires usb-xhci-bus to register the boot-keyboard
/// interface, the manager to spawn usb-hid-keyboard, and that driver to come up
/// (open its device, ask for boot protocol, subscribe) — proof the transfer
/// protocol works class-driver to controller.
fn usbHidTest(boot_information: *const BootInformation) void {
bootServiceTreeTest(boot_information, "usb-hid");
}
/// The USB storage chain: same full-tree boot, but the harness attaches a
/// usb-storage device and the expect regex requires usb-storage to come up
/// (open its device, run the BOT bring-up, read its capacity, and read block 0).
fn usbStorageTest(boot_information: *const BootInformation) void {
bootServiceTreeTest(boot_information, "usb-storage");
}
/// The FAT mount chain: boot the full tree (init spawns the fat server, which
/// brings up the USB storage chain, mounts the FAT volume, and mounts itself into
/// the VFS at /mnt/usb), then spawn a fat-test client that lists and reads through
/// the mount. The harness attaches a usb-storage device; the expect regex requires
/// the fat mount and the client's success.
fn fatMountTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: fat-mount\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over init and the initial_ramdisk", false);
result();
return;
}
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(ramdisk);
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const init_ok = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned (boots the tree, incl. the fat server)", init_ok);
check("fat-test client spawned", spawnNamed(rd, "fat-test"));
result();
}
fn bootServiceTreeTest(boot_information: *const BootInformation, comptime label: []const u8) void {
log("DANOS-TEST-BEGIN: " ++ label ++ "\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over init and the initial_ramdisk", false);
result();
return;
}
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
process.setInitialRamdisk(ramdisk);
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned (boots vfs, input, device-manager, and the USB chain)", spawned);
result();
}
fn orderlyShutdownTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: orderly-shutdown\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
@@ -2015,11 +2095,11 @@ fn acpiReportTest(boot_information: *const BootInformation) void {
result();
}
/// M20.1: the ring-3 AML parse agrees with the kernel's. The manager spawns
/// the discovery service (the acpi build variant); it claims the acpi-tables
/// node, maps the blobs, parses them, and logs its Device count — which must
/// equal what the kernel's own parse produced (the equivalence that licenses
/// retiring the kernel's device build in M20.3).
/// M20.1: the ring-3 AML parse works. The manager spawns the discovery service
/// (the acpi build variant); it claims the acpi-tables node, maps the blobs,
/// parses them, and self-verifies it found at least a floor of Device objects,
/// printing "acpi-parse: ok". The kernel no longer parses AML, so there is no
/// kernel count to compare against — the ring-3 parse is now the only one.
fn acpiParseTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: acpi-parse\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
@@ -2034,23 +2114,18 @@ fn acpiParseTest(boot_information: *const BootInformation) void {
return;
};
// The kernel's own count, from the namespace it already built for \_S5.
const kernel_devices = platform.amlDeviceCount();
check("the kernel namespace has devices to compare against", kernel_devices >= 1);
// Spawn the discovery service directly with that count as argv: it parses
// the same blobs in ring 3 and self-verifies, printing "acpi-parse: ok" iff
// the counts match. The harness's expect regex is that marker — deterministic,
// no racing the shared serial buffer.
// Spawn the discovery service directly with a device-count *floor* as argv:
// it parses the blobs in ring 3 and self-verifies it found at least that many
// Device objects, printing "acpi-parse: ok". A floor of 1 just proves the
// parser ran and produced a namespace (the QEMU q35 DSDT has dozens). The
// marker is deterministic — no racing the shared serial buffer.
process.setInitialRamdisk(image);
var count_text: [16]u8 = undefined;
const count_arg = std.fmt.bufPrint(&count_text, "{d}", .{kernel_devices}) catch "0";
var spawned = false;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "discovery")) continue;
_ = process.spawnProcessSupervised(item.blob, 4, &.{ "discovery", count_arg }, scheduler.currentId(), null) catch 0;
_ = process.spawnProcessSupervised(item.blob, 4, &.{ "discovery", "1" }, scheduler.currentId(), null) catch 0;
spawned = true;
break;
}
@@ -2236,6 +2311,72 @@ fn inputTest(boot_information: *const BootInformation) void {
result();
}
/// D2 — the display service comes up. Spawn it from the initial_ramdisk; it claims the
/// framebuffer the kernel seeded (D1), maps it write-combining, allocates a cacheable
/// back buffer, and proves the double-buffer path by clearing that buffer and presenting
/// it. Its `display: online WxH` + `display: presented frame 0` heartbeats are the
/// markers — seeing them proves a user-space compositor took the framebuffer and pushed
/// a whole composed frame to the screen, without ever drawing straight to the LFB.
fn displayServiceTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: display-service\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
// Spawn the compositor and hand it the core. Its own serial heartbeats — `display:
// online WxH` and `display: presented frame 0` — are what the harness matches (it
// reads serial directly, like the fault cases). We don't poll for them in-kernel: a
// single service that comes up and blocks doesn't reschedule this bring-up context
// (there is no other runnable task to bounce control back through), so the honest
// observation point is the service's output itself, not a check() proxy here.
if (!spawnNamed(rd, "display")) {
log("display-service: could not spawn the display service\n", .{});
result();
return;
}
scheduler.setPriority(1); // below the service, so it runs and comes up first
while (true) scheduler.yield();
}
/// D4 — a separate process drives the compositor. Spawn the display service and the
/// hardware-free `display-demo` client, which creates a wallpaper, a moving rectangle,
/// and a cursor and presents a run of frames. Its `display-demo: ok` heartbeat — printed
/// only after it drove frames of motion through the layer client API and the compositor —
/// is the harness's marker (the visible motion itself is a screenshot away via run-x86-64).
/// The demo keeps presenting, so unlike a lone blocking service the scheduler stays busy;
/// we still match on serial rather than poll, for consistency.
fn displayDemoTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: display-demo\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
if (!spawnNamed(rd, "display")) {
log("display-demo: could not spawn the display service\n", .{});
result();
return;
}
_ = spawnNamed(rd, "display-demo");
scheduler.setPriority(1); // below the service + demo, so they run
while (true) scheduler.yield();
}
/// Process arguments, end to end: spawn args-echo bare (its argv[0] is the
/// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through
/// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument
@@ -2558,8 +2699,8 @@ fn ioPassTest() void {
result();
return;
};
// Map it the way mmio_map does (device grant), then tear the space down.
architecture.mapUserDeviceInto(aspace, process.device_arena_base, frame, abi.page_size);
// Map it the way mmio_map does (device grant, strong-uncacheable), then tear the space down.
architecture.mapUserDeviceInto(aspace, process.device_arena_base, frame, abi.page_size, false);
architecture.destroyAddressSpace(aspace);
// The page tables were reclaimed; the device-granted frame must not have been.
@@ -2569,6 +2710,75 @@ fn ioPassTest() void {
result();
}
/// D1 — the framebuffer handoff primitive. The kernel seeds the loader's framebuffer as
/// a claimable `display` device with a write-combining `memory` resource; a display
/// service reaches it over the ordinary claim + mmio_map path. Prove the whole chain:
/// the node is present and correctly shaped, it maps, and — the point of D1 — the
/// mapping is genuinely write-combining, not the strong-uncacheable default that would
/// make a framebuffer blit glacial.
fn displayTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: display\n", .{});
const fb = boot_information.framebuffer;
if (!fb.present()) {
// Headless: nothing to seed. Not a failure of the mechanism, so pass cleanly.
log("display: no framebuffer (headless); skipping\n", .{});
result();
return;
}
// The kernel seeded a display device in kmain, right after devices_broker.init.
const display_id = devices_broker.displayDevice() orelse {
check("a framebuffer display device was seeded", false);
result();
return;
};
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
check("the seeded display id is enumerable", display_id < n);
if (display_id >= n) {
result();
return;
}
const d = buffer[@intCast(display_id)];
check("the node is class display", d.class == @intFromEnum(device_abi.DeviceClass.display));
check("it carries the framebuffer geometry", d.display.width == fb.width and d.display.height == fb.height and d.display.pitch == fb.pitch);
check("it has exactly one resource", d.resource_count == 1);
const r = d.resources[0];
check("that resource is a memory window", r.kind == @intFromEnum(device_abi.ResourceKind.memory));
check("it spans the whole framebuffer", r.start == fb.base and r.len == @as(u64, fb.height) * fb.pitch);
check("it is flagged write-combining", (r.flags & device_abi.resource_flag_write_combining) != 0);
// Walk the real claim + map path a display service would, into a throwaway address
// space, and confirm the leaf's cache type. We never run this space (no CR3 load) —
// we only read back the page-table entries — so aliasing the same physical page at
// two cache types below is inert.
const aspace = architecture.createAddressSpace() orelse {
check("created a fresh address space", false);
result();
return;
};
defer architecture.destroyAddressSpace(aspace);
const page_base = fb.base & ~@as(u64, abi.page_size - 1);
architecture.mapUserDeviceInto(aspace, process.device_arena_base, page_base, abi.page_size, true);
check(
"the framebuffer maps write-combining (PAT entry 4: PAT bit set, PCD/PWT clear)",
architecture.userLeafIsWriteCombining(aspace, process.device_arena_base) == true,
);
// Regression guard: the strong-uncacheable default is still that, so WC is a real
// choice the flag makes, not the only behaviour.
architecture.mapUserDeviceInto(aspace, process.device_arena_base + abi.page_size, page_base, abi.page_size, false);
check(
"a register window still maps strong-uncacheable",
architecture.userLeafIsWriteCombining(aspace, process.device_arena_base + abi.page_size) == false,
);
log("display: mapped {d}x{d} pitch {d} (write-combining)\n", .{ fb.width, fb.height, fb.pitch });
result();
}
fn faultInvalidOpcode() void {
log("DANOS-TEST-BEGIN: fault-ud\n", .{});
asm volatile ("ud2");
+26
View File
@@ -0,0 +1,26 @@
//! Wall-clock time: the CMOS real-time clock read once at boot and anchored to the
//! monotonic clock, so a query is a cheap arithmetic offset — no per-call CMOS poll,
//! no lock, no SMP hazard on the shared 0x70/0x71 ports.
//!
//! Wall-clock *seconds* are mechanism the kernel owns, exactly like the monotonic
//! clock ([[time-architecture]]): reading the hardware's value is not policy.
//! Calendars, timezones, and formatting layer on top in user space. It exists so the
//! filesystem can stamp real timestamps (mtime) — see docs/zig-self-hosting.md.
const architecture = @import("architecture");
var boot_unix_seconds: u64 = 0;
var boot_nanos: u64 = 0;
/// Read the RTC once and anchor it to the monotonic clock. Call at boot, after the
/// monotonic clock is calibrated.
pub fn init() void {
boot_unix_seconds = architecture.readRtcUnixSeconds();
boot_nanos = architecture.nanos();
}
/// The current wall-clock time in Unix epoch seconds (UTC): the boot RTC value plus
/// the monotonic time elapsed since. Zero until `init` runs.
pub fn nowSeconds() u64 {
return boot_unix_seconds + (architecture.nanos() -% boot_nanos) / 1_000_000_000;
}
+6 -3
View File
@@ -17,10 +17,13 @@ pub const maximum_cpus = 128;
/// Maximum tasks (kernel threads) alive at once — the static task-table size. Each
/// online core consumes one slot for its idle task, plus task 0 on the BSP. Sized
/// for the initial-ramdisk sweep (15 bundled binaries spawned at once) plus the
/// for the initial-ramdisk sweep (the bundled binaries spawned at once) plus the
/// device manager's supervised children with room to grow — at 16 the sweep
/// started failing spawns once the bundle passed a dozen binaries.
pub const maximum_tasks = 32;
/// started failing spawns once the bundle passed a dozen binaries. Raised to 48
/// for the USB stack: the xHCI bus driver spawns a supervised class-driver instance
/// per matched interface (keyboard, mouse, mass storage), on top of the FAT and
/// block servers and the growing ramdisk bundle.
pub const maximum_tasks = 48;
/// Each task's kernel stack (also each AP's bring-up stack), in bytes.
pub const kernel_stack_size = 16 * 1024;
+8 -6
View File
@@ -103,9 +103,11 @@ fn findTablesNode(buffer: []device.DeviceDescriptor) ?device.DeviceDescriptor {
}
pub fn main(init: runtime.process.Init) void {
// When the acpi-parse scenario spawns this directly, argv[1] is the kernel's
// own device count to self-verify against — deterministic, no log-scraping.
const expected: ?usize = if (init.arguments.get(1)) |a| (std.fmt.parseInt(usize, a, 10) catch null) else null;
// When the acpi-parse scenario spawns this directly, argv[1] is a device-count
// *floor* to self-verify against. The kernel no longer parses AML, so there is
// no exact count to match — proving the ring-3 parse found at least a floor of
// devices is the check. Deterministic, no log-scraping.
const floor: ?usize = if (init.arguments.get(1)) |a| (std.fmt.parseInt(usize, a, 10) catch null) else null;
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("/system/services/acpi: out of memory\n");
@@ -162,11 +164,11 @@ pub fn main(init: runtime.process.Init) void {
var namespace = result.namespace;
const devices = aml.deviceCount(&namespace);
writeLine("/system/services/acpi: parsed {d} AML blob(s), {d} namespace devices\n", .{ block_count, devices });
if (expected) |want| {
if (devices == want) {
if (floor) |minimum| {
if (devices >= minimum) {
_ = runtime.system.write("acpi-parse: ok\n");
} else {
writeLine("acpi-parse: mismatch (ring-3 {d} vs kernel {d})\n", .{ devices, want });
writeLine("acpi-parse: too few (ring-3 {d} < floor {d})\n", .{ devices, minimum });
}
// Self-verify mode is standalone (no manager); stop before reporting.
while (true) runtime.system.sleep(1000);
+44
View File
@@ -0,0 +1,44 @@
//! The block-device wire protocol — what a filesystem (the FAT server) says to a
//! block driver (usb-storage) over its well-known `.block` endpoint. A protocol
//! module like vfs-protocol / usb-transfer-protocol: extern-struct messages, an
//! `Operation` tag, everything in one IPC message.
//!
//! Data path: read and write move whole blocks to or from a **caller-owned DMA
//! buffer**, named by its physical address — the same physical-address handoff
//! usb-storage already uses toward the controller, one layer up. So a 512-byte
//! sector never has to cross the 256-byte IPC boundary; only the small request /
//! reply headers do. (Safe while the IOMMU is unenforced; see docs/driver-model.md.)
pub const Operation = enum(u32) {
/// geometry() -> { block_size, block_count }
geometry = 0,
/// read(lba, count, physical): read `count` blocks from `lba` into the buffer
read = 1,
/// write(lba, count, physical): write `count` blocks at `lba` from the buffer
write = 2,
/// flush(): commit any device write cache to stable media (no data transfer).
/// A filesystem calls this to make prior writes durable — e.g. before power-off,
/// so a shutdown-time write isn't lost in the USB flash controller's cache.
flush = 3,
};
pub const Request = extern struct {
operation: u32,
reserved: u32 = 0,
lba: u64,
count: u32, // number of blocks (read/write)
reserved2: u32 = 0,
physical: u64, // caller's DMA buffer physical address (read/write)
};
pub const Reply = extern struct {
status: i32, // 0 on success, negative on failure
reserved: u32 = 0,
block_size: u32, // geometry: bytes per block (512)
reserved2: u32 = 0,
block_count: u64, // geometry: total blocks; read/write: blocks moved
};
pub const message_maximum: usize = 256;
pub const request_size: usize = @sizeOf(Request);
pub const reply_size: usize = @sizeOf(Reply);
@@ -19,6 +19,7 @@ const std = @import("std");
const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids");
const pci_class = @import("pci-class");
const usb_ids = @import("usb-ids");
const protocol = runtime.device_manager_protocol;
const device = runtime.device;
const system = runtime.system;
@@ -61,6 +62,36 @@ fn hidDriverFor(hid: []const u8) ?[]const u8 {
return null;
}
/// The driver that serves a *reported* USB interface by its (class, subclass,
/// protocol) triple — the third bus after PCI and ACPI (docs/device-manager.md:
/// matching stays code until the third bus). The xHCI bus driver reports each
/// interface with this packed triple as its identity; the matched class driver is
/// spawned with the interface's registered id as argv[1], which it presents to the
/// bus driver to open the device.
fn usbDriverForIdentity(identity: u64) ?[]const u8 {
const keyboard = comptime usb_ids.packTriple(
@intFromEnum(usb_ids.Class.hid),
@intFromEnum(usb_ids.hid.SubClass.boot),
@intFromEnum(usb_ids.hid.Protocol.keyboard),
);
const mouse = comptime usb_ids.packTriple(
@intFromEnum(usb_ids.Class.hid),
@intFromEnum(usb_ids.hid.SubClass.boot),
@intFromEnum(usb_ids.hid.Protocol.mouse),
);
const storage = comptime usb_ids.packTriple(
@intFromEnum(usb_ids.Class.mass_storage),
@intFromEnum(usb_ids.mass_storage.SubClass.scsi),
@intFromEnum(usb_ids.mass_storage.Protocol.bulk_only),
);
return switch (identity) {
keyboard => "usb-hid-keyboard",
mouse => "usb-hid-mouse",
storage => "usb-storage",
else => null,
};
}
/// Whether some driver entry already serves registered device `device_id` —
/// a re-report after a bus restart must not spawn a second instance.
fn driverForDevice(device_id: u64) bool {
@@ -411,6 +442,11 @@ fn onChildAdded(message: []const u8, reply: []u8, sender: u32) usize {
if (pciDriverForIdentity(report.identity)) |child_driver| {
if (!driverForDevice(report.device_id)) addDriver(child_driver, report.device_id, true);
}
// USB interface match: the reported identity is the packed class triple,
// and the class driver is spawned with the interface's registered id.
if (usbDriverForIdentity(report.identity)) |usb_driver| {
if (!driverForDevice(report.device_id)) addDriver(usb_driver, report.device_id, true);
}
// ACPI _HID match (M20.3): ps2-bus is a singleton that finds its own
// devices by hid, so spawn it once, without a device assignment.
const hid_len = std.mem.indexOfScalar(u8, &report.hid, 0) orelse report.hid.len;
@@ -0,0 +1,66 @@
//! system/services/display-demo — a hardware-free client of the display service, the
//! `input-source` analog for the compositor. It creates a wallpaper, a rectangle it moves
//! each frame, and a small cursor, then drives the compositor in a present loop — proof
//! that a *separate process* can compose a moving scene through the display service over
//! IPC, exercising the layer client API and damage-driven present end to end
//! (docs/display.md). It logs `display-demo: ok` once it has driven a run of frames.
const runtime = @import("runtime");
const display = runtime.display;
const system = runtime.system;
const time = runtime.time;
pub fn main() void {
const mode = display.info() orelse {
_ = system.write("display-demo: no display service\n");
return;
};
// A full-screen wallpaper under everything.
const wallpaper = display.createLayer(0, 0, mode.width, mode.height, 0) orelse return createFailed();
_ = wallpaper.fill(0, 0, mode.width, mode.height, display.color(0x10, 0x18, 0x28));
// A rectangle that slides back and forth.
const box_w: u32 = 140;
const box_h: u32 = 100;
const box_y: i32 = 200;
const box = display.createLayer(0, box_y, box_w, box_h, 1) orelse return createFailed();
_ = box.fill(0, 0, box_w, box_h, display.color(0xE0, 0x60, 0x40));
// A little cursor on top.
const cursor = display.createLayer(40, 40, 12, 12, 2) orelse return createFailed();
_ = cursor.fill(0, 0, 12, 12, display.color(0xF0, 0xF0, 0xF0));
_ = display.present();
_ = system.write("display-demo: scene up; animating\n");
const span: i32 = @as(i32, @intCast(mode.width)) - @as(i32, @intCast(box_w));
var x: i32 = 0;
var dx: i32 = 8;
var frame: u32 = 0;
while (true) : (frame += 1) {
x += dx;
if (x <= 0) {
x = 0;
dx = -dx;
} else if (x >= span) {
x = span;
dx = -dx;
}
_ = box.configure(x, box_y, 1, true); // move it; the compositor repaints old + new
_ = display.present();
// A run of frames drawn through the compositor is the automated proof (the visible
// motion is a screenshot away via `zig build run-x86-64`).
if (frame == 20) _ = system.write("display-demo: ok\n");
time.sleep(time.Duration.fromMillis(30));
}
}
fn createFailed() void {
_ = system.write("display-demo: create failed\n");
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+192
View File
@@ -0,0 +1,192 @@
//! The compositor's pure core: rectangle math and the three blitting primitives the
//! display service composes frames from — fill a rectangle of a surface, composite one
//! surface onto another clipped to a damage rectangle, and copy a client-supplied pixel
//! tile in. Deliberately free of any syscall or `runtime` dependency (it takes plain
//! pixel pointers), so it is host-tested under `zig build test`. The service
//! (system/services/display/display.zig) wires real mmap'd surfaces and the framebuffer
//! to it. Pixels are opaque native 32-bit values — v1 layers don't alpha-blend, and
//! channel order (rgbx/bgrx) is the caller's concern (see protocol.pack).
const std = @import("std");
/// An axis-aligned rectangle in pixels. Signed, so a surface partly off-screen (a layer
/// dragged past an edge) clips with plain arithmetic. Half-open: covers [x, x+w) × [y, y+h).
pub const Rect = struct {
x: i32,
y: i32,
w: i32,
h: i32,
pub const empty = Rect{ .x = 0, .y = 0, .w = 0, .h = 0 };
pub fn init(x: i32, y: i32, w: i32, h: i32) Rect {
return .{ .x = x, .y = y, .w = w, .h = h };
}
pub fn isEmpty(r: Rect) bool {
return r.w <= 0 or r.h <= 0;
}
pub fn right(r: Rect) i32 {
return r.x + r.w;
}
pub fn bottom(r: Rect) i32 {
return r.y + r.h;
}
/// The overlap of two rectangles, or an empty rectangle if they don't touch.
pub fn intersect(a: Rect, b: Rect) Rect {
const x0 = @max(a.x, b.x);
const y0 = @max(a.y, b.y);
const x1 = @min(a.right(), b.right());
const y1 = @min(a.bottom(), b.bottom());
return .{ .x = x0, .y = y0, .w = x1 - x0, .h = y1 - y0 };
}
/// The bounding box of two rectangles. An empty operand contributes nothing (returns
/// the other), so folding damage rectangles with `unite` from `empty` yields their
/// bounding box.
pub fn unite(a: Rect, b: Rect) Rect {
if (a.isEmpty()) return b;
if (b.isEmpty()) return a;
const x0 = @min(a.x, b.x);
const y0 = @min(a.y, b.y);
const x1 = @max(a.right(), b.right());
const y1 = @max(a.bottom(), b.bottom());
return .{ .x = x0, .y = y0, .w = x1 - x0, .h = y1 - y0 };
}
};
/// A block of 32-bit pixels: `pixels` addressed row-major with `stride` pixels between
/// row starts (≥ width — the framebuffer's stride is pitch/4, a layer's is its width).
pub const Surface = struct {
pixels: [*]u32,
stride: u32, // pixels per row
width: u32,
height: u32,
pub fn bounds(s: Surface) Rect {
return .{ .x = 0, .y = 0, .w = @intCast(s.width), .h = @intCast(s.height) };
}
inline fn row(s: Surface, y: u32) [*]u32 {
return s.pixels + @as(usize, y) * s.stride;
}
};
/// Fill `rect` of `s` with the native pixel `colour`, clipped to `s`'s bounds.
pub fn fillRect(s: Surface, rect: Rect, colour: u32) void {
const c = rect.intersect(s.bounds());
if (c.isEmpty()) return;
var y: i32 = c.y;
while (y < c.bottom()) : (y += 1) {
const r = s.row(@intCast(y));
var x: i32 = c.x;
while (x < c.right()) : (x += 1) r[@intCast(x)] = colour;
}
}
/// Composite the whole of `layer` onto `dst` with the layer's top-left at (`dx`, `dy`),
/// painting only the pixels that fall inside `clip` (a `dst`-space rectangle) and inside
/// `dst`. Opaque copy. This is the primitive `present` repeats over the visible layer
/// stack, bottom to top, for each damaged region.
pub fn composite(dst: Surface, dx: i32, dy: i32, layer: Surface, clip: Rect) void {
const on_screen = Rect{ .x = dx, .y = dy, .w = @intCast(layer.width), .h = @intCast(layer.height) };
const region = on_screen.intersect(clip).intersect(dst.bounds());
if (region.isEmpty()) return;
var y: i32 = region.y;
while (y < region.bottom()) : (y += 1) {
const src = layer.row(@intCast(y - dy));
const d = dst.row(@intCast(y));
var x: i32 = region.x;
while (x < region.right()) : (x += 1) {
d[@intCast(x)] = src[@intCast(x - dx)];
}
}
}
/// Copy a `w`×`h` tile of native pixels from `src` (raw little-endian bytes, row-major,
/// tightly packed) into `dst` at (`dx`, `dy`), clipped to `dst`'s bounds. `src` is read
/// with `readInt` because it comes straight out of an IPC message buffer and carries no
/// alignment guarantee. Returns without touching anything if `src` is short.
pub fn blitTile(dst: Surface, dx: i32, dy: i32, src: []const u8, w: u32, h: u32) void {
if (src.len < @as(usize, w) * h * 4) return;
var ty: u32 = 0;
while (ty < h) : (ty += 1) {
const yy = dy + @as(i32, @intCast(ty));
if (yy < 0 or yy >= dst.height) continue;
const drow = dst.row(@intCast(yy));
var tx: u32 = 0;
while (tx < w) : (tx += 1) {
const xx = dx + @as(i32, @intCast(tx));
if (xx < 0 or xx >= dst.width) continue;
const off = (@as(usize, ty) * w + tx) * 4;
drow[@intCast(xx)] = std.mem.readInt(u32, src[off..][0..4], .little);
}
}
}
// --- tests ------------------------------------------------------------------
test "rect intersect: overlap and disjoint" {
try std.testing.expectEqual(Rect.init(5, 5, 5, 5), Rect.init(0, 0, 10, 10).intersect(Rect.init(5, 5, 10, 10)));
try std.testing.expect(Rect.init(0, 0, 10, 10).intersect(Rect.init(20, 20, 5, 5)).isEmpty());
}
test "rect unite: bounding box, empty is identity" {
const a = Rect.init(2, 2, 4, 4);
try std.testing.expectEqual(Rect.init(2, 1, 10, 5), a.unite(Rect.init(10, 1, 2, 2)));
try std.testing.expectEqual(a, a.unite(Rect.empty));
try std.testing.expectEqual(a, Rect.empty.unite(a));
}
test "fillRect clips to surface and honours stride padding" {
// A 4×3 surface inside a 6-wide allocation (stride 6 > width 4), like pitch padding.
var mem = [_]u32{0} ** (6 * 3);
const s = Surface{ .pixels = &mem, .stride = 6, .width = 4, .height = 3 };
fillRect(s, Rect.init(-1, -1, 3, 3), 0xAB); // straddles the top-left corner
try std.testing.expectEqual(@as(u32, 0xAB), mem[0 * 6 + 0]);
try std.testing.expectEqual(@as(u32, 0xAB), mem[1 * 6 + 1]);
try std.testing.expectEqual(@as(u32, 0), mem[0 * 6 + 2]); // beyond the 2-wide fill
try std.testing.expectEqual(@as(u32, 0), mem[2 * 6 + 0]); // row 2 untouched
try std.testing.expectEqual(@as(u32, 0), mem[0 * 6 + 4]); // stride padding untouched
}
test "composite: overlap shows the top layer, clipped to damage" {
var back = [_]u32{0} ** (8 * 8);
const dst = Surface{ .pixels = &back, .stride = 8, .width = 8, .height = 8 };
var lo = [_]u32{0x11} ** (4 * 4);
var hi = [_]u32{0x22} ** (4 * 4);
const low = Surface{ .pixels = &lo, .stride = 4, .width = 4, .height = 4 };
const high = Surface{ .pixels = &hi, .stride = 4, .width = 4, .height = 4 };
composite(dst, 0, 0, low, dst.bounds()); // bottom at (0,0)
composite(dst, 2, 2, high, dst.bounds()); // top overlaps at (2,2)
try std.testing.expectEqual(@as(u32, 0x11), back[0 * 8 + 0]); // bottom-only
try std.testing.expectEqual(@as(u32, 0x22), back[3 * 8 + 3]); // overlap → top wins
try std.testing.expectEqual(@as(u32, 0x22), back[5 * 8 + 5]); // top-only
try std.testing.expectEqual(@as(u32, 0), back[7 * 8 + 7]); // neither
}
test "composite honours the damage rectangle" {
var back = [_]u32{0} ** (8 * 8);
const dst = Surface{ .pixels = &back, .stride = 8, .width = 8, .height = 8 };
var fill = [_]u32{0x33} ** (8 * 8);
const layer = Surface{ .pixels = &fill, .stride = 8, .width = 8, .height = 8 };
composite(dst, 0, 0, layer, Rect.init(2, 2, 2, 2)); // only this damage region
try std.testing.expectEqual(@as(u32, 0x33), back[2 * 8 + 2]);
try std.testing.expectEqual(@as(u32, 0x33), back[3 * 8 + 3]);
try std.testing.expectEqual(@as(u32, 0), back[1 * 8 + 1]); // outside damage
try std.testing.expectEqual(@as(u32, 0), back[4 * 8 + 4]); // outside damage
}
test "blitTile copies a packed tile, clipping and reading unaligned bytes" {
var back = [_]u32{0} ** (4 * 4);
const dst = Surface{ .pixels = &back, .stride = 4, .width = 4, .height = 4 };
// A 2×2 tile in a byte buffer offset by one byte, so reads are unaligned.
var raw = [_]u8{0} ** (1 + 2 * 2 * 4);
const tile = raw[1..];
for (0..4) |i| std.mem.writeInt(u32, tile[i * 4 ..][0..4], @intCast(0xA0 + i), .little);
blitTile(dst, 3, 3, tile, 2, 2); // bottom-right corner; only (3,3) lands on-surface
try std.testing.expectEqual(@as(u32, 0xA0), back[3 * 4 + 3]);
try std.testing.expectEqual(@as(u32, 0), back[0]); // nothing else touched
}
+397
View File
@@ -0,0 +1,397 @@
//! /system/services/display — the display service (docs/display.md). A ring-3 process
//! that claims the framebuffer the kernel seeded (docs/display-plan.md D1), owns it as a
//! **write-combining front buffer**, composites an ordered stack of **layers** into a
//! **cacheable back buffer**, and presents finished frames — the GUI track's compositor,
//! the sibling of the input service. Reached by name over `ServiceId.display`.
//!
//! A layer is a server-owned surface (its own cacheable buffer) with a screen position,
//! z-order, and visibility. Clients create layers and draw into them by command
//! (`fill_rect`, `blit_tile`), mark `damage`, and ask for a `present`; the compositor
//! repaints only the damaged region — clear it, paint the visible layers bottom-to-top,
//! flush it to the screen. The pixel math lives in the pure, host-tested
//! [compositor.zig](compositor.zig); this file wires real surfaces and the framebuffer to
//! it. Shared-memory client surfaces are a later milestone (docs/display.md).
const std = @import("std");
const runtime = @import("runtime");
const compositor = @import("compositor.zig");
const protocol = runtime.display_protocol;
const ipc = runtime.ipc;
const system = runtime.system;
const device = runtime.device;
const Rect = compositor.Rect;
const Surface = compositor.Surface;
/// The claimed framebuffer and its off-screen twin. The front buffer is the LFB —
/// write-combining, so it is **only ever written**, never read; all compositing happens
/// in the cacheable back buffer, which is then streamed to the front (docs/display.md).
const Display = struct {
device_id: u64,
front: [*]volatile u8, // the LFB (write-combining)
back: [*]u8, // cacheable, same geometry
width: u32,
height: u32,
pitch: u32, // bytes per row (shared by both buffers)
format: u32, // a device-abi DisplayFormat value
frames: u64 = 0,
};
var display: Display = undefined;
/// The wallpaper the compositor clears damaged regions to before painting layers.
var background: u32 = 0;
/// The layer stack. A fixed table (a compositor has few top-level surfaces during
/// bring-up); each used slot owns an mmap'd surface. `damage` accumulates the dirty
/// screen region since the last `present`, so a present touches only what changed.
const maximum_layers = 16;
const Layer = struct {
used: bool = false,
x: i32 = 0,
y: i32 = 0,
z: u32 = 0,
visible: bool = false,
surface: Surface = undefined,
surface_len: usize = 0, // for munmap on destroy
};
var layers: [maximum_layers]Layer = [_]Layer{.{}} ** maximum_layers;
var damage: Rect = Rect.empty;
/// Enumeration buffer kept off the stack — a `DeviceDescriptor` is large, and this
/// service only ever needs one scan.
var device_table: [64]device.DeviceDescriptor = undefined;
// --- geometry helpers -------------------------------------------------------
fn screenRect() Rect {
return .{ .x = 0, .y = 0, .w = @intCast(display.width), .h = @intCast(display.height) };
}
fn backSurface() Surface {
return .{
.pixels = @ptrCast(@alignCast(display.back)),
.stride = display.pitch / 4, // pitch is bytes; a 32-bpp row is pitch/4 pixels
.width = display.width,
.height = display.height,
};
}
fn layerScreenRect(l: *const Layer) Rect {
return .{ .x = l.x, .y = l.y, .w = @intCast(l.surface.width), .h = @intCast(l.surface.height) };
}
/// Add `r` (screen coordinates) to the pending damage, clipped to the screen.
fn addDamage(r: Rect) void {
damage = damage.unite(r.intersect(screenRect()));
}
// --- layer operations (called from onMessage and the self-check) ------------
fn freeLayer() ?u32 {
for (&layers, 0..) |*l, i| {
if (!l.used) return @intCast(i);
}
return null;
}
/// A used layer by id, or null if the id is out of range or free.
fn layerAt(id: u32) ?*Layer {
if (id >= maximum_layers or !layers[id].used) return null;
return &layers[id];
}
fn createLayer(x: i32, y: i32, w: u32, h: u32, z: u32, visible: bool) ?u32 {
if (w == 0 or h == 0) return null;
const slot = freeLayer() orelse return null;
const len = @as(usize, w) * h * 4;
const base = system.mmap(len, system.PROT_READ | system.PROT_WRITE);
if (system.mmapFailed(base)) return null;
layers[slot] = .{
.used = true,
.x = x,
.y = y,
.z = z,
.visible = visible,
.surface = .{ .pixels = @ptrFromInt(base), .stride = w, .width = w, .height = h },
.surface_len = len,
};
return slot;
}
fn fillLayer(id: u32, local: Rect, colour: u32) bool {
const l = layerAt(id) orelse return false;
compositor.fillRect(l.surface, local, colour);
// Damage in screen space = the fill, translated by the layer origin, within the layer.
const screen = Rect{ .x = l.x + local.x, .y = l.y + local.y, .w = local.w, .h = local.h };
addDamage(screen.intersect(layerScreenRect(l)));
return true;
}
fn blitLayer(id: u32, x: i32, y: i32, w: u32, h: u32, pixels: []const u8) bool {
const l = layerAt(id) orelse return false;
compositor.blitTile(l.surface, x, y, pixels, w, h);
const screen = Rect{ .x = l.x + x, .y = l.y + y, .w = @intCast(w), .h = @intCast(h) };
addDamage(screen.intersect(layerScreenRect(l)));
return true;
}
fn configureLayer(id: u32, x: i32, y: i32, z: u32, visible: bool) bool {
const l = layerAt(id) orelse return false;
addDamage(layerScreenRect(l)); // the old footprint must repaint
l.x = x;
l.y = y;
l.z = z;
l.visible = visible;
addDamage(layerScreenRect(l)); // and the new one
return true;
}
fn destroyLayer(id: u32) bool {
const l = layerAt(id) orelse return false;
addDamage(layerScreenRect(l));
_ = system.munmap(@intFromPtr(l.surface.pixels), l.surface_len);
l.* = .{};
return true;
}
// --- compositing + present --------------------------------------------------
/// Repaint the damaged region `clip` of the back buffer: clear it to the background, then
/// paint every visible layer that overlaps it, bottom to top (ascending z).
fn compositeInto(clip: Rect) void {
const back = backSurface();
compositor.fillRect(back, clip, background);
// z-order the used, visible layers (n ≤ 16; a plain insertion sort of indices).
var order: [maximum_layers]u32 = undefined;
var n: usize = 0;
for (layers, 0..) |l, i| {
if (l.used and l.visible) {
order[n] = @intCast(i);
n += 1;
}
}
var a: usize = 1;
while (a < n) : (a += 1) {
const key = order[a];
var b: usize = a;
while (b > 0 and layers[order[b - 1]].z > layers[key].z) : (b -= 1) order[b] = order[b - 1];
order[b] = key;
}
for (order[0..n]) |i| {
const l = layers[i];
compositor.composite(back, l.x, l.y, l.surface, clip);
}
}
/// Stream the damaged rectangle from the cacheable back buffer to the write-combining
/// front buffer, row by row (sequential writes — what WC memory wants; we never read the
/// front buffer). Only the visible width of each row is touched.
fn flushRect(rect: Rect) void {
const c = rect.intersect(screenRect());
if (c.isEmpty()) return;
var y: i32 = c.y;
while (y < c.bottom()) : (y += 1) {
const off = @as(usize, @intCast(y)) * display.pitch;
const src: [*]const u32 = @ptrCast(@alignCast(display.back + off));
const dst: [*]volatile u32 = @ptrCast(@alignCast(display.front + off));
var x: i32 = c.x;
while (x < c.right()) : (x += 1) dst[@intCast(x)] = src[@intCast(x)];
}
}
/// Composite and flush the accumulated damage, then clear it. A no-op when nothing is
/// dirty. The frame counter advances regardless, so callers can name frames.
fn present() void {
const dirty = damage.intersect(screenRect());
if (!dirty.isEmpty()) {
compositeInto(dirty);
flushRect(dirty);
}
damage = Rect.empty;
display.frames += 1;
}
// --- startup self-check -----------------------------------------------------
/// Prove the compositor wiring on the real framebuffer: two overlapping opaque layers,
/// composited, must show the top layer in the overlap and the bottom layer outside it.
/// Exercises the whole path — mmap surfaces, the z-sort, damage, composite into the back
/// buffer — and reads the composited result back. Cleans up after itself.
fn selfCheck() void {
const red = protocol.pack(display.format, 0xC0, 0x20, 0x20);
const green = protocol.pack(display.format, 0x20, 0xC0, 0x20);
const bottom = createLayer(100, 100, 80, 80, 0, true) orelse return fail_check("create");
const top = createLayer(140, 140, 80, 80, 1, true) orelse return fail_check("create");
_ = fillLayer(bottom, Rect.init(0, 0, 80, 80), red);
_ = fillLayer(top, Rect.init(0, 0, 80, 80), green);
present();
const back = backSurface();
const overlap = back.pixels[@as(usize, 150) * back.stride + 150]; // in both layers → top
const bottom_only = back.pixels[@as(usize, 110) * back.stride + 110]; // bottom only
_ = destroyLayer(top);
_ = destroyLayer(bottom);
present(); // repaint the self-check region back to the background
if (overlap == green and bottom_only == red) {
_ = system.write("display: compositor self-check ok\n");
} else {
_ = system.write("display: compositor self-check FAILED\n");
}
}
fn fail_check(_: []const u8) void {
_ = system.write("display: compositor self-check FAILED (setup)\n");
}
// --- service ----------------------------------------------------------------
/// The framebuffer node the kernel seeded (`DeviceClass.display`), or null if none.
fn findDisplay() ?device.DeviceDescriptor {
const total = device.enumerate(&device_table);
const n = @min(total, device_table.len);
for (device_table[0..n]) |d| {
if (d.class == @intFromEnum(device.DeviceClass.display)) return d;
}
return null;
}
fn initialise(endpoint: ipc.Handle) bool {
_ = endpoint;
// Find the framebuffer, retrying while device discovery catches up with our spawn.
var tries: u32 = 0;
const found = while (tries < 100) : (tries += 1) {
if (findDisplay()) |d| break d;
system.sleep(50);
} else {
_ = system.write("display: no framebuffer device (headless?)\n");
return false; // clean exit: nothing to drive
};
if (!device.claim(found.id)) {
_ = system.write("display: could not claim the framebuffer\n");
return false;
}
// Resource 0 is the framebuffer memory window; the kernel maps it write-combining
// because the resource carries that flag (docs/display-plan.md D1).
const front_base = device.mmioMap(found.id, 0) orelse {
_ = system.write("display: could not map the framebuffer\n");
return false;
};
const geometry = found.display;
const size = @as(usize, geometry.height) * geometry.pitch;
const back_base = system.mmap(size, system.PROT_READ | system.PROT_WRITE);
if (system.mmapFailed(back_base)) {
_ = system.write("display: could not allocate the back buffer\n");
return false;
}
display = .{
.device_id = found.id,
.front = @ptrFromInt(front_base),
.back = @ptrFromInt(back_base),
.width = geometry.width,
.height = geometry.height,
.pitch = geometry.pitch,
.format = geometry.format,
};
background = protocol.pack(display.format, 0x20, 0x30, 0x48); // a dark slate wallpaper
// Clear the whole screen through the back buffer → present path (double buffering:
// no direct-to-LFB drawing).
addDamage(screenRect());
present();
var line: [96]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "display: online {d}x{d} pitch {d} format {d}\n", .{
display.width, display.height, display.pitch, display.format,
}) catch "display: online\n");
_ = system.write("display: presented frame 0\n");
selfCheck();
return true;
}
fn writeReply(reply: []u8, value: protocol.Reply) usize {
const bytes = std.mem.asBytes(&value);
@memcpy(reply[0..bytes.len], bytes);
return bytes.len;
}
fn ok(reply: []u8) usize {
return writeReply(reply, .{ .status = 0 });
}
fn fail(reply: []u8) usize {
return writeReply(reply, .{ .status = -1 });
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize {
_ = sender;
_ = capability;
if (message.len < protocol.request_size) return fail(reply);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..];
// Switch on the raw operation value — an out-of-range one must fail cleanly, not
// panic an `@enumFromInt`.
switch (request.operation) {
@intFromEnum(protocol.Operation.info) => return writeReply(reply, .{
.status = 0,
.width = display.width,
.height = display.height,
.pitch = display.pitch,
.format = display.format,
}),
@intFromEnum(protocol.Operation.create_layer) => {
// x/y are signed coordinates carried in the u32 wire fields — reinterpret the
// bits (@bitCast), don't range-check (@intCast) which a negative would fail.
const slot = createLayer(@bitCast(request.x), @bitCast(request.y), request.width, request.height, request.z, request.visible != 0) orelse return fail(reply);
return writeReply(reply, .{ .status = 0, .layer = slot });
},
@intFromEnum(protocol.Operation.configure_layer) => {
return if (configureLayer(request.layer, @bitCast(request.x), @bitCast(request.y), request.z, request.visible != 0)) ok(reply) else fail(reply);
},
@intFromEnum(protocol.Operation.destroy_layer) => {
return if (destroyLayer(request.layer)) ok(reply) else fail(reply);
},
@intFromEnum(protocol.Operation.fill_rect) => {
const local = Rect.init(@bitCast(request.x), @bitCast(request.y), @intCast(request.width), @intCast(request.height));
return if (fillLayer(request.layer, local, request.colour)) ok(reply) else fail(reply);
},
@intFromEnum(protocol.Operation.blit_tile) => {
return if (blitLayer(request.layer, @bitCast(request.x), @bitCast(request.y), request.width, request.height, payload)) ok(reply) else fail(reply);
},
@intFromEnum(protocol.Operation.damage) => {
const l = layerAt(request.layer) orelse return fail(reply);
const screen = Rect{ .x = l.x + @as(i32, @bitCast(request.x)), .y = l.y + @as(i32, @bitCast(request.y)), .w = @intCast(request.width), .h = @intCast(request.height) };
addDamage(screen.intersect(layerScreenRect(l)));
return ok(reply);
},
@intFromEnum(protocol.Operation.present) => {
present();
return ok(reply);
},
else => return fail(reply),
}
}
pub fn main() void {
runtime.service.run(protocol.message_maximum, .{
.service = .display,
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+90
View File
@@ -0,0 +1,90 @@
//! The display wire protocol — what a client says to the display service over its
//! well-known `.display` endpoint. extern-struct messages with an `Operation` tag, the
//! same shape as block/vfs/input protocols. The compositor owns the framebuffer and an
//! ordered stack of **layers**; a client creates layers, draws into them with these
//! operations, marks damage, and asks for a `present`. v1 surfaces are server-owned (a
//! client draws by command); shared-memory surfaces are a later milestone (docs/display.md).
const std = @import("std");
pub const Operation = enum(u32) {
/// info() -> { width, height, pitch, format }: the display's current mode.
info = 0,
/// create_layer(x, y, width, height, z) -> { layer }: a new server-owned surface.
create_layer = 1,
/// configure_layer(layer, x, y, z, visible): move, restack, show, or hide a layer.
configure_layer = 2,
/// destroy_layer(layer): release a layer.
destroy_layer = 3,
/// fill_rect(layer, x, y, width, height, colour): fill a rectangle of a layer.
fill_rect = 4,
/// blit_tile(layer, x, y, width, height, <inline pixels>): copy a small pixel tile in.
blit_tile = 5,
/// damage(layer, x, y, width, height): mark a region dirty for the next present.
damage = 6,
/// present(): composite the dirty layers and flush to the screen.
present = 7,
};
/// The fixed request header. A `blit_tile`'s pixel payload (width*height 32-bit pixels)
/// follows this header inline in the same message, up to `maximum_payload`.
pub const Request = extern struct {
operation: u32,
layer: u32 = 0, // create/configure/destroy/fill/blit/damage: the target layer
x: u32 = 0,
y: u32 = 0,
width: u32 = 0,
height: u32 = 0,
z: u32 = 0, // create_layer / configure_layer: stacking order (higher = in front)
colour: u32 = 0, // fill_rect: the fill colour (native pixel value)
visible: u32 = 1, // configure_layer: 0 hides the layer
reserved: u32 = 0,
};
pub const Reply = extern struct {
status: i32, // 0 on success, negative on failure
reserved: u32 = 0,
// info():
width: u32 = 0,
height: u32 = 0,
pitch: u32 = 0,
format: u32 = 0, // a device-abi DisplayFormat value (0 = rgbx, 1 = bgrx)
// create_layer():
layer: u32 = 0,
reserved2: u32 = 0,
};
/// The IPC message size — the kernel caps every message at `MESSAGE_MAXIMUM` (256 bytes,
/// system/kernel/ipc-synchronous.zig), so this matches it (a larger receive/reply buffer
/// is rejected with -E2BIG). A `blit_tile` therefore carries only a *small* tile inline —
/// `maximum_payload` bytes = up to 54 pixels, enough for a cursor or small sprite; larger
/// bitmaps are the deferred shared-memory surface path (docs/display.md).
pub const message_maximum: usize = 256;
pub const request_size: usize = @sizeOf(Request);
pub const reply_size: usize = @sizeOf(Reply);
pub const maximum_payload: usize = message_maximum - request_size;
/// Pack an 8-bit-per-channel colour into the display's native 32-bit pixel for `format`
/// (a device-abi `DisplayFormat`: 0 = rgbx, 1 = bgrx). Shared so a `colour` in a
/// `fill_rect` request means the same thing to the client that sends it and the
/// compositor that paints it. Little-endian memory, reserved byte 0: rgbx puts red in
/// the low byte, bgrx puts blue there.
pub fn pack(format: u32, r: u8, g: u8, b: u8) u32 {
const rr: u32 = r;
const gg: u32 = g;
const bb: u32 = b;
return switch (format) {
1 => bb | (gg << 8) | (rr << 16), // bgrx
else => rr | (gg << 8) | (bb << 16), // rgbx
};
}
test "pack encodes native byte order for rgbx and bgrx" {
// rgbx: red in the low byte, blue in byte 2.
try std.testing.expectEqual(@as(u32, 0x0000_00AA), pack(0, 0xAA, 0, 0));
try std.testing.expectEqual(@as(u32, 0x00AA_0000), pack(0, 0, 0, 0xAA));
// bgrx: blue in the low byte, red in byte 2.
try std.testing.expectEqual(@as(u32, 0x0000_00AA), pack(1, 0, 0, 0xAA));
try std.testing.expectEqual(@as(u32, 0x00AA_0000), pack(1, 0xAA, 0, 0));
try std.testing.expectEqual(@as(u32, 0x0000_3020), pack(0, 0x20, 0x30, 0)); // green in byte 1
}
File diff suppressed because it is too large Load Diff
+108
View File
@@ -0,0 +1,108 @@
//! system/services/fat/fat-test — a client that proves the FAT mount end to end:
//! it waits for the fat server to mount the USB volume at /mnt/usb, lists the
//! root directory through the VFS (which routes /mnt/usb to the fat backend), and
//! reads a known file off it. Shipped in the initial_ramdisk; the `fat-mount`
//! kernel test spawns it alongside init.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main(init: runtime.process.Init) void {
_ = init;
// Wait for /mnt/usb to be mounted — the fat server races us at boot (it must
// bring up the whole USB storage chain first).
var opened: ?fs.Directory = null;
var tries: u32 = 0;
while (opened == null and tries < 1400) : (tries += 1) {
opened = fs.openDirectory("/mnt/usb");
if (opened == null) runtime.system.sleep(50);
}
var dir = opened orelse {
_ = runtime.system.write("fat-test: /mnt/usb never became available\n");
return;
};
var count: u32 = 0;
var entry: fs.Entry = .{};
while (dir.next(&entry)) {
writeLine("fat-test: entry '{s}' kind={d} size={d}\n", .{ entry.name(), @intFromEnum(entry.kind), entry.size });
count += 1;
if (count > 32) break;
}
dir.close();
writeLine("fat-test: listed {d} entries\n", .{count});
// Read a known file off the boot volume through the mount (best effort): the
// kernel image is an ELF, so its first bytes are the ELF magic.
if (fs.open("/mnt/usb/system/kernel", .{})) |opened_file| {
var file = opened_file;
var magic: [4]u8 = undefined;
const n = file.read(&magic) orelse 0;
file.close();
if (n == 4 and magic[0] == 0x7F and magic[1] == 'E' and magic[2] == 'L' and magic[3] == 'F') {
_ = runtime.system.write("fat-test: read /mnt/usb/system/kernel ELF magic ok\n");
} else {
writeLine("fat-test: /mnt/usb/system/kernel read {d} bytes (not ELF magic)\n", .{n});
}
}
// Exercise directory + file mutation through the mount: mkdir, create a file
// inside it, read it back, then remove it — proof mkdir/unlink reach the engine.
if (fs.makeDirectory("/mnt/usb/TESTDIR")) {
var wrote = false;
if (fs.open("/mnt/usb/TESTDIR/HELLO.TXT", .{ .create = true, .truncate = true })) |created| {
var f = created;
wrote = (f.writeAll("mutation-ok") orelse 0) == "mutation-ok".len;
f.close();
}
// The created file carries a real modification time (stamped from the RTC).
var mtime_ok = false;
if (fs.attributes("/mnt/usb/TESTDIR/HELLO.TXT")) |attrs| {
writeLine("fat-test: mtime {d}\n", .{attrs.mtime});
mtime_ok = attrs.mtime > 1_577_836_800; // after 2020-01-01
}
if (mtime_ok) _ = runtime.system.write("fat-test: mtime ok\n");
// Rename it, then read from the new name and confirm the old name is gone.
const renamed = fs.rename("/mnt/usb/TESTDIR/HELLO.TXT", "/mnt/usb/TESTDIR/RENAMED.TXT");
const old_gone = !fs.exists("/mnt/usb/TESTDIR/HELLO.TXT");
if (renamed and old_gone) _ = runtime.system.write("fat-test: rename ok\n");
var readback = false;
if (fs.open("/mnt/usb/TESTDIR/RENAMED.TXT", .{})) |reopened| {
var f = reopened;
var buf: [16]u8 = undefined;
const got = f.read(&buf) orelse 0;
f.close();
readback = std.mem.eql(u8, buf[0..got], "mutation-ok");
}
const removed = fs.remove("/mnt/usb/TESTDIR/RENAMED.TXT");
const gone = !fs.exists("/mnt/usb/TESTDIR/RENAMED.TXT");
if (wrote and mtime_ok and renamed and old_gone and readback and removed and gone) {
_ = runtime.system.write("fat-test: mutations ok\n");
} else {
writeLine("fat-test: mutations FAILED (wrote={} mtime={} renamed={} oldgone={} read={} removed={} gone={})\n", .{ wrote, mtime_ok, renamed, old_gone, readback, removed, gone });
}
} else {
_ = runtime.system.write("fat-test: mkdir /mnt/usb/TESTDIR failed\n");
}
if (count > 0) {
while (true) {
_ = runtime.system.write("fat-test: ok\n");
runtime.system.sleep(1000);
}
}
_ = runtime.system.write("fat-test: root listing was empty\n");
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+249
View File
@@ -0,0 +1,249 @@
//! system/services/fat — the FAT filesystem server. Spawned as a boot service, it
//! opens the block device (a USB stick via usb-storage) under `.block`, mounts the
//! FAT filesystem on it (the pure engine in engine.zig), and mounts itself into
//! the VFS at /mnt/usb. From then on the VFS forwards every open/read/write/
//! status/readdir/close under /mnt/usb to this server, which serves the same
//! vfs-protocol as a backend — turning block reads into file reads.
//!
//! The block data path never crosses IPC: a DMA bounce buffer is handed to the
//! block driver by physical address, and the engine copies sectors in and out of
//! it.
const std = @import("std");
const runtime = @import("runtime");
const engine = @import("engine.zig");
const on_disk = @import("on-disk.zig");
const protocol = runtime.vfs_protocol;
const dma = runtime.dma;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
const mount_point = "/mnt/usb";
// The engine's BlockDevice, backed by the `.block` driver plus a DMA bounce
// buffer the driver reads/writes by physical address.
const IpcBlock = struct {
device: runtime.block.Device,
bounce: dma.Region,
fn readBlock(context: *anyopaque, lba: u64, buffer: []u8) bool {
const self: *IpcBlock = @ptrCast(@alignCast(context));
if (!self.device.read(lba, 1, self.bounce.physical)) return false;
const source: [*]const u8 = @ptrFromInt(self.bounce.virtual);
@memcpy(buffer[0..512], source[0..512]);
return true;
}
fn writeBlock(context: *anyopaque, lba: u64, buffer: []const u8) bool {
const self: *IpcBlock = @ptrCast(@alignCast(context));
const destination: [*]u8 = @ptrFromInt(self.bounce.virtual);
@memcpy(destination[0..512], buffer[0..512]);
if (!self.device.write(lba, 1, self.bounce.physical)) return false;
device_dirty = true; // a block reached the device; a close will flush it
return true;
}
};
var ipc_block: IpcBlock = undefined;
// Set whenever a block is written, cleared when the device cache is flushed on a
// file close — so writes are committed to stable media before a power-off.
var device_dirty: bool = false;
var filesystem: engine.FileSystem = undefined;
// Open handles the VFS holds against this backend: each maps a node id to a
// resolved engine node.
const OpenNode = struct { used: bool = false, node: engine.Node = undefined, owner: u32 = 0 };
var open_nodes = [_]OpenNode{.{}} ** 32;
fn allocOpen() ?usize {
for (&open_nodes, 0..) |*o, i| {
if (!o.used) return i;
}
return null;
}
fn openAt(id: u64) ?*OpenNode {
if (id >= open_nodes.len) return null;
const o = &open_nodes[@intCast(id)];
return if (o.used) o else null;
}
fn writeReply(out: []u8, reply: protocol.Reply, payload: []const u8) usize {
@memcpy(out[0..protocol.reply_size], std.mem.asBytes(&reply));
const n = @min(payload.len, out.len - protocol.reply_size);
@memcpy(out[protocol.reply_size..][0..n], payload[0..n]);
return protocol.reply_size + n;
}
fn fail(out: []u8) usize {
return writeReply(out, .{ .status = -1 }, &.{});
}
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = runtime.system.write("/system/services/fat: starting, waiting for a block device\n");
const device = runtime.block.open() orelse {
_ = runtime.system.write("/system/services/fat: no block device (no storage attached)\n");
return false; // clean exit: nothing to serve
};
const geometry = device.geometry() orelse {
_ = runtime.system.write("/system/services/fat: block geometry unavailable\n");
return false;
};
ipc_block = .{ .device = device, .bounce = dma.alloc(4096, dma.coherent) orelse return false };
const block_device = engine.BlockDevice{
.context = &ipc_block,
.block_size = geometry.block_size,
.block_count = geometry.block_count,
.readBlockFn = IpcBlock.readBlock,
.writeBlockFn = IpcBlock.writeBlock,
};
filesystem = engine.FileSystem.mount(block_device) orelse {
_ = runtime.system.write("/system/services/fat: not a FAT filesystem\n");
return false;
};
writeLine("/system/services/fat: mounted FAT ({s}, {d} clusters, partition lba {d})\n", .{ @tagName(filesystem.geometry.fat_type), filesystem.geometry.cluster_count, filesystem.base_lba });
// Mount ourselves into the VFS namespace at /mnt/usb (retry while the VFS
// comes up). From here the VFS routes /mnt/usb/... to this server.
var tries: u32 = 0;
while (tries < 100) : (tries += 1) {
if (runtime.fs.mount(mount_point, endpoint)) {
writeLine("/system/services/fat: mounted {s}\n", .{mount_point});
return true;
}
runtime.system.sleep(50);
}
_ = runtime.system.write("/system/services/fat: could not mount into the VFS\n");
return true; // still serve directly, even if the namespace mount didn't take
}
const ParentLeaf = struct { parent: []const u8, leaf: []const u8 };
// Split a path into its parent directory and final component: "/a/b" -> ("/a",
// "b"); "/b" -> ("/", "b"); "b" -> ("/", "b").
fn splitParent(path: []const u8) ParentLeaf {
const slash = std.mem.lastIndexOfScalar(u8, path, '/');
return .{
.parent = if (slash) |s| (if (s == 0) "/" else path[0..s]) else "/",
.leaf = if (slash) |s| path[s + 1 ..] else path,
};
}
fn handleOpen(out: []u8, path: []const u8, flags: u32) usize {
var node = filesystem.resolve(path);
if (node == null and flags & protocol.create != 0) {
const split = splitParent(path);
const parent = filesystem.resolve(split.parent) orelse return fail(out);
node = filesystem.createFile(parent, split.leaf);
}
var resolved = node orelse return fail(out);
// O_TRUNC: replace an existing file's contents rather than overwriting in place
// (frees the old chain, so a shorter rewrite leaves no stale tail).
if (flags & protocol.truncate != 0 and !resolved.is_directory) {
filesystem.truncate(&resolved);
}
const index = allocOpen() orelse return fail(out);
open_nodes[index] = .{ .used = true, .node = resolved };
return writeReply(out, .{ .status = 0, .node = index }, &.{});
}
fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = capability;
_ = sender;
if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..];
// Stamp create/write with the current wall-clock time (mtime). Cheap, and it
// keeps the engine pure (it takes the time as data, not a syscall).
filesystem.current_time_epoch = runtime.system.wallClock();
switch (request.operation) {
.open => return handleOpen(out, payload[0..@min(payload.len, request.len)], request.flags),
.read => {
const o = openAt(request.node) orelse return fail(out);
var buffer: [protocol.maximum_payload]u8 = undefined;
const want = @min(@as(usize, request.len), buffer.len);
const n = filesystem.readFile(o.node, @intCast(request.offset), buffer[0..want]);
return writeReply(out, .{ .status = 0, .len = @intCast(n) }, buffer[0..n]);
},
.write => {
const o = openAt(request.node) orelse return fail(out);
const data = payload[0..@min(payload.len, request.len)];
const n = filesystem.writeFile(&o.node, @intCast(request.offset), data);
return writeReply(out, .{ .status = 0, .len = @intCast(n) }, &.{});
},
.status => {
const o = openAt(request.node) orelse return fail(out);
const kind: protocol.NodeKind = if (o.node.is_directory) .directory else .regular;
const status = protocol.FileStatus{ .size = o.node.size, .kind = @intFromEnum(kind), .mtime = o.node.mtime };
return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&status));
},
.readdir => {
const o = openAt(request.node) orelse return fail(out);
if (!o.node.is_directory) return writeReply(out, .{ .status = 0, .len = 0 }, &.{});
const listing = filesystem.listEntry(o.node, @intCast(request.offset)) orelse return writeReply(out, .{ .status = 0, .len = 0 }, &.{});
const kind: protocol.NodeKind = if (listing.is_directory) .directory else .regular;
const header = protocol.DirectoryEntry{ .kind = @intFromEnum(kind), .name_len = @intCast(listing.name_len), .size = listing.size };
var buffer: [protocol.maximum_payload]u8 = undefined;
@memcpy(buffer[0..protocol.directory_entry_size], std.mem.asBytes(&header));
const nlen = @min(listing.name_len, buffer.len - protocol.directory_entry_size);
@memcpy(buffer[protocol.directory_entry_size..][0..nlen], listing.name_buffer[0..nlen]);
const total = protocol.directory_entry_size + nlen;
return writeReply(out, .{ .status = 0, .len = @intCast(total) }, buffer[0..total]);
},
.close => {
if (openAt(request.node)) |o| o.used = false;
// Durable-on-close: if any block reached the device since the last
// flush, commit its cache to stable media now (best-effort). This is
// what makes init's shutdown log flush survive a real power-off, and is
// the right default for removable media the user may unplug.
if (device_dirty) {
_ = ipc_block.device.flush();
device_dirty = false;
}
return writeReply(out, .{ .status = 0 }, &.{});
},
.mkdir => {
const split = splitParent(payload[0..@min(payload.len, request.len)]);
const parent = filesystem.resolve(split.parent) orelse return fail(out);
if (filesystem.createDirectory(parent, split.leaf) == null) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
.unlink => {
const split = splitParent(payload[0..@min(payload.len, request.len)]);
const parent = filesystem.resolve(split.parent) orelse return fail(out);
if (!filesystem.removeFile(parent, split.leaf)) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
.rename => {
const both = payload[0..@min(payload.len, request.len)];
const sep = std.mem.indexOfScalar(u8, both, 0) orelse return fail(out);
const old_split = splitParent(both[0..sep]);
const new_split = splitParent(both[sep + 1 ..]);
// Same-directory rename only.
if (!std.mem.eql(u8, old_split.parent, new_split.parent)) return fail(out);
const parent = filesystem.resolve(old_split.parent) orelse return fail(out);
if (!filesystem.rename(parent, old_split.leaf, new_split.leaf)) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
// A backend is never itself a mount target.
.mount, .unmount => return fail(out),
}
}
pub fn main() void {
runtime.service.run(protocol.message_maximum, .{
.service = .fat,
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+310
View File
@@ -0,0 +1,310 @@
//! The on-disk layout of a FAT filesystem — the boot sector / BIOS Parameter
//! Block, directory entries, long-file-name entries, and the FAT32 FSInfo — as
//! `align(1)` extern structs that bit-cast straight out of a 512-byte sector
//! (multi-byte fields are little-endian, like usb-abi.zig). Pure data, plus the
//! cluster-count FAT-type detection. Host-testable.
const std = @import("std");
/// The BIOS Parameter Block, common to FAT12/16/32 (offset 0..36 of the boot
/// sector). The extended part that follows differs by FAT type.
pub const BiosParameterBlock = extern struct {
jump: [3]u8,
oem_name: [8]u8,
bytes_per_sector: u16 align(1),
sectors_per_cluster: u8,
reserved_sector_count: u16 align(1),
fat_count: u8,
root_entry_count: u16 align(1),
total_sectors_16: u16 align(1),
media: u8,
fat_size_16: u16 align(1),
sectors_per_track: u16 align(1),
head_count: u16 align(1),
hidden_sectors: u32 align(1),
total_sectors_32: u32 align(1),
};
/// The FAT12/16 extended boot record (offset 36).
pub const ExtendedBootRecord16 = extern struct {
drive_number: u8,
reserved: u8,
boot_signature: u8,
volume_id: u32 align(1),
volume_label: [11]u8,
filesystem_type: [8]u8,
};
/// The FAT32 extended boot record (offset 36).
pub const ExtendedBootRecord32 = extern struct {
fat_size_32: u32 align(1),
extended_flags: u16 align(1),
filesystem_version: u16 align(1),
root_cluster: u32 align(1),
filesystem_information_sector: u16 align(1),
backup_boot_sector: u16 align(1),
reserved: [12]u8,
drive_number: u8,
reserved1: u8,
boot_signature: u8,
volume_id: u32 align(1),
volume_label: [11]u8,
filesystem_type: [8]u8,
};
/// A 32-byte directory entry (8.3 short name form).
pub const DirectoryEntry = extern struct {
name: [11]u8, // 8 name + 3 extension, space-padded
attributes: u8,
reserved_nt: u8,
creation_time_tenth: u8,
creation_time: u16 align(1),
creation_date: u16 align(1),
last_access_date: u16 align(1),
first_cluster_high: u16 align(1),
write_time: u16 align(1),
write_date: u16 align(1),
first_cluster_low: u16 align(1),
file_size: u32 align(1),
pub fn firstCluster(self: DirectoryEntry) u32 {
return (@as(u32, self.first_cluster_high) << 16) | self.first_cluster_low;
}
pub fn setFirstCluster(self: *DirectoryEntry, cluster: u32) void {
self.first_cluster_low = @truncate(cluster);
self.first_cluster_high = @truncate(cluster >> 16);
}
pub fn isFree(self: DirectoryEntry) bool {
return self.name[0] == 0x00 or self.name[0] == 0xE5;
}
pub fn isEnd(self: DirectoryEntry) bool {
return self.name[0] == 0x00;
}
pub fn isDirectory(self: DirectoryEntry) bool {
return self.attributes & attribute_directory != 0;
}
pub fn isLongName(self: DirectoryEntry) bool {
return self.attributes & attribute_long_name_mask == attribute_long_name;
}
pub fn isVolumeLabel(self: DirectoryEntry) bool {
return self.attributes & attribute_volume_id != 0 and !self.isLongName();
}
};
/// A 32-byte long-file-name entry (attributes == 0x0F). A sequence of these
/// precedes the 8.3 entry they name, each carrying 13 UTF-16 code units.
pub const LongNameEntry = extern struct {
order: u8,
name1: [5]u16 align(1),
attributes: u8,
kind: u8,
checksum: u8,
name2: [6]u16 align(1),
first_cluster_low: u16 align(1),
name3: [2]u16 align(1),
};
/// The FAT32 FSInfo sector (usually sector 1): advisory free-cluster bookkeeping.
pub const FileSystemInformation = extern struct {
lead_signature: u32 align(1), // 0x41615252
reserved1: [480]u8,
struct_signature: u32 align(1), // 0x61417272
free_count: u32 align(1),
next_free: u32 align(1),
reserved2: [12]u8,
trail_signature: u32 align(1), // 0xAA550000
};
// Directory-entry attribute bits.
pub const attribute_read_only: u8 = 0x01;
pub const attribute_hidden: u8 = 0x02;
pub const attribute_system: u8 = 0x04;
pub const attribute_volume_id: u8 = 0x08;
pub const attribute_directory: u8 = 0x10;
pub const attribute_archive: u8 = 0x20;
pub const attribute_long_name: u8 = 0x0F; // read_only|hidden|system|volume_id
pub const attribute_long_name_mask: u8 = 0x3F;
// FSInfo signatures.
pub const fsinfo_lead_signature: u32 = 0x41615252;
pub const fsinfo_struct_signature: u32 = 0x61417272;
pub const fsinfo_trail_signature: u32 = 0xAA550000;
/// End-of-chain markers (a cluster value >= these ends a chain).
pub const end_of_chain_12: u32 = 0xFF8;
pub const end_of_chain_16: u32 = 0xFFF8;
pub const end_of_chain_32: u32 = 0x0FFFFFF8;
pub const bad_cluster_32: u32 = 0x0FFFFFF7;
pub const free_cluster: u32 = 0;
pub const boot_signature_offset: usize = 510; // 0x55 0xAA at the end of the boot sector
pub const FatType = enum { fat12, fat16, fat32 };
/// The geometry derived from the BPB, plus the FAT type (by the Microsoft
/// cluster-count rule: <4085 FAT12, <65525 FAT16, else FAT32).
pub const Geometry = struct {
fat_type: FatType,
bytes_per_sector: u32,
sectors_per_cluster: u32,
reserved_sector_count: u32,
fat_count: u32,
fat_size_sectors: u32, // per FAT
root_entry_count: u32, // FAT12/16
root_cluster: u32, // FAT32
first_data_sector: u32,
total_sectors: u32,
cluster_count: u32,
fsinfo_sector: u32, // FAT32
};
/// Derive the geometry (and FAT type) from a boot sector's first 512 bytes.
/// Returns null if the sector is not a plausible FAT boot sector.
pub fn geometryOf(sector: []const u8) ?Geometry {
if (sector.len < 512) return null;
if (sector[boot_signature_offset] != 0x55 or sector[boot_signature_offset + 1] != 0xAA) return null;
const bpb = std.mem.bytesToValue(BiosParameterBlock, sector[0..@sizeOf(BiosParameterBlock)]);
if (bpb.bytes_per_sector == 0 or bpb.sectors_per_cluster == 0 or bpb.fat_count == 0) return null;
const fat_size_16: u32 = bpb.fat_size_16;
var fat_size: u32 = fat_size_16;
var root_cluster: u32 = 0;
var fsinfo_sector: u32 = 0;
if (fat_size_16 == 0) {
const ebr = std.mem.bytesToValue(ExtendedBootRecord32, sector[36 .. 36 + @sizeOf(ExtendedBootRecord32)]);
fat_size = ebr.fat_size_32;
root_cluster = ebr.root_cluster;
fsinfo_sector = ebr.filesystem_information_sector;
}
const total_sectors: u32 = if (bpb.total_sectors_16 != 0) bpb.total_sectors_16 else bpb.total_sectors_32;
const root_dir_sectors = (@as(u32, bpb.root_entry_count) * 32 + bpb.bytes_per_sector - 1) / bpb.bytes_per_sector;
const first_data_sector = bpb.reserved_sector_count + bpb.fat_count * fat_size + root_dir_sectors;
if (total_sectors < first_data_sector) return null;
const data_sectors = total_sectors - first_data_sector;
const cluster_count = data_sectors / bpb.sectors_per_cluster;
const fat_type: FatType = if (cluster_count < 4085) .fat12 else if (cluster_count < 65525) .fat16 else .fat32;
return .{
.fat_type = fat_type,
.bytes_per_sector = bpb.bytes_per_sector,
.sectors_per_cluster = bpb.sectors_per_cluster,
.reserved_sector_count = bpb.reserved_sector_count,
.fat_count = bpb.fat_count,
.fat_size_sectors = fat_size,
.root_entry_count = bpb.root_entry_count,
.root_cluster = root_cluster,
.first_data_sector = first_data_sector,
.total_sectors = total_sectors,
.cluster_count = cluster_count,
.fsinfo_sector = fsinfo_sector,
};
}
// --- DOS date/time <-> Unix epoch --------------------------------------------
//
// FAT stamps a file's modification time as two 16-bit DOS fields. There is no
// timezone, so danos treats them as UTC. `date`: year-1980(7)|month(4)|day(5);
// `time`: hour(5)|minute(6)|(second/2)(5).
fn isLeapYear(year: u32) bool {
return (year % 4 == 0 and year % 100 != 0) or (year % 400 == 0);
}
const days_in_month = [_]u8{ 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31 };
/// Convert a FAT date+time to Unix epoch seconds (UTC). Returns 0 for an unset
/// (zero) date.
pub fn fatToEpoch(date: u16, time: u16) u64 {
if (date == 0) return 0;
const day: u32 = date & 0x1F;
const month: u32 = (date >> 5) & 0x0F;
const year: u32 = 1980 + (date >> 9);
if (month < 1 or month > 12 or day < 1) return 0;
const second: u32 = @as(u32, time & 0x1F) * 2;
const minute: u32 = (time >> 5) & 0x3F;
const hour: u32 = (time >> 11) & 0x1F;
var days: u64 = 0;
var y: u32 = 1970;
while (y < year) : (y += 1) days += if (isLeapYear(y)) 366 else 365;
var m: u32 = 1;
while (m < month) : (m += 1) {
days += days_in_month[m - 1];
if (m == 2 and isLeapYear(year)) days += 1;
}
days += day - 1;
return ((days * 24 + hour) * 60 + minute) * 60 + second;
}
pub const FatDateTime = struct { date: u16, time: u16 };
/// Convert Unix epoch seconds (UTC) to a FAT date+time. Returns {0,0} for epoch 0 or
/// any time before 1980 (which DOS cannot represent).
pub fn epochToFatDateTime(epoch: u64) FatDateTime {
if (epoch == 0) return .{ .date = 0, .time = 0 };
var remaining = epoch;
const second: u32 = @intCast(remaining % 60);
remaining /= 60;
const minute: u32 = @intCast(remaining % 60);
remaining /= 60;
const hour: u32 = @intCast(remaining % 24);
remaining /= 24;
var days: u32 = @intCast(remaining); // whole days since 1970-01-01
var year: u32 = 1970;
while (true) {
const y_days: u32 = if (isLeapYear(year)) 366 else 365;
if (days < y_days) break;
days -= y_days;
year += 1;
}
if (year < 1980) return .{ .date = 0, .time = 0 };
var month: u32 = 1;
while (true) {
var m_days: u32 = days_in_month[month - 1];
if (month == 2 and isLeapYear(year)) m_days += 1;
if (days < m_days) break;
days -= m_days;
month += 1;
}
const day = days + 1;
return .{
.date = @intCast(((year - 1980) << 9) | (month << 5) | day),
.time = @intCast((hour << 11) | (minute << 5) | (second / 2)),
};
}
test "FAT date/time <-> Unix epoch round trip" {
// Even-second UTC times (FAT stores seconds/2, so even seconds round-trip exactly).
for ([_]u64{ 1_577_836_800, 1_700_000_000, 1_262_304_000, 1_783_971_244 }) |epoch| {
const fat = epochToFatDateTime(epoch);
try std.testing.expectEqual(epoch, fatToEpoch(fat.date, fat.time));
}
// Absolute check: 1577836800 is 2020-01-01 00:00:00 UTC.
const y2020 = epochToFatDateTime(1_577_836_800);
try std.testing.expectEqual(@as(u16, 2020), 1980 + (y2020.date >> 9));
try std.testing.expectEqual(@as(u16, 1), (y2020.date >> 5) & 0x0F); // month
try std.testing.expectEqual(@as(u16, 1), y2020.date & 0x1F); // day
// 0 is "unset" both ways.
try std.testing.expectEqual(@as(u64, 0), fatToEpoch(0, 0));
try std.testing.expectEqual(@as(u16, 0), epochToFatDateTime(0).date);
}
test "on-disk struct sizes match the specification" {
try std.testing.expectEqual(@as(usize, 36), @sizeOf(BiosParameterBlock));
try std.testing.expectEqual(@as(usize, 26), @sizeOf(ExtendedBootRecord16));
try std.testing.expectEqual(@as(usize, 54), @sizeOf(ExtendedBootRecord32));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(DirectoryEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(LongNameEntry));
try std.testing.expectEqual(@as(usize, 512), @sizeOf(FileSystemInformation));
}
test "directory entry cluster split/join" {
var entry = std.mem.zeroes(DirectoryEntry);
entry.setFirstCluster(0x01234567);
try std.testing.expectEqual(@as(u16, 0x4567), entry.first_cluster_low);
try std.testing.expectEqual(@as(u16, 0x0123), entry.first_cluster_high);
try std.testing.expectEqual(@as(u32, 0x01234567), entry.firstCluster());
}
+42 -4
View File
@@ -21,11 +21,16 @@ const std = @import("std");
const runtime = @import("runtime");
const power = runtime.power_protocol;
/// Where the kernel boot log is persisted on the USB FAT volume — an 8.3 name at
/// the mount root (see system/services/log-flush). init writes it at shutdown;
/// the log-flush one-shot writes it once at boot.
const log_path = "/mnt/usb/DANOS.LOG";
/// The system services init brings up at boot, in order. This is init's policy — the
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent
/// on purpose: the device manager owns those. (A future init reads this from a
/// manifest under /system/services instead of a hardcoded list.)
const boot_services = [_][]const u8{ "vfs", "input", "device-manager" };
const boot_services = [_][]const u8{ "vfs", "input", "device-manager", "fat", "display" };
var children: [boot_services.len]u32 = .{0} ** boot_services.len;
var child_count: usize = 0;
@@ -65,6 +70,14 @@ pub fn main() void {
}
}
// Once the storage stack is up, a one-shot copies the boot log to the USB
// volume (/mnt/usb/DANOS.LOG) so it can be read on another machine — the only
// way to see it on a headless/real board with no host capturing serial. Fire
// and forget: it polls for the mount itself, and is deliberately NOT one of
// init's supervised children (a transient one-shot must not be stopped-and-
// waited-for during shutdown).
_ = runtime.system.spawn("log-flush");
// Subscribe to power events (retry: the power service registers well after
// init starts). Best-effort — without it, a `terminate` signal still
// triggers the same shutdown path.
@@ -115,11 +128,36 @@ fn subscribePower() void {
_ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {};
}
/// The stop sequence: terminate each child in reverse spawn order (vfs last —
/// other services may flush through it), waiting up to a deadline for each to
/// exit before killing it, then ask the power service to enter S5.
/// Copy the whole kernel log to /mnt/usb/DANOS.LOG (the same file log-flush
/// writes at boot), so a poweroff captures the fullest log. Best-effort: if the
/// USB volume is not mounted, the open fails and it does nothing. Must run while
/// the storage services are still alive (see shutDown).
fn flushKernelLog() void {
// Truncate on open so this fuller flush replaces the boot-time one cleanly.
var file = runtime.fs.open(log_path, .{ .create = true, .truncate = true }) orelse return; // no USB volume
defer file.close();
var chunk: [4096]u8 = undefined;
var offset: usize = 0;
while (true) {
const got = runtime.system.klogRead(offset, &chunk);
if (got == 0) break; // reached the end of the accumulated log
if (file.writeAll(chunk[0..got]) == null) break; // storage went away
offset += got;
}
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, "/system/services/init: flushed log to {s} ({d} bytes)\n", .{ log_path, offset }) catch "");
}
/// The stop sequence: persist the log while storage is still up, then terminate
/// each child in reverse spawn order (vfs last — other services may flush through
/// it), waiting up to a deadline for each to exit before killing it, then ask the
/// power service to enter S5.
fn shutDown() void {
_ = runtime.system.write("/system/services/init: shutting down\n");
// Persist the fullest log to the USB volume BEFORE tearing anything down: the
// reverse-order stop loop below kills the fat server (children[3]) first, so
// /mnt/usb must be written while it is still mounted.
flushKernelLog();
var i = child_count;
while (i > 0) {
i -= 1;
+67
View File
@@ -0,0 +1,67 @@
//! system/services/log-flush — a one-shot that copies the kernel's in-memory
//! diagnostic log to a file on the mounted USB FAT volume, so the boot log
//! survives to be read on another machine. On a headless or real board there is
//! no host capturing serial, so without this the log is lost at power-off; this
//! is the on-disk equivalent of QEMU's `-serial file:`.
//!
//! It reads the whole kernel log back through `klog_read` (the RAM sink in
//! system/kernel/log.zig) and writes it to /mnt/usb/DANOS.LOG. The name is 8.3
//! (FAT short-name rule: base <= 8, extension <= 3) and lives at the mount root
//! (there is no mkdir on the FAT path yet). init spawns this once the boot
//! services are up; init itself repeats the flush at shutdown for a fuller log.
//!
//! If no USB volume is mounted — no stick, or the initial-ramdisk sweep that
//! spawns every bundled binary bare with no VFS — it waits briefly, then exits
//! silently, deranging no other test's output.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
const log_path = "/mnt/usb/DANOS.LOG";
/// Copy the whole kernel log to the open file, looping klog_read -> write until
/// the log is exhausted. Returns the number of bytes written.
fn drainKernelLog(file: *fs.File) usize {
var chunk: [4096]u8 = undefined;
var offset: usize = 0;
while (true) {
const got = runtime.system.klogRead(offset, &chunk);
if (got == 0) break; // reached the end of the accumulated log
if (file.writeAll(chunk[0..got]) == null) break; // storage went away
offset += got;
}
return offset;
}
pub fn main() void {
// Wait for the fat server to mount /mnt/usb (it must bring up the whole USB
// storage chain first, so it races us at boot). Bounded: if the mount never
// appears — no volume, or the no-VFS ramdisk sweep — give up silently.
var ready = false;
var tries: u32 = 0;
while (tries < 1400) : (tries += 1) {
if (fs.openDirectory("/mnt/usb")) |directory| {
var dir = directory;
dir.close();
ready = true;
break;
}
runtime.system.sleep(50);
}
if (!ready) return; // /mnt/usb never became available — nothing to persist to
// Truncate on open: each flush replaces the file, so a shorter log on a later
// boot of the same stick leaves no stale tail from a previous, longer one.
var file = fs.open(log_path, .{ .create = true, .truncate = true }) orelse return;
const written = drainKernelLog(&file);
file.close();
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, "log-flush: wrote {d} bytes to {s}\n", .{ written, log_path }) catch return);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+39
View File
@@ -0,0 +1,39 @@
//! Pure path utilities for the VFS mount router — no IPC, no state, so they are
//! host-testable in isolation. The router uses these to decide whether an opened
//! path lies under a mount point and, if so, what it looks like relative to that
//! mount.
const std = @import("std");
/// If `path` lies under `mount_prefix` — equal to it, or the prefix followed by a
/// path separator — return the path relative to the mount ("/" for an exact
/// match, otherwise the tail beginning with '/'). Returns null when `path` is not
/// under the mount, so a prefix like "/mnt/usb" never captures "/mnt/usbextra".
pub fn underMount(path: []const u8, mount_prefix: []const u8) ?[]const u8 {
if (path.len < mount_prefix.len) return null;
if (!std.mem.eql(u8, path[0..mount_prefix.len], mount_prefix)) return null;
if (path.len == mount_prefix.len) return "/";
if (path[mount_prefix.len] != '/') return null;
return path[mount_prefix.len..];
}
/// Whether `path` is absolute (rooted at '/'). Bare names — what the flat ramfs
/// uses — are relative and never route through a mount.
pub fn isAbsolute(path: []const u8) bool {
return path.len > 0 and path[0] == '/';
}
test "underMount matches only at path boundaries" {
try std.testing.expectEqualStrings("/", underMount("/mnt/usb", "/mnt/usb").?);
try std.testing.expectEqualStrings("/system/kernel", underMount("/mnt/usb/system/kernel", "/mnt/usb").?);
try std.testing.expect(underMount("/mnt/usbextra", "/mnt/usb") == null); // not a boundary
try std.testing.expect(underMount("/mnt", "/mnt/usb") == null); // shorter than the prefix
try std.testing.expect(underMount("/other", "/mnt/usb") == null);
try std.testing.expect(underMount("greeting", "/mnt/usb") == null); // a bare name
}
test "isAbsolute distinguishes paths from bare names" {
try std.testing.expect(isAbsolute("/mnt/usb"));
try std.testing.expect(!isAbsolute("greeting"));
try std.testing.expect(!isAbsolute(""));
}
+60 -5
View File
@@ -4,12 +4,11 @@
//! header followed by an inline payload (read bytes, or a FileStatus). Everything fits
//! in one IPC message (<= ipc MESSAGE_MAXIMUM = 256 bytes).
//!
//! This is a danos-native contract, so it uses danos names throughout — the POSIX
//! spellings (`stat`, `O_CREAT`, ...) live only in the POSIX layer
//! (library/posix/unistd.zig), which translates to these.
//! This is a danos-native contract, so it uses danos names throughout. The client
//! side is `runtime.fs` (library/runtime/fs.zig), which programs use directly.
//!
//! This is user-space only — the kernel knows nothing of files or paths; it only moves the bytes.
//! Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server).
//! Shared by library/runtime/fs.zig (client) and system/services/vfs/vfs.zig (server).
pub const Operation = enum(u32) {
open, // open(path) -> node id
@@ -17,8 +16,43 @@ pub const Operation = enum(u32) {
read, // read(node, offset, len) -> bytes
write, // write(node, offset, bytes) -> count
status, // status(node) -> FileStatus
// Appended for the mount router (M5). Values stay stable, so existing clients
// and the flat-ramfs tests are unaffected.
readdir, // readdir(dir_node, cursor=offset) -> one DirectoryEntry (len==0 => EOF)
mount, // mount(prefix payload, capability = backend endpoint)
unmount, // unmount(prefix payload)
// Appended for filesystem mutation (Phase 2). Path-based (the path is the
// payload); a mounted backend handles them, the flat ramfs refuses them.
mkdir, // mkdir(path payload) -> status
unlink, // unlink(path payload) -> status
// rename: the payload is the old path, a single 0x00 separator, then the new
// path. Same-directory rename only (the router requires both under one mount).
rename, // rename(old\0new payload) -> status
};
/// The type of a filesystem node, aligned to the FSH file-type table
/// (docs/danos-file-system-hierarchy-FSH.md). Fills `FileStatus.kind` and
/// `DirectoryEntry.kind`; `regular = 0` keeps the historical hardcoded value.
pub const NodeKind = enum(u32) {
regular = 0,
directory = 1,
character_device = 2,
block_device = 3,
symbolic_link = 4,
fifo = 5,
socket = 6,
};
/// One directory entry, returned by `readdir`: a fixed header followed inline in
/// the reply payload by `name_len` bytes of name. A zero-length reply is EOF.
pub const DirectoryEntry = extern struct {
kind: u32, // a NodeKind
name_len: u32,
size: u64,
};
pub const directory_entry_size: usize = @sizeOf(DirectoryEntry);
/// Request header. `node` is the server-side open-file id (from a prior open);
/// for `open` the path is the payload and `len` is its length. `offset`/`len`
/// carry the read/write position and count.
@@ -47,6 +81,9 @@ pub const FileStatus = extern struct {
size: u64,
kind: u32,
_padding: u32 = 0,
/// Modification time — Unix epoch seconds, UTC. 0 if the backend has none (the
/// flat ramfs). Filled from the FAT directory entry's write date/time.
mtime: u64 = 0,
};
pub const message_maximum: usize = 256;
@@ -55,5 +92,23 @@ pub const reply_size: usize = @sizeOf(Reply);
/// Largest inline payload that still fits one IPC message alongside a header.
pub const maximum_payload: usize = message_maximum - request_size;
/// Open flags (danos-native; the POSIX layer maps `O_CREAT` onto `create`).
/// Open flags (danos-native; `runtime.fs.OpenOptions` maps its booleans onto these).
pub const create: u32 = 1;
/// Open a directory (for readdir) rather than a file. A mounted backend uses
/// this to open a directory node; the flat ramfs ignores it.
pub const directory: u32 = 2;
/// Truncate the file to zero length on open (O_TRUNC): replace its contents rather
/// than overwriting in place, so a shorter new file leaves no stale tail. A mounted
/// backend frees the old cluster chain; the flat ramfs ignores it.
pub const truncate: u32 = 4;
test "protocol struct sizes and node kinds" {
const std = @import("std");
try std.testing.expectEqual(@as(u32, 0), @intFromEnum(NodeKind.regular));
try std.testing.expectEqual(@as(u32, 1), @intFromEnum(NodeKind.directory));
try std.testing.expectEqual(@as(usize, 16), @sizeOf(DirectoryEntry));
// The appended operations keep the original values.
try std.testing.expectEqual(@as(u32, 0), @intFromEnum(Operation.open));
try std.testing.expectEqual(@as(u32, 4), @intFromEnum(Operation.status));
try std.testing.expectEqual(@as(u32, 5), @intFromEnum(Operation.readdir));
}
+18 -18
View File
@@ -1,26 +1,26 @@
//! /system/services/vfs/vfs-test — a client that proves the VFS round trip end to end: open a
//! file through the `runtime` file API, write to it, seek back, read it, and compare.
//! file through the `runtime.fs` file API, write to it, seek back, read it, and compare.
//! On success it heartbeats "vfstest: ok" so the kernel test can observe it;
//! on failure it reports what went wrong. Shipped in the initial_ramdisk alongside vfs.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
pub fn main(init: runtime.process.Init) void {
const u = @import("posix").unistd;
const payload = "hello-vfs";
// The "park" role (the vfs-client-death test): open a file, then hold the
// handle forever without closing — the kill and the VFS's release-on-death
// are the point.
if (init.arguments.count > 1) {
var fd: i32 = -1;
var parked: ?fs.File = null;
var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) {
fd = u.open("parked", u.O_CREAT);
if (fd < 0) runtime.system.sleep(20);
while (parked == null and tries < 200) : (tries += 1) {
parked = fs.open("parked", .{ .create = true });
if (parked == null) runtime.system.sleep(20);
}
if (fd < 0) {
if (parked == null) {
_ = runtime.system.write("vfstest: park open failed\n");
return;
}
@@ -31,28 +31,28 @@ pub fn main(init: runtime.process.Init) void {
}
// The VFS server may not have registered yet — retry open until it's up.
var fd: i32 = -1;
var opened: ?fs.File = null;
var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) {
fd = u.open("greeting", u.O_CREAT);
if (fd < 0) runtime.system.sleep(20);
while (opened == null and tries < 200) : (tries += 1) {
opened = fs.open("greeting", .{ .create = true });
if (opened == null) runtime.system.sleep(20);
}
if (fd < 0) {
var greeting = opened orelse {
_ = runtime.system.write("vfstest: open failed\n");
return;
}
};
if (u.write(fd, payload) != @as(isize, payload.len)) {
if ((greeting.write(payload) orelse 0) != payload.len) {
_ = runtime.system.write("vfstest: write failed\n");
return;
}
_ = u.lseek(fd, 0, u.SEEK_SET);
greeting.seekTo(0);
var buffer: [32]u8 = undefined;
const n = u.read(fd, &buffer);
u.close(fd);
const n = greeting.read(&buffer) orelse 0;
greeting.close();
if (n == @as(isize, payload.len) and std.mem.eql(u8, buffer[0..@intCast(n)], payload)) {
if (n == payload.len and std.mem.eql(u8, buffer[0..n], payload)) {
while (true) {
_ = runtime.system.write("vfstest: ok\n");
runtime.system.sleep(1000);
+242 -16
View File
@@ -3,14 +3,24 @@
//! file API marshals open/read/write/stat/close into calls to this server's
//! endpoint, published under the well-known `vfs` service id).
//!
//! For now the namespace is a small in-memory ramfs (opening a name creates it):
//! enough to prove the whole path — client file API -> IPC -> server dispatch ->
//! reply. Device nodes backed by user-space drivers (/device) layer on top in M10,
//! where `open` on a /device name forwards to the owning driver's endpoint.
//! Two namespaces meet here (M5):
//! - a small in-memory **ramfs** — opening a bare name creates it — enough to
//! prove the round trip and to back the existing tests;
//! - **mounted filesystems**: a mount table maps an absolute path prefix (e.g.
//! `/mnt/usb`) to a backend server's endpoint. An open of a path under a mount
//! is *forwarded* to that backend (which speaks this same protocol), and every
//! later read/write/status/readdir/close on the resulting handle is relayed to
//! it. The VFS is the router; a filesystem (FAT) is the backend.
//!
//! A path routes through a mount only when it is absolute and lies under a mount
//! prefix; bare names always resolve in the flat ramfs — the backward-compat
//! contract the `vfs` / `vfs-client-death` tests rely on.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.vfs_protocol;
const path = @import("path.zig");
const ipc = runtime.ipc;
const Node = struct {
used: bool = false,
@@ -22,15 +32,29 @@ const Node = struct {
const OpenFile = struct {
used: bool = false,
// For a local handle: an index into `nodes`. For a forwarding handle: the
// node id the backend returned. (usize == u64 here, so it holds either.)
node: usize = 0,
// Non-null for a handle that forwards to a mounted backend.
backend: ?ipc.Handle = null,
// The client (task id — an IPC badge is one) that opened this handle. What
// release-on-death sweeps by: a service must never depend on its clients
// cleaning up after themselves (docs/process-lifecycle.md).
owner: u32 = 0,
};
// One mounted filesystem: an absolute path prefix and the backend endpoint that
// serves everything under it.
const Mount = struct {
used: bool = false,
prefix: [64]u8 = undefined,
prefix_len: usize = 0,
backend: ipc.Handle = 0,
};
var nodes = [_]Node{.{}} ** 8;
var opens = [_]OpenFile{.{}} ** 16;
var mounts = [_]Mount{.{}} ** 8;
fn findNode(name: []const u8) ?usize {
for (&nodes, 0..) |*n, i| {
@@ -57,6 +81,26 @@ fn openAt(id: u64) ?*OpenFile {
return if (o.used) o else null;
}
/// The mount whose prefix most specifically contains `name`, and the path
/// relative to it. Only absolute paths route; bare names never match.
const MountMatch = struct { backend: ipc.Handle, relative: []const u8 };
fn longestMount(name: []const u8) ?MountMatch {
if (!path.isAbsolute(name)) return null;
var best: ?MountMatch = null;
var best_len: usize = 0;
for (&mounts) |*m| {
if (!m.used) continue;
const prefix = m.prefix[0..m.prefix_len];
if (path.underMount(name, prefix)) |relative| {
if (best == null or prefix.len >= best_len) {
best_len = prefix.len;
best = .{ .backend = m.backend, .relative = relative };
}
}
}
return best;
}
/// Serialise a reply header + payload into `out`; returns the total length.
fn writeReply(out: []u8, reply: protocol.Reply, payload: []const u8) usize {
@memcpy(out[0..protocol.reply_size], std.mem.asBytes(&reply));
@@ -76,13 +120,134 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// --- mount routing ----------------------------------------------------------
/// Forward an open under a mount to its backend and, on success, allocate a local
/// forwarding handle that remembers the backend's node id.
fn forwardOpen(out: []u8, backend: ipc.Handle, relative: []const u8, flags: u32, sender: u32) usize {
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = flags };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const rel = relative[0..@min(relative.len, protocol.maximum_payload)];
@memcpy(message[protocol.request_size..][0..rel.len], rel);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + rel.len], &reply) catch return fail(out);
if (n < protocol.reply_size) return fail(out);
const backend_reply = std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]);
if (backend_reply.status != 0) return writeReply(out, .{ .status = backend_reply.status }, &.{});
for (&opens, 0..) |*o, i| {
if (!o.used) {
o.* = .{ .used = true, .node = @intCast(backend_reply.node), .backend = backend, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{});
}
}
return fail(out);
}
/// Relay a read/write/status/readdir/close on a forwarding handle to the backend
/// (the node already rewritten to the backend's id) and copy its reply out.
fn forwardRequest(out: []u8, backend: ipc.Handle, request: protocol.Request, payload: []const u8) usize {
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const plen = @min(payload.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..plen], payload[0..plen]);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + plen], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Forward a path-based operation (mkdir, unlink) under a mount to its backend and
/// relay the reply. No handle is created — these operate by path and return only a
/// status.
fn forwardPath(out: []u8, backend: ipc.Handle, operation: protocol.Operation, relative: []const u8) usize {
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = 0 };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const rel = relative[0..@min(relative.len, protocol.maximum_payload)];
@memcpy(message[protocol.request_size..][0..rel.len], rel);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + rel.len], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Forward a rename to its backend: the payload is the mount-relative old path, a
/// 0x00 separator, then the mount-relative new path. Relays the backend's reply.
fn forwardRename(out: []u8, backend: ipc.Handle, old_relative: []const u8, new_relative: []const u8) usize {
const total = old_relative.len + 1 + new_relative.len;
var message: [protocol.message_maximum]u8 = undefined;
if (protocol.request_size + total > message.len) return fail(out);
const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 };
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
var p = protocol.request_size;
@memcpy(message[p..][0..old_relative.len], old_relative);
p += old_relative.len;
message[p] = 0;
p += 1;
@memcpy(message[p..][0..new_relative.len], new_relative);
p += new_relative.len;
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0..p], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Best-effort close of a backend node (used when a dead client's forwarding
/// handles are swept — the backend must not leak the vfs's opens).
fn forwardClose(backend: ipc.Handle, backend_node: u64) void {
const request = protocol.Request{ .operation = .close, .node = backend_node, .offset = 0, .len = 0, .flags = 0 };
var reply: [protocol.message_maximum]u8 = undefined;
_ = ipc.call(backend, std.mem.asBytes(&request), &reply) catch {};
}
fn doMount(out: []u8, prefix: []const u8, backend: ipc.Handle) usize {
for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefix[0..m.prefix_len], prefix)) {
m.backend = backend;
writeLine("/system/services/vfs: remounted {s}\n", .{prefix});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
for (&mounts) |*m| {
if (!m.used) {
const l = @min(prefix.len, m.prefix.len);
m.used = true;
@memcpy(m.prefix[0..l], prefix[0..l]);
m.prefix_len = l;
m.backend = backend;
writeLine("/system/services/vfs: mounted {s}\n", .{prefix[0..l]});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
return fail(out);
}
fn doUnmount(out: []u8, prefix: []const u8) usize {
for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefix[0..m.prefix_len], prefix)) {
m.used = false;
writeLine("/system/services/vfs: unmounted {s}\n", .{prefix});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
return fail(out);
}
/// Release every open handle `client` held — called on that client's published
/// exit event. The nodes (the files) stay: ramfs contents outlive their writers,
/// only the dead client's handles go.
/// exit event. Forwarding handles also tell their backend to release; local
/// nodes (the ramfs files) stay, since ramfs contents outlive their writers.
fn releaseClientHandles(client: u32) void {
var released: u32 = 0;
for (&opens) |*o| {
if (o.used and o.owner == client) {
if (o.backend) |backend| forwardClose(backend, o.node);
o.used = false;
released += 1;
}
@@ -91,19 +256,32 @@ fn releaseClientHandles(client: u32) void {
}
/// Handle one request from `sender`; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = capability;
fn handle(message: []const u8, out: []u8, sender: u32, capability: ?ipc.Handle) usize {
if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..];
switch (request.operation) {
.mount => {
const prefix = payload[0..@min(payload.len, request.len)];
const backend = capability orelse return fail(out);
return doMount(out, prefix, backend);
},
.unmount => {
const prefix = payload[0..@min(payload.len, request.len)];
return doUnmount(out, prefix);
},
.open => {
const name = payload[0..@min(payload.len, request.len)];
if (longestMount(name)) |m| return forwardOpen(out, m.backend, m.relative, request.flags, sender);
// An absolute path with no matching mount is simply not found — only
// bare names live in the flat ramfs. (Else /mnt/usb would be silently
// created as a flat file when its filesystem is not yet mounted.)
if (path.isAbsolute(name)) return fail(out);
const ni = findNode(name) orelse createNode(name) orelse return fail(out);
for (&opens, 0..) |*o, i| {
if (!o.used) {
o.* = .{ .used = true, .node = ni, .owner = sender };
o.* = .{ .used = true, .node = ni, .backend = null, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{});
}
}
@@ -111,7 +289,12 @@ fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.
},
.read => {
const of = openAt(request.node) orelse return fail(out);
const nd = &nodes[of.node];
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const nd = &nodes[@intCast(of.node)];
const off: usize = @intCast(request.offset);
if (off >= nd.size) return writeReply(out, .{ .status = 0, .len = 0 }, &.{}); // EOF
const n = @min(@min(nd.size - off, request.len), protocol.maximum_payload);
@@ -119,7 +302,12 @@ fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.
},
.write => {
const of = openAt(request.node) orelse return fail(out);
const nd = &nodes[of.node];
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const nd = &nodes[@intCast(of.node)];
const off: usize = @intCast(request.offset);
if (off > nd.data.len) return fail(out);
const n = @min(@min(payload.len, request.len), nd.data.len - off);
@@ -129,20 +317,58 @@ fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.
},
.status => {
const of = openAt(request.node) orelse return fail(out);
const st = protocol.FileStatus{ .size = nodes[of.node].size, .kind = 0 };
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const st = protocol.FileStatus{ .size = nodes[@intCast(of.node)].size, .kind = @intFromEnum(protocol.NodeKind.regular) };
return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&st));
},
.readdir => {
const of = openAt(request.node) orelse return fail(out);
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
// The flat ramfs has no directories: report EOF.
return writeReply(out, .{ .status = 0, .len = 0 }, &.{});
},
.close => {
if (request.node < opens.len) opens[@intCast(request.node)].used = false;
const of = openAt(request.node);
if (of) |o| {
if (o.backend) |backend| forwardClose(backend, o.node);
o.used = false;
}
return writeReply(out, .{ .status = 0 }, &.{});
},
.mkdir, .unlink => {
const name = payload[0..@min(payload.len, request.len)];
if (longestMount(name)) |m| return forwardPath(out, m.backend, request.operation, m.relative);
// Only a mounted backend has real directories; the flat ramfs cannot
// create or remove them (and a bare-name path is not a mount target).
return fail(out);
},
.rename => {
const both = payload[0..@min(payload.len, request.len)];
const sep = std.mem.indexOfScalar(u8, both, 0) orelse return fail(out);
const old_path = both[0..sep];
const new_path = both[sep + 1 ..];
const mo = longestMount(old_path) orelse return fail(out);
const mn = longestMount(new_path) orelse return fail(out);
// Both paths must live under the same mount — cross-filesystem rename is
// not supported.
if (mo.backend != mn.backend) return fail(out);
return forwardRename(out, mo.backend, mo.relative, mn.relative);
},
}
}
/// Startup, under the harness: subscribe to the published exit events — when a
/// client dies holding open handles, the exit notification is how the VFS learns
/// to release them (docs/process-lifecycle.md).
fn initialise(endpoint: runtime.ipc.Handle) bool {
fn initialise(endpoint: ipc.Handle) bool {
if (!runtime.process.subscribeExits(endpoint)) {
_ = runtime.system.write("/system/services/vfs: exit subscription failed\n");
}
@@ -152,8 +378,8 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
/// A non-signal notification: the only kind the VFS subscribes to is exit events.
fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_exit_bit != 0) {
releaseClientHandles(@intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit)));
if (badge & ipc.notify_exit_bit != 0) {
releaseClientHandles(@intCast(badge & ~(ipc.notify_badge_bit | ipc.notify_exit_bit)));
}
}
+131 -23
View File
@@ -25,6 +25,7 @@ import shutil
import socket
import subprocess
import sys
import tempfile
import time
REPO = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
@@ -66,7 +67,14 @@ ARCHES = {
"-machine", "q35", "-m", "128M",
"-drive", f"if=pflash,format=raw,readonly=on,file={a['ovmf_code']}",
"-drive", f"if=pflash,format=raw,file={vars_fd}",
"-drive", f"format=raw,file=fat:rw:{boot_volume}",
# Boot off a FAT USB device: the boot volume is a mass-storage device on
# the xHCI bus (usb-kbd/usb-mouse ride the same controller). `boot_volume`
# is the FAT image the build produces. bootindex=0 steers OVMF to it.
"-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0",
"-drive", f"if=none,id=bootusb,format=raw,file={boot_volume}",
"-device", "usb-storage,bus=xhci.0,drive=bootusb,removable=on,bootindex=0",
"-net", "none",
"-vga", "none", "-device", "VGA,edid=on,xres=1280,yres=720",
"-display", "none",
@@ -100,6 +108,11 @@ CASES = [
{"name": "clock",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Wall-clock: the CMOS RTC read at boot gives a plausible current epoch (the
# foundation for filesystem mtime).
{"name": "wall-clock",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "vmm",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -152,6 +165,26 @@ CASES = [
{"name": "ioport",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Display handoff (D1): the kernel seeds the loader's framebuffer as a claimable
# `display` device with a write-combining memory resource; the claim + mmio_map path
# maps it, and the leaf is genuinely write-combining (PAT entry 4), not the UC default.
{"name": "display",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Display service (D2/D3): the user-space compositor claims the framebuffer, allocates
# a cacheable back buffer, clears it, and presents that composed frame (double-buffer
# path); then a startup self-check composites two overlapping layers and confirms the
# overlap shows the top layer (D3). Matched on the service's own heartbeats.
{"name": "display-service",
"expect": r"display: online \d+x\d+ pitch \d+[\s\S]*display: presented frame 0[\s\S]*display: compositor self-check ok",
"fail": r"display: could not|self-check FAILED|CPU EXCEPTION|KERNEL PANIC"},
# Display demo (D4): a separate process (display-demo) drives the compositor over the
# layer client API — wallpaper + a moving rectangle + a cursor, presented in a loop.
# `display-demo: ok` is printed only after it drove a run of frames of motion through
# the service (the visible motion is a screenshot via `zig build run-x86-64`).
{"name": "display-demo",
"expect": r"display-demo: scene up[\s\S]*display-demo: ok",
"fail": r"display-demo: (no display|create failed)|display: could not|CPU EXCEPTION|KERNEL PANIC"},
# Monotonic clock (clock() syscall source): calibrated, advancing, never backwards.
{"name": "clock",
"expect": r"DANOS-TEST-RESULT: PASS",
@@ -274,9 +307,8 @@ CASES = [
{"name": "usb-report",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
# The xHCI bus + usb-kbd/usb-mouse come from the default boot config now
# (every case boots off a usb-storage device on that bus).
"expect": r"device-manager: child added[\s\S]*"
r"device-manager: child added[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
@@ -284,13 +316,76 @@ CASES = [
r"device-manager: restarting usb-xhci-bus[\s\S]*"
r"device-manager: child added",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# USB HID end to end: boot the full tree, enumerate the xHCI, and let the
# manager spawn the USB keyboard driver, which opens its device over the
# transfer protocol, asks for boot protocol, subscribes to its interrupt
# endpoint, and comes up — proof the class-driver <-> controller path works.
{"name": "usb-hid",
"smp": 4,
"timeout": 150,
# usb-kbd/usb-mouse ride the default boot xHCI bus (see qemu_args).
"expect": r"(?=[\s\S]*usb-hid/keyboard: ok)(?=[\s\S]*usb-hid/mouse: ok)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# USB mass storage end to end: the boot usb-storage device (the FAT32 image,
# which has a real 0x55AA boot sector) is enough — the manager spawns
# usb-storage, which opens the device, runs the Bulk-Only / SCSI bring-up,
# reads its capacity, and reads block 0 (the 0x55AA boot sig). Proof of the
# bulk transfer path + BOT + SCSI end to end.
{"name": "usb-storage",
"smp": 4,
"timeout": 150,
"expect": r"usb-storage: ready[\s\S]*usb-storage: block 0 signature 0x55aa",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# FAT mount end to end: the fat server mounts the boot usb-storage device (the
# FAT32 image) into the VFS at /mnt/usb. A fat-test client then lists and reads
# through the mount — proof of the whole stack: block device -> FAT parse ->
# VFS routing -> file read.
{"name": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat: mounted /mnt/usb[\s\S]*fat-test: ok",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the
# fat-test client, after listing, makes a directory, writes+reads a file inside
# it, then removes the file, exercising the whole VFS -> fat mutation path.
{"name": "fat-mutations",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: mutations ok",
"fail": r"fat-test: mutations FAILED|fat-test: mkdir .* failed|DANOS-TEST-RESULT: FAIL"},
# Phase 2c: rename through the mount — fat-test renames the file it created
# before removing it, and confirms the old name is gone.
{"name": "fat-rename",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: rename ok",
"fail": r"fat-test: mutations FAILED|DANOS-TEST-RESULT: FAIL"},
# Phase 2d: filesystem timestamps — a freshly-created file's mtime is a real
# current wall-clock time (stamped from the RTC), read back through stat.
{"name": "fat-mtime",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: mtime ok",
"fail": r"fat-test: mutations FAILED|DANOS-TEST-RESULT: FAIL"},
# Boot-from-USB smoke: the whole system now boots off the FAT32 image on a
# usb-storage device (OVMF -> \EFI\BOOT\BOOTX64.efi -> kernel), so the kernel
# reaching its PASS marker at all proves the USB boot path end to end. Reuses
# the smoke kernel build; the value is the explicit, named regression guard.
{"name": "usb-boot",
"build_case": "smoke",
"qmp_after": {"delay": 2, "command": "query-status"},
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.1: the ring-3 AML parse (the acpi service maps the blobs and parses
# them) finds exactly the Device count the kernel's own parse produced.
{"name": "acpi-parse",
"smp": 4,
"timeout": 60,
"expect": r"acpi-parse: ok",
"fail": r"acpi-parse: mismatch|DANOS-TEST-RESULT: FAIL"},
"fail": r"acpi-parse: too few|DANOS-TEST-RESULT: FAIL"},
# M20.3: the flip — ps2-bus now comes up from the acpi service's report, not
# a kernel-built node. Ordered: report -> spawn -> the driver attaches its
# keyboard, proving discovery runs entirely in ring 3 (docs/discovery.md).
@@ -323,15 +418,26 @@ CASES = [
r"init: shutting down[\s\S]*"
r"power: entering S5",
"fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"},
# M8: the boot log is persisted to the USB FAT volume. Reuses the orderly-
# shutdown build (full tree + power button): init spawns log-flush at boot,
# which copies the kernel log to /mnt/usb/DANOS.LOG once /mnt/usb is mounted
# (first marker); then the power button drives init's own pre-teardown flush
# (second marker), proving both triggers write the file while storage is up.
{"name": "log-flush",
"build_case": "orderly-shutdown",
"smp": 4,
"timeout": 150,
"qmp_after": {"delay": 8, "command": "system_powerdown"},
"expect": r"log-flush: wrote \d+ bytes to /mnt/usb/DANOS\.LOG[\s\S]*"
r"init: flushed log to /mnt/usb/DANOS\.LOG[\s\S]*"
r"power: entering S5",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers +
# reports its _HID devices — the two PS/2 nodes must appear with resources
# (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/discovery.md).
{"name": "acpi-report",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"acpi: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*"
r"acpi: reported PNP0F13 \(device \d+, 1 resources\)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -348,9 +454,6 @@ CASES = [
{"name": "device-list",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-list: \d+ devices[\s\S]*"
r"device-list: subscribed[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
@@ -363,9 +466,6 @@ CASES = [
{"name": "driver-restart",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"usb-xhci-bus: hello acknowledged[\s\S]*"
r"device-manager: restarting crash-test[\s\S]*"
r"device-manager: crash-test is failing repeatedly",
@@ -412,9 +512,8 @@ CASES = [
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The ACPI power path succeeds by QEMU *exiting* (S5 off / reset), so match the
# pre-transition marker; the FAIL line only appears if the transition didn't take.
{"name": "poweroff",
"expect": r"DANOS-POWER: attempting poweroff",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Soft-off (S5) is owned by the ring-3 acpi service now (see orderly-shutdown);
# the kernel keeps only reboot (FADT reset register, no AML).
{"name": "reboot",
"expect": r"DANOS-POWER: attempting reboot",
"fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -424,7 +523,10 @@ TIMEOUT = 30 # seconds per case
def build(arch, case):
cmd = ["zig", "build", f"-Dtest-case={case}"] + arch["zig_flags"]
# -Dserial: the harness asserts on markers the kernel writes to serial0, so the
# serial log sink must be compiled in. It is off by default (a flashed real-
# hardware image keeps its log in RAM instead; see build.zig / serial.zig).
cmd = ["zig", "build", f"-Dtest-case={case}", "-Dserial=true"] + arch["zig_flags"]
r = subprocess.run(cmd, cwd=REPO, capture_output=True, text=True)
if r.returncode != 0:
return r.stderr.strip() or r.stdout.strip()
@@ -471,12 +573,15 @@ def qmp_send(path, command):
def run_case(arch, case):
err = build(arch, case["name"])
# A case's kernel build defaults to its name; `build_case` decouples the two
# so a case can reuse another's kernel (e.g. usb-boot reuses smoke's).
err = build(arch, case.get("build_case", case["name"]))
if err:
return False, "build failed:\n" + err
# zig-out is the FHS boot volume; hand it to the guest as-is (see qemu_args).
boot_volume = os.path.join(REPO, "zig-out")
# The bootable FAT32 USB image the build produced (tools/make-fat-image.py),
# presented to the guest as a usb-storage device (see qemu_args).
boot_volume = os.path.join(REPO, "zig-out", "danos-usb.img")
vars_fd = os.path.join(WORK, "vars.fd")
shutil.copy(arch["ovmf_vars"], vars_fd)
serial = os.path.join(WORK, "serial.log")
@@ -492,8 +597,11 @@ def run_case(arch, case):
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
cmd += case["qemu_extra"]
# A QMP control socket, always present (additive): how a case's `qmp_after`
# hook injects host-side events into the guest mid-run.
qmp_path = os.path.join(WORK, "qmp.sock")
# hook injects host-side events into the guest mid-run. Kept under a short temp
# dir, not WORK: a unix socket path is capped at ~104 bytes (sun_path), and a
# deep worktree path (e.g. .claude/worktrees/<name>/zig-out/qemu-test/qmp.sock)
# blows that limit on macOS, so QEMU fails to bind and exits before booting.
qmp_path = os.path.join(tempfile.gettempdir(), f"danos-qmp-{os.getpid()}.sock")
if os.path.exists(qmp_path):
os.remove(qmp_path)
cmd += ["-qmp", f"unix:{qmp_path},server,nowait"]
+362
View File
@@ -0,0 +1,362 @@
#!/usr/bin/env python3
"""Format a real FAT32 image from a set of host files — the danos boot volume.
Mirrors tools/make-initial-ramdisk.py in spirit: pure Python 3 standard library,
no external tools (no mkfs.fat / mtools). It writes a valid FAT32 filesystem — a
boot sector + BPB, an FSInfo sector, a backup boot sector, two FATs, and a
directory tree of clusters — so UEFI/OVMF boots \\EFI\\BOOT\\BOOTX64.efi off it
and the danos FAT driver mounts the same image.
make-fat-image.py <out.img> <size-MiB> [<dest-path> <host-file>]...
make-fat-image.py --verify <out.img>
Each <dest-path> is a forward-slash path inside the image (e.g.
"EFI/BOOT/BOOTX64.efi"); intermediate directories are created. Names that do not
fit 8.3 get a mangled short name plus long-file-name (LFN) entries.
"""
import struct
import sys
SECTOR = 512
SECTORS_PER_CLUSTER = 1 # 512-byte clusters keep the cluster count high for FAT32
RESERVED_SECTORS = 32
NUM_FATS = 2
CLUSTER_BYTES = SECTOR * SECTORS_PER_CLUSTER
END_OF_CHAIN = 0x0FFFFFFF
BAD_CLUSTER = 0x0FFFFFF7
ATTR_ARCHIVE = 0x20
ATTR_DIRECTORY = 0x10
ATTR_LONG_NAME = 0x0F
VALID_83 = set("ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789$%'-_@~!(){}^#& ")
def fat32_geometry(total_sectors):
"""Solve for the FAT size (sectors per FAT) and cluster count that fit."""
fat_size = 1
while True:
data_sectors = total_sectors - RESERVED_SECTORS - NUM_FATS * fat_size
cluster_count = data_sectors // SECTORS_PER_CLUSTER
needed = ((cluster_count + 2) * 4 + SECTOR - 1) // SECTOR
if needed <= fat_size:
return fat_size, cluster_count
fat_size = needed
class Fat32Image:
def __init__(self, total_sectors):
self.total_sectors = total_sectors
self.fat_size, self.cluster_count = fat32_geometry(total_sectors)
if self.cluster_count < 65525:
sys.exit(f"error: image too small for FAT32 ({self.cluster_count} clusters "
f"< 65525); use a larger size")
self.first_data_sector = RESERVED_SECTORS + NUM_FATS * self.fat_size
# The FAT, in memory: entry 0 media, entry 1 EOC, entry 2 the root dir.
self.fat = [0] * (self.cluster_count + 2)
self.fat[0] = 0x0FFFFFF8
self.fat[1] = END_OF_CHAIN
self.fat[2] = END_OF_CHAIN
self.next_free = 3
self.cluster_data = {} # cluster number -> bytes (one cluster's worth)
def alloc(self):
cluster = self.next_free
if cluster >= self.cluster_count + 2:
sys.exit("error: image out of clusters")
self.next_free += 1
self.fat[cluster] = END_OF_CHAIN
return cluster
def store_chain(self, content):
"""Allocate a cluster chain holding `content` and return its first cluster."""
length = max(1, (len(content) + CLUSTER_BYTES - 1) // CLUSTER_BYTES)
clusters = [self.alloc() for _ in range(length)]
for i in range(length - 1):
self.fat[clusters[i]] = clusters[i + 1]
for i, cluster in enumerate(clusters):
chunk = content[i * CLUSTER_BYTES:(i + 1) * CLUSTER_BYTES]
self.cluster_data[cluster] = chunk + b"\x00" * (CLUSTER_BYTES - len(chunk))
return clusters[0]
def store_directory(self, first_cluster, entries):
"""Write directory `entries` (bytes) into `first_cluster`, extending the chain."""
length = max(1, (len(entries) + CLUSTER_BYTES - 1) // CLUSTER_BYTES)
clusters = [first_cluster]
for _ in range(length - 1):
clusters.append(self.alloc())
for i in range(len(clusters) - 1):
self.fat[clusters[i]] = clusters[i + 1]
for i, cluster in enumerate(clusters):
chunk = entries[i * CLUSTER_BYTES:(i + 1) * CLUSTER_BYTES]
self.cluster_data[cluster] = chunk + b"\x00" * (CLUSTER_BYTES - len(chunk))
def cluster_sector(self, cluster):
return self.first_data_sector + (cluster - 2) * SECTORS_PER_CLUSTER
def serialize(self):
image = bytearray(self.total_sectors * SECTOR)
image[0:SECTOR] = self.boot_sector()
image[SECTOR:2 * SECTOR] = self.fsinfo_sector()
image[6 * SECTOR:7 * SECTOR] = self.boot_sector() # backup boot sector
# Both FATs.
fat_bytes = b"".join(struct.pack("<I", entry & 0x0FFFFFFF) for entry in self.fat)
fat_bytes += b"\x00" * (self.fat_size * SECTOR - len(fat_bytes))
for copy in range(NUM_FATS):
base = (RESERVED_SECTORS + copy * self.fat_size) * SECTOR
image[base:base + len(fat_bytes)] = fat_bytes
# The data region (clusters).
for cluster, data in self.cluster_data.items():
base = self.cluster_sector(cluster) * SECTOR
image[base:base + len(data)] = data
return bytes(image)
def boot_sector(self):
sector = bytearray(SECTOR)
# BPB.
struct.pack_into(
"<3s8sHBHBHHBHHHII", sector, 0,
b"\xEB\x58\x90", # jump
b"MSWIN4.1", # OEM name (widest firmware compatibility)
SECTOR, # bytes per sector
SECTORS_PER_CLUSTER, # sectors per cluster
RESERVED_SECTORS, # reserved sector count
NUM_FATS, # number of FATs
0, # root entry count (0 for FAT32)
0, # total sectors 16 (0 -> use 32)
0xF8, # media descriptor
0, # FAT size 16 (0 for FAT32)
32, # sectors per track
2, # heads
0, # hidden sectors
self.total_sectors, # total sectors 32
)
# FAT32 extended BPB (offset 36).
struct.pack_into(
"<IHHIHH12sBBBI11s8s", sector, 36,
self.fat_size, # FAT size 32
0, # extended flags
0, # filesystem version
2, # root cluster
1, # FSInfo sector
6, # backup boot sector
b"\x00" * 12, # reserved
0x80, # drive number
0, # reserved
0x29, # extended boot signature
0x12345678, # volume id
b"DANOS ", # volume label
b"FAT32 ", # filesystem type
)
sector[510] = 0x55
sector[511] = 0xAA
return bytes(sector)
def fsinfo_sector(self):
sector = bytearray(SECTOR)
struct.pack_into("<I", sector, 0, 0x41615252) # lead signature
struct.pack_into("<I", sector, 484, 0x61417272) # struct signature
free = self.cluster_count - (self.next_free - 2)
struct.pack_into("<I", sector, 488, free) # free count
struct.pack_into("<I", sector, 492, self.next_free) # next free hint
struct.pack_into("<I", sector, 508, 0xAA550000) # trail signature
return bytes(sector)
def lfn_checksum(short_name):
checksum = 0
for byte in short_name:
checksum = (((checksum & 1) << 7) + (checksum >> 1) + byte) & 0xFF
return checksum
def short_name_for(name, used):
"""Return (raw 11-byte 8.3 name, needs_lfn)."""
if "." in name and not name.startswith("."):
base, ext = name.rsplit(".", 1)
else:
base, ext = name, ""
upper_base, upper_ext = base.upper(), ext.upper()
# A name fits 8.3 if it is short enough and uses valid characters; a lowercase
# name is simply stored uppercased (FAT is case-insensitive, so the bootloader
# and the danos driver still find it). Only genuinely non-8.3 names (too long,
# e.g. initial-ramdisk.img) get a mangled short name plus LFN entries.
fits = (1 <= len(base) <= 8 and len(ext) <= 3
and all(c in VALID_83 for c in upper_base + upper_ext))
if fits:
return (upper_base.ljust(8) + upper_ext.ljust(3)).encode("ascii"), False
# Mangle to STEM~N.EXT.
stem = "".join(c for c in upper_base if c in VALID_83 and c != " ")[:6] or "FILE"
index = 1
while True:
candidate = f"{stem}~{index}".ljust(8)[:8] + upper_ext.ljust(3)[:3]
raw = candidate.encode("ascii")
if raw not in used:
used.add(raw)
return raw, True
index += 1
def lfn_entries(name, short_raw):
checksum = lfn_checksum(short_raw)
units = list(name.encode("utf-16-le"))
pairs = [bytes(units[i:i + 2]) for i in range(0, len(units), 2)]
pairs.append(b"\x00\x00") # null terminator
while len(pairs) % 13 != 0:
pairs.append(b"\xff\xff")
count = len(pairs) // 13
out = bytearray()
for sequence in range(count, 0, -1): # stored last-logical-first
piece = pairs[(sequence - 1) * 13:sequence * 13]
entry = bytearray(32)
entry[0] = sequence | (0x40 if sequence == count else 0)
for i in range(5):
entry[1 + i * 2:1 + i * 2 + 2] = piece[i]
entry[11] = ATTR_LONG_NAME
entry[12] = 0
entry[13] = checksum
for i in range(6):
entry[14 + i * 2:14 + i * 2 + 2] = piece[5 + i]
entry[26:28] = b"\x00\x00"
for i in range(2):
entry[28 + i * 2:28 + i * 2 + 2] = piece[11 + i]
out += entry
return bytes(out)
def short_entry(raw11, attributes, cluster, size):
return struct.pack(
"<11sBBBHHHHHHHI",
raw11, attributes, 0, 0, 0, 0, 0,
(cluster >> 16) & 0xFFFF, 0, 0, cluster & 0xFFFF, size,
)
def write_directory(image, cluster, children, parent_cluster, is_root):
"""Recursively lay out a directory: allocate child clusters, build entries."""
entries = bytearray()
if not is_root:
entries += short_entry(b". ", ATTR_DIRECTORY, cluster, 0)
parent = 0 if parent_cluster == 2 else parent_cluster
entries += short_entry(b".. ", ATTR_DIRECTORY, parent, 0)
used_short_names = set()
for name, child in children.items():
raw, needs_lfn = short_name_for(name, used_short_names)
used_short_names.add(raw)
if child["type"] == "dir":
child_cluster = image.alloc()
if needs_lfn:
entries += lfn_entries(name, raw)
entries += short_entry(raw, ATTR_DIRECTORY, child_cluster, 0)
write_directory(image, child_cluster, child["children"], cluster, False)
else:
data = child["data"]
first = image.store_chain(data) if data else 0
if needs_lfn:
entries += lfn_entries(name, raw)
entries += short_entry(raw, ATTR_ARCHIVE, first, len(data))
image.store_directory(cluster, bytes(entries))
def build_tree(pairs):
root = {}
for dest, host in pairs:
with open(host, "rb") as handle:
data = handle.read()
parts = [p for p in dest.replace("\\", "/").split("/") if p]
node = root
for part in parts[:-1]:
node = node.setdefault(part, {"type": "dir", "children": {}})["children"]
node[parts[-1]] = {"type": "file", "data": data}
return root
def build(out_path, size_mib, pairs):
total_sectors = size_mib * 1024 * 1024 // SECTOR
image = Fat32Image(total_sectors)
tree = build_tree(pairs)
write_directory(image, 2, tree, 0, True)
with open(out_path, "wb") as handle:
handle.write(image.serialize())
print(f"make-fat-image: wrote {out_path} "
f"({size_mib} MiB FAT32, {image.cluster_count} clusters)")
def verify(path):
with open(path, "rb") as handle:
data = handle.read()
if len(data) < SECTOR or data[510] != 0x55 or data[511] != 0xAA:
sys.exit("verify: missing 0x55AA boot signature")
bytes_per_sector, sectors_per_cluster = struct.unpack_from("<HB", data, 11)
reserved, num_fats = struct.unpack_from("<H", data, 14)[0], data[16]
fat_size_32, root_cluster = struct.unpack_from("<I", data, 36)[0], struct.unpack_from("<I", data, 44)[0]
total_sectors = struct.unpack_from("<I", data, 32)[0]
if bytes_per_sector != SECTOR or sectors_per_cluster == 0 or num_fats == 0 or fat_size_32 == 0:
sys.exit("verify: implausible BPB")
first_data = reserved + num_fats * fat_size_32
cluster_count = (total_sectors - first_data) // sectors_per_cluster
if cluster_count < 65525:
sys.exit(f"verify: not FAT32 ({cluster_count} clusters)")
# Resolve EFI/BOOT/BOOTX64.efi through the directory tree to prove it is present.
if not _resolve(data, ["EFI", "BOOT", "BOOTX64.EFI"], root_cluster,
reserved, num_fats, fat_size_32, first_data, sectors_per_cluster):
sys.exit("verify: EFI/BOOT/BOOTX64.efi not found")
print(f"verify: {path} is FAT32 ({cluster_count} clusters); EFI/BOOT/BOOTX64.efi present")
def _read_fat(data, cluster, reserved):
offset = reserved * SECTOR + cluster * 4
return struct.unpack_from("<I", data, offset)[0] & 0x0FFFFFFF
def _resolve(data, parts, cluster, reserved, num_fats, fat_size, first_data, spc):
for part in parts:
cluster = _find(data, cluster, part, reserved, first_data, spc)
if cluster is None:
return False
return True
def _find(data, dir_cluster, name, reserved, first_data, spc):
target = name.upper()
cluster = dir_cluster
guard = 0
while cluster >= 2 and cluster < BAD_CLUSTER and guard < 100000:
sector = first_data + (cluster - 2) * spc
for s in range(spc):
base = (sector + s) * SECTOR
for i in range(SECTOR // 32):
entry = data[base + i * 32:base + i * 32 + 32]
if entry[0] == 0x00:
return None
if entry[0] == 0xE5 or (entry[11] & ATTR_LONG_NAME) == ATTR_LONG_NAME:
continue
raw = entry[0:11]
short = (raw[0:8].rstrip().decode("latin1") +
("." + raw[8:11].rstrip().decode("latin1") if raw[8:11].strip() else "")).upper()
if short == target:
return ((entry[20] | (entry[21] << 8)) << 16) | (entry[26] | (entry[27] << 8))
cluster = _read_fat(data, cluster, reserved)
guard += 1
return None
def main(argv):
if len(argv) == 3 and argv[1] == "--verify":
verify(argv[2])
return 0
if len(argv) < 3 or (len(argv) - 3) % 2 != 0:
sys.exit("usage: make-fat-image.py <out.img> <size-MiB> [<dest> <host>]...\n"
" make-fat-image.py --verify <out.img>")
out_path = argv[1]
size_mib = int(argv[2])
rest = argv[3:]
pairs = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)]
build(out_path, size_mib, pairs)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))