Author SHA1 Message Date
Daniel Samson 67702fa250 Phase 2d (ii): filesystem modification time (mtime)
Completes Phase 2: the FAT filesystem now stamps and reports a real modification
time, built on the Phase 2d(i) kernel wall-clock. This is the last stat field the
compiler's build cache needs to reason about (source vs cached output).

- on-disk.zig: fatToEpoch / epochToFatDateTime convert between the two 16-bit DOS
  date/time fields and Unix epoch seconds (UTC — FAT has no timezone). Host-tested
  round-trip + an absolute check (1577836800 == 2020-01-01).
- engine: a settable current_time_epoch that create/write stamp into the entry's
  write (and creation) date/time; Node/Listing gained an mtime decoded from those
  fields on read. Host test: a create stamps the mtime, read back through resolve
  and listEntry.
- vfs protocol FileStatus + runtime.fs.Attributes gained an mtime field; the fat
  server sets current_time_epoch from runtime.system.wallClock() per request and
  returns mtime from stat. The flat ramfs reports 0 (it has no timestamps).
- fat-test reads the created file's mtime through stat and checks it is a real
  current time, behind a new `fat-mtime` QEMU case.

Verified against the host: the guest stamped mtime 1783971676 while the host clock
was 1783971680 (boot+test lag) — the file's mtime is real current time. zig build,
zig build test (the epoch<->DOS conversions + the engine mtime test),
zig build check-fat-image, and a sequential QEMU sweep — fat-mount, fat-mutations,
fat-rename, fat-mtime, vfs, vfs-client-death, log-flush, orderly-shutdown,
initial-ramdisk, smoke, wall-clock, usb-storage — all green. mode/inode remain.
2026-07-13 20:43:40 +01:00
Daniel Samson 8a38540312 Phase 2d (i): a kernel wall-clock from the CMOS RTC
Adds real (calendar) time, the foundation for filesystem timestamps. Monotonic
time (`clock`) says how long since boot; this says what time it actually is.

- cpu.zig (x86_64): readRtcUnixSeconds() reads the CMOS real-time clock (ports
  0x70/0x71) — waits out an update-in-progress, reads twice until stable, handles
  BCD-vs-binary and 12-vs-24-hour per status register B — and converts to Unix
  epoch seconds (UTC).
- kernel/wall-clock.zig: reads the RTC once at boot and anchors it to the monotonic
  clock, so a query is a cheap arithmetic offset — no per-call CMOS poll, no lock,
  no SMP hazard on the shared ports. kmain calls init() once the monotonic clock is
  final and logs the epoch.
- wall_clock() syscall (33) -> Unix epoch seconds, wrapped by runtime.system
  .wallClock(). Wall-clock *seconds* are mechanism the kernel owns like the
  monotonic clock; calendars/timezones are user-space policy (the stale comment on
  systemClock that called wall-clock a "user-space service" is updated in spirit by
  the new handler's doc).
- A `wall-clock` kernel test asserts the boot RTC read is a plausible current epoch.

Verified against the host: the guest read epoch 1783971244 while `date -u +%s` gave
1783971245 (one second of boot lag) — the CMOS read + epoch conversion are correct to
the second. zig build, zig build test, and smoke/clock/init are green.
2026-07-13 20:35:03 +01:00
Daniel Samson 54635eecf5 Phase 2c: rename, wired through the VFS to runtime.fs
Completes the Phase 2 FAT mutation set (truncate, mkdir, unlink, rename).

- engine: rename(dir, old_name, new_name) rewrites an existing entry's 8.3 name in
  place within the same directory. Refuses a missing source, a non-8.3 target, or a
  name that already exists; drops any long-name entries on the old file (it takes
  its new 8.3 name), LFN-aware like removeFile. Cross-directory and long-name-
  preserving rename are noted limitations. Host-tested (rename keeps contents;
  collision, non-8.3, and missing-source are refused).
- vfs protocol: a `rename` operation whose payload is old-path, a 0x00 separator,
  then new-path.
- VFS router: a forwardRename helper + a `.rename` case that requires both paths
  under the same mount (cross-filesystem rename is refused) and forwards the
  mount-relative old+new.
- fat server: a `.rename` handler that requires the same parent directory and calls
  engine.rename.
- runtime.fs: rename(old_path, new_path).
- fat-test now renames the file it created (before removing it) and asserts the old
  name is gone, behind a new `fat-rename` QEMU case.

Verified: zig build, zig build test (the engine rename unit test), zig build
check-fat-image, and a sequential QEMU sweep — fat-mount, fat-mutations, fat-rename,
vfs, vfs-client-death, log-flush, orderly-shutdown, initial-ramdisk, smoke — green.
2026-07-13 20:22:32 +01:00
Daniel Samson 184d90c2c6 Phase 2b: wire mkdir + unlink through the VFS to runtime.fs
The engine gained mkdir/unlink in Phase 2a; this exposes them as first-class
filesystem operations so programs can use them.

- vfs protocol: two new path-based operations, mkdir and unlink (appended, so
  existing opcodes/offsets are unchanged).
- VFS router: a forwardPath helper relays a path-based op under a mount to its
  backend; the mkdir/unlink cases forward to the mounted filesystem (the flat
  ramfs refuses them — it has no directories).
- fat server: mkdir -> engine.createDirectory, unlink -> engine.removeFile, each
  resolving the parent via a shared splitParent helper (also used by open-create).
- runtime.fs: makeDirectory(path) and remove(path).
- fat-test now exercises the whole path — mkdir /mnt/usb/TESTDIR, create + write +
  read a file inside it, then remove it — behind a new `fat-mutations` QEMU case.

Verified: zig build, zig build test, zig build check-fat-image, and a sequential
QEMU sweep — fat-mount, fat-mutations (mkdir/write/read/unlink through the mount),
vfs, vfs-client-death, log-flush, orderly-shutdown, initial-ramdisk, smoke — green.
2026-07-13 20:12:07 +01:00
Daniel Samson a32eed877d Phase 2a: FAT engine mutations + O_TRUNC (fix the overwrite corruption)
The FAT engine could create, read, write, and grow files, but never free clusters
or make directories — so overwriting a shorter file left a stale tail (a real bug:
corrupt boot-log re-flushes, and later corrupt compiler cache/.o files). This adds
the mutation half of the engine, with the corruption fix wired all the way through.

Engine (system/services/fat/engine.zig), all host-tested:
- freeChain: return a cluster chain to the pool (bounded against a corrupt cycle) —
  the shared primitive under truncate and remove.
- truncate: free the chain and zero the entry's size/first-cluster (O_TRUNC).
- createDirectory (mkdir): allocate + initialise a cluster with "." and ".." and add
  the directory entry to the parent.
- removeFile (unlink): free the chain and mark the 8.3 entry plus any preceding
  long-name entries deleted, so a reused slot can't inherit an orphaned long name.
  Refuses directories.
- createFile now shares a common addEntry helper with createDirectory.

O_TRUNC wired end to end: a truncate open-flag (vfs protocol) that the router already
forwards; runtime.fs.OpenOptions.truncate; and fat's handleOpen calls engine.truncate
on an existing file. The boot-log flush (log-flush + init) now opens with truncate, so
a shorter log on a later boot of the same stick leaves no stale tail — closing the
caveat from the boot-log work.

mkdir/unlink are engine-complete and host-tested but not yet exposed as VFS
operations / runtime.fs methods (they are new path-based ops needing router cases);
rename and the richer stat (mtime/mode, blocked on wall-clock) remain. See
docs/zig-self-hosting.md (Phase 2).

Verified: zig build, zig build test (8 engine host tests, incl. truncate, the
overwrite-no-stale-tail regression, remove, and mkdir), zig build check-fat-image, and
a sequential QEMU sweep — fat-mount, log-flush (DANOS.LOG read back at 11804 bytes,
clean, with the truncate-based final flush), vfs, vfs-client-death, usb-storage,
orderly-shutdown, initial-ramdisk, smoke — all green.
2026-07-13 20:02:01 +01:00
Daniel Samson 347a041d85 Phase 1a: add the native runtime.fs, retire the posix shim
The first step of the Zig self-hosting roadmap (docs/zig-self-hosting.md): give
danos programs a danos-native file API and remove the premature POSIX compatibility
shim. This also resolves the earlier misplacement of a full-write helper into the
compat layer — that behaviour now lives natively in runtime.fs.File.writeAll.

- library/runtime/fs.zig: the danos-native file client over the VFS (open/read/
  write/writeAll/seekTo/attributes/close, directory listing, mount). Handles are
  *values* — a File/Directory owns its VFS node id and byte offset — so there is no
  per-process fd table or descriptor limit, unlike the POSIX fd model the shim
  emulated. This is where the operations that later become std.os.danos are staged.
- Retire library/posix/ (unistd, stdio): only five call sites used it, all file
  operations, all migrated to runtime.fs — fat (mount), the vfs-test and fat-test
  clients, and init/log-flush (the boot-log flush). stdio was already dead.
- build.zig: drop the posix module, its addUserBinary parameter, the per-binary
  import, and the ~26 call-site arguments.
- Docs: the VFS protocol's client is now runtime.fs; the docs index and
  coding-standards note posix is retired and the foreign-ABI naming exception now
  applies to the future std.os.danos seam; the process-lifecycle note points the
  future musl layer at that same seam rather than the deleted directory.

Deferred by design (see the roadmap): the C-ABI runtime.os errno seam is built at
fork time (its shape must match std/os/danos.zig); truncate/mkdir/rename are
Phase 2; stdio-byte fds and cwd are later slices.

Verified: zig build, zig build test, zig build check-fat-image, and a sequential
QEMU sweep — vfs, vfs-client-death (the park/hold-handle path), fat-mount, log-flush,
orderly-shutdown, initial-ramdisk (log-flush silent in the bare sweep), smoke, init,
usb-storage, device-manager — all green.
2026-07-13 19:43:56 +01:00
Daniel Samson 53e42837e0 Docs: add the Zig self-hosting roadmap
A forward-looking design note on making danos a real Zig target
(-target x86_64-danos) and eventually running the compiler on it, focused on
the standard-library surface (not the editor/terminal).

The core realisation: Zig 0.16 (post-writergate) collapses an OS port to ONE
seam — std.fs is gone, everything routes through the std.Io vtable, and
std.posix is generic over a single per-OS `system` module (std.os.<tag>). So the
port is "write std.os.danos once" and the whole fs/process/Io tower lights up,
rather than reimplementing the namespaces.

Records the decisions this shapes now: build runtime.os (the seam, promoted into
a forked std/os/danos.zig later) plus a thin runtime.fs; retire the premature
library/posix shim (only 5 unistd call sites); do NOT hand-mirror the high-level
std namespaces; do NOT emulate the Linux ABI; defer musl. Covers the host/target/
self-host roles and the four-part compiler fork, a coverage table of what danos
has vs the gaps (mkdir/unlink/rename/truncate, richer stat, wall-clock, env, cwd,
entropy, stdio bytes), a phased plan (target -> read-side+retire-posix ->
fs-mutation+stat -> single-threaded self-linked compiler), and the risks
(fork rebase treadmill, -fsingle-threaded and -fno-llvm/-fno-lld being
load-bearing, the "w"-does-not-truncate corruption bug). Linked from the docs index.
2026-07-13 19:26:51 +01:00
Daniel Samson f52c591f5e Persist the kernel boot log to the USB FAT volume
On a headless or real board nothing captures serial, so the boot log — the whole
diagnostic stream — is lost at power-off. This retains it in the kernel and copies
it to the boot USB volume as /mnt/usb/DANOS.LOG, the on-disk equivalent of QEMU's
`-serial file:`. Pull the stick, read DANOS.LOG on another machine.

How it fits together:
- Kernel RAM sink (log.zig): a fixed 256 KiB in-image buffer registered as a log
  sink in kmain, right after serial. Because userspace debug_write funnels through
  log.write, it captures the entire stream — kernel lines and every service's
  output — from the first line. Fills linearly and stops when full (earliest boot
  output, the most valuable, is kept); no allocation, so it is panic-safe.
- klog_read syscall (32): copies that buffer out to a user buffer, the mirror of
  debug_write — same overflow-safe user-half bounds check, kernel -> user copy,
  under the kernel lock so the snapshot can't grow mid-copy. Wrapped by
  runtime.system.klogRead.
- log-flush (new one-shot, in the initial-ramdisk): waits for the fat server to
  mount /mnt/usb, then copies the whole log to /mnt/usb/DANOS.LOG. init spawns it
  once the boot services are up (fire-and-forget; it polls the mount itself). If
  no volume is mounted — no stick, or the no-VFS ramdisk sweep — it exits silently.
- init shutdown flush: init repeats the copy inline at the top of shutDown(),
  BEFORE it tears down the storage services (the fat server is stopped first), so
  a clean poweroff captures the fullest log while /mnt/usb is still writable.
- unistd.writeAll: loops write() past the 224-byte VFS payload cap; both flush
  paths use it.

The filename is 8.3 (DANOS.LOG) at the mount root — the FAT short-name rule, and
there is no mkdir on the FAT path yet. Extend-only writes mean the two same-session
flushes never leave stale bytes (the shutdown log is a superset of the boot log);
a shorter log on a later boot of the same stick can leave a stale tail — a noted,
cosmetic limitation, not worth pulling O_TRUNC into the FAT write path for now.

Verified end to end under QEMU: a new `log-flush` case (reusing the orderly-shutdown
build) asserts both markers then S5, and DANOS.LOG is read back out of the image
afterwards (11804 bytes, containing the kernel init line, the FAT mount line, and
the boot flush marker). Regression stays green: zig build, zig build test, zig
build check-fat-image, and a sequential QEMU sweep — smoke, init, initial-ramdisk
(log-flush silent in the bare sweep), orderly-shutdown, fat-mount, usb-storage,
usb-hid, vfs, input, device-manager, process, signals, dma, fault-pf.
2026-07-13 18:18:34 +01:00
Daniel Samson 77d2e22ed1 M7: in-repo FAT32 image builder + boot the whole system off a USB stick
The system now boots off a real FAT32 filesystem on a USB mass-storage device
instead of QEMU's synthesized VVFAT drive. A new in-repo image builder formats
that filesystem from the FHS boot tree, and both QEMU call sites (the run step
and the test harness) attach it as a usb-storage device on the xHCI bus, so
every boot exercises the full USB path OVMF -> BOOTX64.efi -> kernel.

- tools/make-fat-image.py: a Python 3 stdlib-only FAT32 formatter (mirrors
  tools/make-initial-ramdisk.py — no external host dependencies). It lays down
  the boot sector + BPB/EBPB32, FSInfo, backup boot sector, two FATs, and the
  root/subdir/file cluster chains, emitting long-name entries where a name is
  not 8.3. Packs the four boot inputs (EFI/BOOT/BOOTX64.efi, system/kernel,
  system/services/init, boot/initial-ramdisk.img) into their boot paths. A
  --verify subcommand re-checks the 0xAA55 signature, recomputes the cluster
  count -> FAT32, and resolves EFI/BOOT/BOOTX64.efi, all with no dependencies.

- build.zig: a mk_fat step builds zig-out/danos-usb.img from the four boot
  artifacts (so changing -Dtest-case rebuilds the image with that kernel), a
  check-fat-image step runs --verify, and run-x86-64 boots the image on a
  usb-storage device (if=none,id=bootusb + usb-storage,bus=xhci.0,bootindex=0),
  keeping usb-kbd/usb-mouse on the same controller.

- test/qemu_test.py: the default boot config now boots off danos-usb.img on a
  usb-storage device (xHCI + usb-kbd + usb-mouse + the boot stick). The seven
  per-case qemu_extra blocks that added their own qemu-xhci/usb-kbd/usb-mouse
  (or a VVFAT stick) collided on id=xhci and are removed — the default provides
  the bus and the boot device. usb-storage and fat-mount now exercise the real
  FAT32 boot image (usb-storage reads its 0x55AA boot sector; fat mounts it at
  /mnt/usb). A build_case override lets a case reuse another's kernel, used by a
  new usb-boot case: an explicit, named boot-from-USB regression guard.

Verified: zig build, zig build test, and zig build check-fat-image are green
(FAT32, 128992 clusters, BOOTX64.efi present); a broad sequential QEMU sweep
passes — smoke, init, vfs, input, device-manager, usb-report, usb-hid,
usb-storage, fat-mount, device-list, driver-restart, acpi-report, iommu,
orderly-shutdown, usb-boot, dma, msi, initial-ramdisk, args, process — proving
the boot switch holds across kernel tests, the full init tree, the USB stack,
the FAT mount, and orderly shutdown.
2026-07-13 15:34:18 +01:00
Daniel Samson a64a01a6a9 M6: FAT read/write filesystem server, mounted into the VFS
Add a FAT12/16/32 filesystem the VFS mounts at /mnt/usb, reading and writing a
USB stick through the block device. Verified end to end under QEMU: the fat
server mounts the volume, the VFS routes /mnt/usb to it, and a client lists the
root and reads a file (the ELF magic of /mnt/usb/system/kernel).

- engine.zig: the FAT engine over a BlockDevice interface — mount (a bare FAT or,
  as QEMU's VVFAT and most real sticks present it, an MBR-partitioned disk), FAT
  chain walk (12/16/32), cluster allocation, directory traversal with long-name
  read, and file read / write / create. Host-tested against a RAM-backed FAT16
  image (create, cluster-spanning write, mid-file overwrite, read-back, list).
- on-disk.zig: the align(1) boot-sector / directory / long-name / FSInfo structs
  and the cluster-count FAT-type detection.
- fat.zig: the server — wraps the .block device (a DMA bounce buffer) in a
  BlockDevice, mounts the FAT, serves the vfs-protocol as a backend, and mounts
  itself into the VFS at /mnt/usb. Spawned by init as a boot service.
- runtime.block: the block-device client (geometry / read / write by physical
  address, so whole sectors never cross IPC).
- Raise the kernel service-name registry (maximum_services) 8 -> 16: it is
  indexed directly by ServiceId, and fat = 8 was being rejected, so the fat
  server exited before registering.
- VFS: an absolute path with no matching mount is now not-found rather than
  silently created in the flat ramfs — so /mnt/usb fails cleanly until mounted.

Tests: fat-mount (the full stack: block -> FAT -> VFS mount -> list + file read)
passes; host units cover the engine and on-disk structs; the vfs, shutdown, and
USB regression suite stays green (10/10).
2026-07-13 15:05:53 +01:00
Daniel Samson 35e8921de8 M5: VFS mount support — mount table + forwarding router
Turn the flat-ramfs VFS into a router: a mount table maps an absolute path prefix
(e.g. /mnt/usb) to a backend server's endpoint, and open/read/write/status/
readdir/close on a path under a mount are forwarded to that backend, which speaks
the same vfs-protocol. This is what a FAT filesystem mounts into.

- protocol: append readdir / mount / unmount operations, a NodeKind enum (the FSH
  file types) that now fills FileStatus.kind, a DirectoryEntry record, and a
  directory open flag. Appended values keep existing clients and tests unchanged.
- vfs.zig: a mount table, longest-prefix routing, forwarding of every op on a
  backend handle, mount/unmount handlers (the backend arrives as the call's
  capability), and release-on-death that also closes the backend's handles.
- path.zig: pure, host-tested mount-prefix matching that never captures a
  non-boundary like /mnt/usbextra.
- unistd: mount(), opendir / readdir / closedir clients.

Bare names still resolve in the flat ramfs — the backward-compat contract; the
vfs and vfs-client-death tests pass unchanged. End-to-end mount+read is exercised
by the FAT server (M6). Host units cover path matching and protocol sizes.
2026-07-13 14:21:30 +01:00
32 changed files with 3741 additions and 412 deletions
+75 -39
View File
@@ -58,7 +58,6 @@ fn addUserBinary(
b: *std.Build, b: *std.Build,
target: std.Build.ResolvedTarget, target: std.Build.ResolvedTarget,
runtime_module: *std.Build.Module, runtime_module: *std.Build.Module,
posix_module: *std.Build.Module,
mmio_module: *std.Build.Module, mmio_module: *std.Build.Module,
xkeyboard_config_module: *std.Build.Module, xkeyboard_config_module: *std.Build.Module,
acpi_ids_module: *std.Build.Module, acpi_ids_module: *std.Build.Module,
@@ -78,9 +77,6 @@ fn addUserBinary(
.stack_protector = false, .stack_protector = false,
.imports = &.{ .imports = &.{
.{ .name = "runtime", .module = runtime_module }, .{ .name = "runtime", .module = runtime_module },
// POSIX/C compatibility layer, available to any program that wants it
// (danos-native code uses `runtime` directly). See library/posix/.
.{ .name = "posix", .module = posix_module },
// Typed volatile MMIO + memory barriers, for drivers. See library/mmio/. // Typed volatile MMIO + memory barriers, for drivers. See library/mmio/.
.{ .name = "mmio", .module = mmio_module }, .{ .name = "mmio", .module = mmio_module },
// Keyboard layouts (keycode + modifiers -> keysym/character), available // Keyboard layouts (keycode + modifiers -> keysym/character), available
@@ -250,6 +246,8 @@ pub fn build(b: *std.Build) void {
// The USB transfer protocol, so runtime.usb (the class-driver client) can speak // The USB transfer protocol, so runtime.usb (the class-driver client) can speak
// it, the way runtime.input speaks the input protocol. // it, the way runtime.input speaks the input protocol.
runtime_module.addImport("usb-transfer-protocol", usb_transfer_protocol_module); runtime_module.addImport("usb-transfer-protocol", usb_transfer_protocol_module);
// The block protocol, so runtime.block (the block-device client) can speak it.
runtime_module.addImport("block-protocol", block_protocol_module);
// The power protocol: system power's domain-named surface (docs/power.md). // The power protocol: system power's domain-named surface (docs/power.md).
const power_protocol_module = b.addModule("power-protocol", .{ const power_protocol_module = b.addModule("power-protocol", .{
@@ -278,18 +276,6 @@ pub fn build(b: *std.Build) void {
}, },
}); });
// The POSIX / C compatibility layer, a separate library layered strictly over the
// runtime (it calls the runtime's IPC/heap, never system calls directly). This is
// the one place POSIX/C spellings are allowed verbatim — see docs/coding-standards.md
// and library/posix/posix.zig.
const posix_module = b.addModule("posix", .{
.root_source_file = b.path("library/posix/posix.zig"),
.imports = &.{
.{ .name = "runtime", .module = runtime_module },
.{ .name = "vfs-protocol", .module = vfs_protocol_module },
},
});
// The initial_ramdisk container format, shared by the kernel (unpacks it) and the // The initial_ramdisk container format, shared by the kernel (unpacks it) and the
// build-time packer tools/make-initial-ramdisk.py (produces it). No dependencies. // build-time packer tools/make-initial-ramdisk.py (produces it). No dependencies.
const initial_ramdisk_module = b.addModule("initial-ramdisk", .{ const initial_ramdisk_module = b.addModule("initial-ramdisk", .{
@@ -361,7 +347,7 @@ pub fn build(b: *std.Build) void {
// Built by the shared user-binary recipe (see addUserBinary): freestanding, // Built by the shared user-binary recipe (see addUserBinary): freestanding,
// linked into the kernel's user region against the `runtime` runtime library, and // linked into the kernel's user region against the `runtime` runtime library, and
// started in ring 3 by the kernel's user-ELF loader. // started in ring 3 by the kernel's user-ELF loader.
const init_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "init", "system/services/init/init.zig"); const init_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "init", "system/services/init/init.zig");
const init_install = b.addInstallArtifact(init_exe, .{ .dest_dir = .{ .override = .{ .custom = "system/services" } } }); const init_install = b.addInstallArtifact(init_exe, .{ .dest_dir = .{ .override = .{ .custom = "system/services" } } });
b.getInstallStep().dependOn(&init_install.step); b.getInstallStep().dependOn(&init_install.step);
@@ -369,12 +355,12 @@ pub fn build(b: *std.Build) void {
// Each is built by the same user-binary recipe, then packed into one image by // Each is built by the same user-binary recipe, then packed into one image by
// the host-side make-initial-ramdisk tool. The bootloader ferries the image to the kernel, // the host-side make-initial-ramdisk tool. The bootloader ferries the image to the kernel,
// which unpacks it and spawns each program (system/initial-ramdisk.zig). // which unpacks it and spawns each program (system/initial-ramdisk.zig).
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig"); const vfs_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig"); const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig"); const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig");
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig"); const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig"); const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig"); const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
// The xHCI bus driver builds chapter-9 requests and decodes descriptors from // The xHCI bus driver builds chapter-9 requests and decodes descriptors from
// usb-abi, and reports each interface's (class,subclass,protocol) identity via // usb-abi, and reports each interface's (class,subclass,protocol) identity via
// usb-ids.packTriple. // usb-ids.packTriple.
@@ -384,22 +370,26 @@ pub fn build(b: *std.Build) void {
// The USB HID class drivers: keyboard and mouse. They own no hardware — each // The USB HID class drivers: keyboard and mouse. They own no hardware — each
// opens its device through runtime.usb (the transfer protocol) and publishes to // opens its device through runtime.usb (the transfer protocol) and publishes to
// the input service. They build chapter-9 class requests from usb-abi. // the input service. They build chapter-9 class requests from usb-abi.
const usb_hid_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-keyboard", "system/drivers/usb-hid/keyboard.zig"); const usb_hid_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-keyboard", "system/drivers/usb-hid/keyboard.zig");
usb_hid_keyboard_exe.root_module.addImport("usb-abi", usb_abi_module); usb_hid_keyboard_exe.root_module.addImport("usb-abi", usb_abi_module);
const usb_hid_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-mouse", "system/drivers/usb-hid/mouse.zig"); const usb_hid_mouse_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-mouse", "system/drivers/usb-hid/mouse.zig");
usb_hid_mouse_exe.root_module.addImport("usb-abi", usb_abi_module); usb_hid_mouse_exe.root_module.addImport("usb-abi", usb_abi_module);
// The USB mass-storage class driver: opens its device via runtime.usb, drives it // The USB mass-storage class driver: opens its device via runtime.usb, drives it
// with Bulk-Only Transport + SCSI, and serves the block protocol under `.block`. // with Bulk-Only Transport + SCSI, and serves the block protocol under `.block`.
const usb_storage_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-storage", "system/drivers/usb-storage/usb-storage.zig"); const usb_storage_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-storage", "system/drivers/usb-storage/usb-storage.zig");
usb_storage_exe.root_module.addImport("block-protocol", block_protocol_module); usb_storage_exe.root_module.addImport("block-protocol", block_protocol_module);
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig"); // The FAT filesystem server: mounts the block device and serves it into the VFS
// at /mnt/usb. Its engine (engine.zig / on-disk.zig) is imported relatively.
const fat_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat", "system/services/fat/fat.zig");
const fat_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat-test", "system/services/fat/fat-test.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
// The PCI bus driver decodes each function's class triple to human names in its // The PCI bus driver decodes each function's class triple to human names in its
// boot log (class/subclass/prog-IF), so pull in the shared pci-class reference. // boot log (class/subclass/prog-IF), so pull in the shared pci-class reference.
pci_bus_exe.root_module.addImport("pci-class", pci_class_module); pci_bus_exe.root_module.addImport("pci-class", pci_class_module);
// A test fixture, not a real driver: hellos to the device manager, then faults — // A test fixture, not a real driver: hellos to the device manager, then faults —
// what the driver-restart scenario drives the crash-loop cap with. // what the driver-restart scenario drives the crash-loop cap with.
const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig"); const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig");
const device_list_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig"); const device_list_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig");
// The discovery service: one swappable process per firmware // The discovery service: one swappable process per firmware
// (docs/discovery.md), bundled under the neutral ramdisk name // (docs/discovery.md), bundled under the neutral ramdisk name
// "discovery" so the device manager never learns which firmware it is on. // "discovery" so the device manager never learns which firmware it is on.
@@ -413,9 +403,9 @@ pub fn build(b: *std.Build) void {
.acpi => "system/services/acpi/acpi.zig", .acpi => "system/services/acpi/acpi.zig",
.fdt => "system/services/fdt/fdt.zig", .fdt => "system/services/fdt/fdt.zig",
}; };
const discovery_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source); const discovery_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source);
if (discovery == .acpi) discovery_exe.root_module.addImport("aml", aml_module); if (discovery == .acpi) discovery_exe.root_module.addImport("aml", aml_module);
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig"); const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// Names the xHCI PCI class triple from the shared taxonomy instead of a bare 0x0C0330. // Names the xHCI PCI class triple from the shared taxonomy instead of a bare 0x0C0330.
device_manager_exe.root_module.addImport("pci-class", pci_class_module); device_manager_exe.root_module.addImport("pci-class", pci_class_module);
// The manager matches reported USB interfaces by their (class,subclass,protocol) // The manager matches reported USB interfaces by their (class,subclass,protocol)
@@ -423,11 +413,12 @@ pub fn build(b: *std.Build) void {
device_manager_exe.root_module.addImport("usb-ids", usb_ids_module); device_manager_exe.root_module.addImport("usb-ids", usb_ids_module);
// The input service and its exercisers: the fan-out server, a hardware-free synthetic // The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md. // source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
const input_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig"); const input_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig");
const input_source_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-source", "system/services/input-source/input-source.zig"); const input_source_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-source", "system/services/input-source/input-source.zig");
const input_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig"); const input_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig");
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig"); const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig");
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig"); const process_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig");
const log_flush_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "log-flush", "system/services/log-flush/log-flush.zig");
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool // Pack the user binaries into the initial_ramdisk image with the host-side Python tool
// (the container format is trivial, and Python sidesteps std API churn). Args: // (the container format is trivial, and Python sidesteps std API churn). Args:
@@ -453,6 +444,10 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(usb_hid_mouse_exe.getEmittedBin()); mk_run.addFileArg(usb_hid_mouse_exe.getEmittedBin());
mk_run.addArg("usb-storage"); mk_run.addArg("usb-storage");
mk_run.addFileArg(usb_storage_exe.getEmittedBin()); mk_run.addFileArg(usb_storage_exe.getEmittedBin());
mk_run.addArg("fat");
mk_run.addFileArg(fat_exe.getEmittedBin());
mk_run.addArg("fat-test");
mk_run.addFileArg(fat_test_exe.getEmittedBin());
mk_run.addArg("pci-bus"); mk_run.addArg("pci-bus");
mk_run.addFileArg(pci_bus_exe.getEmittedBin()); mk_run.addFileArg(pci_bus_exe.getEmittedBin());
mk_run.addArg("crash-test"); mk_run.addArg("crash-test");
@@ -473,6 +468,8 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(args_echo_exe.getEmittedBin()); mk_run.addFileArg(args_echo_exe.getEmittedBin());
mk_run.addArg("process-test"); mk_run.addArg("process-test");
mk_run.addFileArg(process_test_exe.getEmittedBin()); mk_run.addFileArg(process_test_exe.getEmittedBin());
mk_run.addArg("log-flush");
mk_run.addFileArg(log_flush_exe.getEmittedBin());
// Also install the packed binaries to their FHS homes, so zig-out is a true image // Also install the packed binaries to their FHS homes, so zig-out is a true image
// of the filesystem — even though at boot they arrive inside the initial-ramdisk. // of the filesystem — even though at boot they arrive inside the initial-ramdisk.
@@ -487,6 +484,8 @@ pub fn build(b: *std.Build) void {
.{ usb_hid_keyboard_exe, "system/drivers" }, .{ usb_hid_keyboard_exe, "system/drivers" },
.{ usb_hid_mouse_exe, "system/drivers" }, .{ usb_hid_mouse_exe, "system/drivers" },
.{ usb_storage_exe, "system/drivers" }, .{ usb_storage_exe, "system/drivers" },
.{ fat_exe, "system/services" },
.{ log_flush_exe, "system/services" },
}) |entry| { }) |entry| {
const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } }); const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } });
b.getInstallStep().dependOn(&step.step); b.getInstallStep().dependOn(&step.step);
@@ -520,6 +519,36 @@ pub fn build(b: *std.Build) void {
const efi_install = b.addInstallArtifact(efiexe, .{ .dest_dir = .{ .override = .{ .custom = "EFI/BOOT" } } }); const efi_install = b.addInstallArtifact(efiexe, .{ .dest_dir = .{ .override = .{ .custom = "EFI/BOOT" } } });
b.getInstallStep().dependOn(&efi_install.step); b.getInstallStep().dependOn(&efi_install.step);
// --- danos-usb.img: the bootable FAT32 USB image ---
// Format a real FAT32 image (the in-repo Python builder, no external tools)
// holding exactly what the firmware and bootloader need off the ESP: the EFI
// stub, the kernel, init, and the initial-ramdisk. QEMU presents this image as
// a USB mass-storage device the guest boots from (see run-x86-64 and the test
// harness), and the danos fat driver mounts the same image at /mnt/usb.
const mk_fat = b.addSystemCommand(&.{"python3"});
mk_fat.addFileArg(b.path("tools/make-fat-image.py"));
const fat_image = mk_fat.addOutputFileArg("danos-usb.img");
mk_fat.addArg("64"); // MiB
mk_fat.addArg("EFI/BOOT/BOOTX64.efi");
mk_fat.addFileArg(efiexe.getEmittedBin());
mk_fat.addArg("system/kernel");
mk_fat.addFileArg(exe.getEmittedBin());
mk_fat.addArg("system/services/init");
mk_fat.addFileArg(init_exe.getEmittedBin());
mk_fat.addArg("boot/initial-ramdisk.img");
mk_fat.addFileArg(initial_ramdisk_img);
const fat_image_install = b.addInstallFile(fat_image, "danos-usb.img");
b.getInstallStep().dependOn(&fat_image_install.step);
// `zig build check-fat-image` — validate the produced image is a real FAT32
// with the EFI stub present (the builder's own --verify, no external tools).
const check_fat = b.addSystemCommand(&.{"python3"});
check_fat.addFileArg(b.path("tools/make-fat-image.py"));
check_fat.addArg("--verify");
check_fat.addFileArg(fat_image);
const check_fat_step = b.step("check-fat-image", "Verify the FAT32 USB image is valid and bootable");
check_fat_step.dependOn(&check_fat.step);
// --- run-x86-64: boot the x86-64 kernel in QEMU via UEFI/OVMF --- // --- run-x86-64: boot the x86-64 kernel in QEMU via UEFI/OVMF ---
// Firmware lives in different places per OS/distro, so probe the known // Firmware lives in different places per OS/distro, so probe the known
// layouts (Architecture, Debian/Ubuntu, Fedora, macOS Homebrew) and use the first // layouts (Architecture, Debian/Ubuntu, Fedora, macOS Homebrew) and use the first
@@ -580,10 +609,13 @@ pub fn build(b: *std.Build) void {
}); });
run_efi.addArg("-drive"); run_efi.addArg("-drive");
run_efi.addPrefixedFileArg("if=pflash,format=raw,file=", vars_out); run_efi.addPrefixedFileArg("if=pflash,format=raw,file=", vars_out);
// Present the FHS zig-out to the guest as a FAT drive — it is the boot volume. // Boot off the FAT32 USB image: a mass-storage device on the same xHCI bus as
// the keyboard and mouse. OVMF finds \EFI\BOOT\BOOTX64.efi on it and boots.
run_efi.addArg("-drive");
run_efi.addPrefixedFileArg("if=none,id=bootusb,format=raw,file=", fat_image);
run_efi.addArgs(&.{ run_efi.addArgs(&.{
"-drive", "-device",
b.fmt("format=raw,file=fat:rw:{s}", .{b.install_path}), "usb-storage,bus=xhci.0,drive=bootusb,removable=on,bootindex=0",
"-net", "-net",
"none", "none",
// Emulated display advertising 1280x720 as its native (EDID preferred) // Emulated display advertising 1280x720 as its native (EDID preferred)
@@ -638,6 +670,10 @@ pub fn build(b: *std.Build) void {
"system/drivers/usb-hid/hid-report.zig", // HID boot-report keyboard/mouse decode "system/drivers/usb-hid/hid-report.zig", // HID boot-report keyboard/mouse decode
"system/drivers/usb-storage/bulk-only-transport.zig", // CBW/CSW wrapper sizes "system/drivers/usb-storage/bulk-only-transport.zig", // CBW/CSW wrapper sizes
"system/drivers/usb-storage/scsi.zig", // SCSI CDB encodings (big-endian) "system/drivers/usb-storage/scsi.zig", // SCSI CDB encodings (big-endian)
"system/services/vfs/path.zig", // mount-prefix path matching
"system/services/vfs/protocol.zig", // NodeKind / DirectoryEntry sizes + op values
"system/services/fat/on-disk.zig", // FAT on-disk struct sizes + type detection
"system/services/fat/engine.zig", // FAT read/write over a RAM-backed image
}) |root| { }) |root| {
const mod_tests = b.addTest(.{ const mod_tests = b.addTest(.{
.root_module = b.createModule(.{ .root_module = b.createModule(.{
+17 -11
View File
@@ -82,6 +82,12 @@ Start with the north star:
- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on - **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on
fault isolation + live restart — the reincarnation-server + capability model that fault isolation + live restart — the reincarnation-server + capability model that
makes "if I break it, I can restart it" real. danos's core motivation. makes "if I break it, I can restart it" real. danos's core motivation.
- **[zig-self-hosting.md](zig-self-hosting.md) — running Zig on danos.** A design note
(not built yet) on making danos a real Zig target (`-target x86_64-danos`) and
eventually running the compiler on it. The key realisation: Zig 0.16 reduces an OS
port to **one seam** (`std.os.danos`), so we build `runtime.os` (→ that seam) plus a
thin `runtime.fs`, retire the `posix` shim, and follow a phased path to
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
Cutting across all of these: Cutting across all of these:
@@ -191,21 +197,22 @@ system/ → /system danos's own internals (the self-representation)
services/ init/ vfs/ device-manager/ system servers → /system/services (vfs/ holds services/ init/ vfs/ device-manager/ system servers → /system/services (vfs/ holds
vfs.zig, vfs-test.zig, protocol.zig) vfs.zig, vfs-test.zig, protocol.zig)
library/ → /lib libraries, one sub-directory each library/ → /lib libraries, one sub-directory each
runtime/ the danos-native runtime — the stable application ABI runtime/ the danos-native runtime + file API (fs) — the stable application ABI
posix/ POSIX/C compatibility, layered over runtime
boot/ → /boot the loaders boot/ → /boot the loaders
tools/ test/ host-side build + QEMU test harness tools/ test/ host-side build + QEMU test harness
``` ```
A sub-project exposes its **public interface as a module**: `system/services/vfs/` owns A sub-project exposes its **public interface as a module**: `system/services/vfs/` owns
the VFS wire protocol (`protocol.zig`, the `vfs-protocol` module), which the POSIX the VFS wire protocol (`protocol.zig`, the `vfs-protocol` module), which the runtime's
layer imports by name. `usb`/`block` drivers will expose their protocols the same way. file API (`runtime.fs`) imports by name. `usb`/`block` drivers expose their protocols the
same way.
`library/posix/` is special: it is the **one place** POSIX/C spellings are allowed There is **no POSIX/C compatibility layer today**: danos programs do file I/O through the
verbatim (`stat`, `O_CREAT`, `fopen`, `errno`). Everywhere else follows the danos danos-native `runtime.fs` (open/read/write/list over the VFS). A hand-rolled POSIX shim
naming rule with no exception — see [coding-standards.md](coding-standards.md). The (`library/posix/`) was retired as premature — the real POSIX/C surface will come later
POSIX layer calls the runtime, never the kernel's system calls directly, so it never from the `std.os.danos` seam (and, eventually, musl) when danos becomes a Zig target (see
appears in the private-ABI path. [zig-self-hosting.md](zig-self-hosting.md)). When it does, the foreign-ABI naming
exception in [coding-standards.md](coding-standards.md) applies to that seam.
## Source map ## Source map
@@ -229,8 +236,7 @@ appears in the private-ABI path.
| Framebuffer text console (mirrors to serial) | `system/kernel/console.zig` | | Framebuffer text console (mirrors to serial) | `system/kernel/console.zig` |
| In-kernel test cases | `system/kernel/tests.zig` | | In-kernel test cases | `system/kernel/tests.zig` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/IO-APIC/timer, serial, linker script) | `system/kernel/architecture/x86_64/` | | Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/IO-APIC/timer, serial, linker script) | `system/kernel/architecture/x86_64/` |
| danos-native runtime (`runtime`): syscall wrappers, heap, IPC, device access — the stable application ABI | `library/runtime/` | | danos-native runtime (`runtime`): syscall wrappers, heap, IPC, device access, the file API (`fs`) — the stable application ABI | `library/runtime/` |
| POSIX/C compatibility (`posix`): unistd, stdio — the one place POSIX names are allowed | `library/posix/` |
| System services (init, the VFS server + `protocol`, the device-manager) | `system/services/` | | System services (init, the VFS server + `protocol`, the device-manager) | `system/services/` |
| Device drivers, one sub-project each (`pci-bus`, `ps2-bus`, `usb-xhci-bus` bus drivers) | `system/drivers/` | | Device drivers, one sub-project each (`pci-bus`, `ps2-bus`, `usb-xhci-bus` bus drivers) | `system/drivers/` |
| Build + `run-x86-64` (QEMU/OVMF) | `build.zig` | | Build + `run-x86-64` (QEMU/OVMF) | `build.zig` |
+12 -9
View File
@@ -65,15 +65,18 @@ Three, and only three.
`errno`, `O_CREAT`. We don't get to rename `fwrite` to `fileWrite` — it wouldn't be `errno`, `O_CREAT`. We don't get to rename `fwrite` to `fileWrite` — it wouldn't be
`fwrite` any more. `fwrite` any more.
**This exception is scoped to one place: `library/posix/`.** A file under **This exception is scoped to a file that *is* a foreign ABI, and nothing else.**
`library/posix/` *is* the foreign ABI, so it keeps the ABI's spellings — that is the danos has no such file today: the old `library/posix/` compatibility shim was retired
whole rule for that directory. **Everywhere else, Zig/danos naming applies with no once its callers moved to the danos-native `runtime.fs`, since a hand-rolled POSIX
POSIX exception**, so there is nothing to get wrong: if you're not in layer is premature until danos actually needs it (see
`library/posix/`, expand it. A concept POSIX also has gets a danos name outside that [zig-self-hosting.md](zig-self-hosting.md)). The exception will apply again to the
layer — the VFS wire protocol carries a `FileStatus`, not a `Stat`, and a `create` `std.os.danos` seam when danos becomes a real Zig target — that module *is* the C-ABI
flag, not `O_CREAT`; `library/posix/` is what maps `stat`→`status` and `system` interface, so it keeps `open`/`read`/`errno`/`O_CREAT`. **Everywhere else,
`O_CREAT`→`create` at the boundary. (The `syscall` *wrappers* elsewhere are not an Zig/danos naming applies with no exception**: a concept POSIX also has gets a danos
exception to this — they wrap the private danos ABI, so they use danos names.) name — the VFS wire protocol carries a `FileStatus`, not a `Stat`, and a `create`
flag, not `O_CREAT`; the boundary is where `stat`→`status` and `O_CREAT`→`create` get
mapped. (The `syscall` *wrappers* elsewhere are not an exception — they wrap the
private danos ABI, so they use danos names.)
2. **Zig idioms are spelled the way Zig spells them.** Three names are the language's, 2. **Zig idioms are spelled the way Zig spells them.** Three names are the language's,
not ours, and are left alone: not ours, and are left alone:
+6 -6
View File
@@ -17,12 +17,12 @@ was a mistake) without inheriting the mechanism, the API, or the names. The nami
rule is danos's own and it is strict: plain words that communicate intent rule is danos's own and it is strict: plain words that communicate intent
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks (`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for (`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
a concept that already has one. Literal POSIX arrives later and lives elsewhere: a a concept that already has one. Literal POSIX arrives later and lives elsewhere: the
**musl-based C layer** (growing out of library/posix) that wires C programs to the `std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C
danos runtime — musl's syscall surface retargeted at danos system calls and IPC layer** on the same native surface (see [zig-self-hosting.md](zig-self-hosting.md)) —
protocols (files onto the VFS protocol, `sigaction`/`wait` onto this lifecycle, musl's syscall surface retargeted at danos system calls and IPC protocols (files onto
sockets onto whatever networking becomes). Ported programs see POSIX; the system the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever
underneath never does. networking becomes). Ported programs see POSIX; the system underneath never does.
## Why a standard vocabulary ## Why a standard vocabulary
+349
View File
@@ -0,0 +1,349 @@
# Running Zig on danos: the self-hosting roadmap
A design note (not built yet) on the path to making danos a **real Zig target** — a
target you can name (`-target x86_64-danos`) and, eventually, run the Zig compiler
itself on. It is forward-looking, like [vision.md](vision.md): it sets a direction
and the decisions that follow from it, so the code we write now bends toward it
instead of away.
This note deliberately does **not** cover a text editor or terminal. Those are
easier (single-process, I/O-bound) and fall out of the early phases here almost for
free; the hard, shaping problem is the standard-library surface, so that is what
this roadmap is about.
The analysis behind it was done against **Zig 0.16** (the pinned toolchain). Zig's
standard library moves between releases — especially the parts described here — so
treat upstream references as "the shape in 0.16.x," and expect to re-check them on a
toolchain bump.
## The win condition
danos runs the Zig compiler when a bare
```
zig build-exe hello.zig
```
completes **on danos** and produces a runnable danos binary. Note the milestone is
`build-exe`, not `zig build`: the `zig build` runner spawns child processes (the
build steps), which needs a whole process-control surface danos does not have yet.
A single `build-exe` needs none of that (see Phase 3). Reaching `build-exe` is
"self-hosting"; reaching `zig build` is a later, separate lift.
### Non-goals
- **No Linux syscall/ABI emulation.** danos will not implement the Linux `syscall`
interface so that stock `x86_64-linux` binaries run. That is a permanent
compatibility treadmill and it inverts the microkernel design — explicitly out.
- **No musl port yet.** A musl libc port is a reasonable *later* effort (it unlocks
the C ecosystem), but it is not on the critical path to Zig-on-danos, and it is
deferred. The roadmap below is arranged so the work still pays off if musl ever
happens (see "The same surface, twice").
- **Editor/terminal are out of scope for this note** (they are downstream of Phase 1).
**On FFI.** Foreign-function interop splits the same way as the doors below. Zig-level
and C-ABI-*exposing* FFI (`extern`, `callconv(.c)`, C-ABI structs) work on a real target
immediately — and the `std.os.danos` seam is C-ABI-shaped by construction, so it is
FFI-friendly from the start. *Consuming* C libraries (`@cImport`, linking archives) is
the part that needs a libc + headers, i.e. the deferred musl door. So an eventual FFI
need reinforces keeping that door open; it does not change the plan.
## The realization that shapes everything: 0.16 gives us *one* seam
The instinct "to target Zig we'd have to reimplement all the `std` namespaces" was
how older Zig worked. Zig 0.16 (post-"writergate") is far kinder:
- **`std.fs` is essentially gone.** It is now path helpers plus deprecated aliases;
there is no `std.fs.File`, `std.fs.Dir`, or `std.fs.cwd()`. File and directory
work goes through **`std.Io`** — a single runtime **vtable** (`Io.zig`) of
function pointers handed to `main` as `std.process.Init.io`. `std.Io.File` and
`std.Io.Dir` are thin forwarders to that vtable. `Io.zig` and the `fs` shim carry
**zero** per-OS branches.
- **`std.posix` is one generic body** parameterised over a single `system` module.
With no libc, `system` resolves **per target OS**: `.linux => std.os.linux`,
`.plan9 => std.os.plan9`, and so on. The generic `std.posix.read`/`write`/`open`
bodies are just `system.read(...)` plus an errno switch — *identical for every
OS*. The only variable is what `system` binds to.
- **`std.os.<tag>`** (e.g. `std/os/linux.zig`) is therefore the real porting seam: a
low-level, C-ABI-shaped module of `read/write/open/close/lseek/mmap/clock/exit/…`
plus an `errno` enum and the constant tables (`O_*`, `CLOCK_*`, `S_*`).
Put together: **to port danos we write `std.os.danos` once** — the ~30-operation
seam — and the whole `std.posix` / `std.fs` / `std.Io` tower above it lights up
generically, because none of it branches on the OS. That is a dramatically smaller
and more contained target than "reimplement the namespaces."
## Three doors, and why we take the first
| Door | What it is | Verdict |
|------|-----------|---------|
| **1. Implement the std seam** (`std.os.danos`) | Write the ~30-op `system` module over danos's native ABI + VFS; the generic std tower lights up. | **Take this.** The only door that touches neither C nor the Linux ABI. |
| **2. Port musl** | Port musl libc to danos, link Zig against it. | Defer. Good later for the *C* ecosystem; barely helps *Zig* (std only uses libc on the libc-linked path). |
| **3. Emulate the Linux ABI** | Implement Linux syscalls so stock linux binaries run. | Reject. Bottomless compatibility treadmill; against the design. |
### The same surface, twice
Doors 1 and 2 are the **same native surface at different layers**. `std.posix.read`
is `system.read(...)` + an errno switch *regardless of OS* — the only question is
whether `system` is **`std.os.danos` (Zig)** or **musl (C)**. Either way, the set of
danos-facing operations you must implement is the *same* ~30 ops, all bottoming out
in danos's native syscalls + the VFS/FAT server.
So the runtime work below is **not throwaway** if musl ever happens: you are building
the danos-native implementations of that surface either way. Door 1 just packages
them as Zig; a future musl re-uses the identical kernel/VFS operations underneath. The
two symmetries worth keeping in mind: doors 1 and 2 converge at the **top** (identical
POSIX surface); doors 2 and 3 converge at the **bottom** (unmodified musl needs the
Linux syscall ABI). Door 1 is the only one that avoids both C and Linux.
### A fork is table stakes — for any door
`std.Target.Os.Tag` is a **closed enum** baked into the compiler binary *and* into
the `std` linked with every program; `-target x86_64-danos` resolves through it. So
adding `danos` as a name requires patching and rebuilding the compiler — even the
musl door needs this. "Fork Zig" is therefore not an extra cost unique to door 1; it
is the price of admission for *any* real target. What door 1 adds on top is small and
localised (below).
## The architecture decision: `runtime.os` + `runtime.fs`, and retire `posix`
danos already has the right split ([the private-ABI boundary](../README.md)): the
kernel exposes a minimal syscall ABI ([syscall.md](syscall.md)); the **`runtime`**
library is the stable, danos-native application ABI. What this roadmap adds:
- **`runtime.os` — the seam.** A C-ABI-shaped module of the ~30 operations
(`read/write/open/close/lseek/mmap/munmap/clock/exit/…`) + an errno enum + the
constant tables, each backed by danos's native syscalls and the VFS. **Structure it
to mirror `std/os/linux.zig`.** This is the load-bearing, *non-throwaway* artifact:
when we fork Zig, `runtime.os` is copy-pasted (near-verbatim) into `std.os.danos`.
- **`runtime.fs` — the thin native file API** danos programs use *today*, layered
over `runtime.os`. It is also the concrete backing for the `std.Io` vtable's
file-write entry once we're a real target, which is why program stdout, diagnostics,
and file writes should all be *decided once at that seam* rather than as bespoke
per-call helpers (see "How this informs decisions now").
**Do not hand-mirror the high-level std namespaces.** `std.fs`/`std.Io`/`std.process`
are generic and OS-agnostic; once `std.os.danos` exists and we fork, upstream *gives*
them to danos for free. Hand-writing `runtime.std.fs` to imitate them would be
redundant the day the fork works, and it would chase a moving target (0.16's `std.Io`
is large and still shifting). Build the seam well; take the tower for free.
**Why not a library called `std`?** Because `@import("std")` resolves to the
compiler-provided standard library; a user module named `std` would *shadow* it for
anything that imports it that way. That is the real reason the seam lives *inside* a
forked std as `std/os/danos.zig`, not as a `runtime.std` library — and why danos's end
state (`@import("std")` just working, and knowing danos) is the most natively Zig it can
be. `runtime.os` is only the interim staging ground: developed against the stock
toolchain so Phase 1 need not wait on the fork, then promoted near-verbatim into the
fork's `std/os/danos.zig`.
### Retire `library/posix`
The `posix` compatibility layer (`unistd`, `stdio`) was the right instinct too early.
Its whole value is POSIX *spellings* for POSIX software — and danos has no POSIX
software; every current caller is danos-native code that could use `runtime.fs`
directly. The real POSIX story arrives later and from elsewhere (musl, or upstream
`std`'s own posix over `std.os.danos`), which supersedes a hand-rolled shim. So it is
premature abstraction that adds a "which layer do I use?" fork with no payoff yet.
Its footprint is tiny: **five** call sites, all `unistd` file operations —
`system/services/fat/fat.zig` (`mount`), the `vfs-test` and `fat-test` clients, and
(from the boot-log work) `init.zig` and `log-flush.zig`. `stdio.zig` is dead — nothing
imports it. The plan: build `runtime.fs`, migrate those five to it, delete
`library/posix/`, and drop the `posix` module from `build.zig`'s `addUserBinary`.
## Where danos stands: coverage vs. the gaps
What the seam needs, and what danos already provides:
| std need | danos today | Gap |
|----------|-------------|-----|
| open / read / write / close / lseek | VFS (via the current `unistd`, → `runtime.fs`) | none — repackage |
| directory read (`getdents`) | VFS `readdir` | none — repackage |
| mmap / munmap | native syscalls ([abi.zig](../system/abi.zig)) | none |
| page allocator | over `mmap`, via `root.os.heap.page_allocator` override | ~30-line hook |
| monotonic clock | `clock` syscall | none |
| args / argv | SysV entry stack ([sysv.md](sysv.md)), `runtime.process.Init` | none |
| stdout / stderr | `debug_write` today | wire fd 1/2 to a console **byte** stream |
| mkdir / unlink / rename / truncate | done — engine + VFS + `runtime.fs` (Phase 2) | — |
| stat fields | `{size, kind, mtime}` | **mode / inode** still missing (cache validity) |
| wall-clock / realtime | done — `wall_clock` syscall (CMOS RTC, Phase 2d) | — |
| **environment variables** | `Init` has no env field | missing (can start empty) |
| **cwd / chdir** | paths are absolute or bare | missing (no cwd anchor) |
| **entropy / random** | — | missing (needed behind `vtable.random`) |
| process spawn + exit status | `system_spawn` starts a *named ramdisk binary*; `ExitReason` is a *category* | no exec-of-path, no numeric `WEXITSTATUS` |
| threads | one thread per process | avoided via `-fsingle-threaded` (below) |
| symlinks | `NodeKind` has the tag; unimplemented | low priority |
The clustering is clear: reads and memory are basically done; the real work is
**filesystem mutation + richer stat + wall-clock**, and a few small seam pieces
(page-allocator hook, stdio bytes, entropy). Process spawning and threads are
side-stepped entirely for a single `build-exe`.
## The roadmap
### Phase 0 — Make `danos` a real target
**Host, target, self-host — keep the three roles straight.** The *host* is where the
compiler runs (your mac + linux dev machines); the *target* is what it emits (`danos`);
and eventually danos becomes a host too (self-hosting — the win condition). So the move
is: fork the compiler, build it **for** your dev hosts, and teach it to **cross-compile
to** danos. You already do this — danos is cross-compiled `freestanding` from your dev
host today; Phase 0 swaps that `freestanding` target for a real `x86_64-danos` one, which
is what unlocks the native `std`.
**Why a compiler fork, not just a `--zig-lib-dir` override.** `std.Target.Os.Tag` is a
*closed enum compiled into the compiler binary*, so `-target x86_64-danos` will not even
parse unless the compiler itself knows the tag. Overriding the std lib directory alone
cannot add a target — and there is no libc-only shortcut (a future musl needs the same
patch). The only alternative, staying on `freestanding` + hand-shims, is exactly the
non-native feel we are leaving: `@import("std")` there is stubbed, not real.
**The fork.** Clone `ziglang/zig` at the pinned 0.16 tag; build it with a stock
same-version `zig` (`zig build` in the tree — a standard, LLVM-pulling, roughly one-time
build); point danos's `build.zig`/CI at the resulting binary. Four localised patches:
- add `danos` to `std.Target.Os.Tag`, in the "no version range" group alongside
plan9/serenity;
- add `danos` to the freestanding/other **no-op `_start` list** in `std`'s `start.zig`,
so std does *not* emit its own System-V `_start` — danos keeps owning the entry shim
and `Init`/argv construction it already builds ([sysv.md](sysv.md));
- wire the `system` selector `.danos => std.os.danos` in `std.posix`;
- add `std/os/danos.zig` — **the seam itself**, promoted near-verbatim from the
`runtime.os` developed first in Phase 1 (against the stock toolchain, so the fork is
not a prerequisite for starting).
This is the fork treadmill we accept once. Keep the patch set tiny and `else`-friendly,
pin to one 0.16.x, and rebase on point releases.
### Phase 1 — `runtime.os` read-side + allocator + stdio + cwd; retire `posix`
Author `runtime.os` (→ `std.os.danos`): the `errno` enum, the constant tables, and
the C-convention `read / write / open / openat / close / lseek / mmap / munmap /
exit`, each returning result-or-`-errno`. Most backing already exists (VFS + native
mmap + clock).
- Provide `page_allocator` via `root.os.heap.page_allocator` (a thin override over
danos `mmap`). This sits **outside** the `std.Io` vtable, so it is wired separately.
- Wire fd 0/1/2 to a console **byte** stream (today output only reaches `debug_write`;
input is structured `InputEvent` IPC — a byte tty is a new, small thing in both
directions).
- Add a `getcwd`/`chdir` anchor so `std.fs.cwd()`-style resolution has something to
resolve against.
- Build `runtime.fs` over `runtime.os`; migrate the five `posix` callers to it; delete
`library/posix/` and drop its build module.
After Phase 1, the surface an editor or terminal needs (open/read/write/close/lseek/
readdir/isatty/args/exit) exists. Those are downstream and out of scope here.
### Phase 2 — Filesystem mutation + real stat (the compiler's cache tower)
danos's biggest genuine gap, and the correctness-critical one:
- Add **mkdir / unlink / rename / truncate** to *both* the VFS wire protocol
([protocol.zig](../system/services/vfs/protocol.zig)) and the FAT engine
([engine.zig](../system/services/fat/engine.zig)), then expose them via `runtime.os`.
- Extend `stat` beyond `{size, kind}` to carry **mtime + inode + mode** — `std`'s file
stat needs them for build-cache validity — which in turn needs **wall-clock** time
(danos is monotonic-only today; an RTC/time service is the dependency).
Because `std.fs`/`std.Io` have no per-OS branches, finishing this in `runtime.os`
lights up the whole file tower for the compiler at once. Environment can stay an empty
map until the kernel populates a non-empty `envp`.
**Status — Phase 2 complete.** `truncate` (O_TRUNC, closing the boot-log stale-tail
bug), `mkdir`, `unlink`, and `rename` are all wired through the FAT engine, the VFS
protocol + router, and `runtime.fs` (`makeDirectory` / `remove` / `rename`) —
host-tested and QEMU-tested (`fat-mutations` + `fat-rename` make a directory, write+read
a file in it, rename it, then remove it through the mount). `removeFile` and `rename`
are LFN-aware; `rename` is same-directory + 8.3 (cross-directory and long-name-
preserving rename are noted limitations). Wall-clock is now a kernel syscall
(`wall_clock`, a CMOS-RTC read anchored to the monotonic clock), and the FAT engine
stamps and reports **mtime** — `stat` / `runtime.fs.Attributes` carry a real
modification time (the `fat-mtime` case reads it back within seconds of the host clock).
The remaining `stat` fields, `mode`/`inode`, are deferred (not needed until the
compiler's cache layer wants them). **Everything past here is gated on Phase 0 (the
fork):** the `runtime.os` seam, `cwd`, stdio-as-fds, and the compiler bring-up.
### Phase 3 — Single-threaded, self-linked compiler bring-up
Build the compiler with **two load-bearing flags**:
- **`-fsingle-threaded`** removes `std.Thread` entirely — `Thread.spawn` is a hard
compile error under it, and `std.Io`'s threaded backend runs inline. danos being
one-thread-per-process is therefore **not** a blocker. Parallel codegen is a
throughput optimisation, not a correctness requirement.
- **`-fno-llvm -fno-lld`** keeps codegen and linking **in-process** (the self-hosted
x86-64 backend + self-linker), so a single `build-exe` **never forks a child**. That
is what lets us defer the entire spawn/exec/wait surface.
Then supply the few remaining seam pieces: `now` (wrap the danos clock), an entropy
source behind `vtable.random` (`randomSecure` can alias it initially — low volume, for
temp-file names and hashmap seeds), and the Phase-2 mkdir/rename/unlink for cache dir
trees and atomic temp-then-rename output.
**Explicitly deferred** (not on the `build-exe` path): child-process spawn/exec (only
`zig build` and external tools need it), `std.Thread`, `fsync` (FAT is write-through
today), symlinks, and musl.
## Risks and gotchas
- **The std-fork rebase treadmill is the main ongoing cost.** A new OS tag touches the
same broad file set plan9/serenity touch (hundreds of `native_os` sites, plus
"unsupported OS" `@compileError` dead-ends a new tag must be routed around), and the
entire `std.Io` layer is new in 0.16 and still moving. Stay pinned to one 0.16.x,
keep additions localised and `else`-friendly. Watch the closed-enum gotcha: adding
`danos` to `Os.Tag` can break existing *exhaustive* switches that lack an `else`, so
expect to touch switch sites beyond the ones you implement.
- **Single-threaded is load-bearing.** The "no `std.Thread`" simplification rests
entirely on `-fsingle-threaded`. If a dependency or flag flips threading back on, you
inherit an unescapable compile error (no root-hook exists) — the only outs are a full
thread-impl fork or linking libc for pthreads. Keep `single_threaded` asserted end to
end.
- **In-process linking is load-bearing.** Reaching the compiler without fork/exec
depends on `-fno-llvm -fno-lld`. The moment you shell out to LLD/`ld`, you need the
full `spawn`/`wait` surface — the hardest microkernel piece — and danos's
`system_spawn` only starts a *named ramdisk binary*, not exec of an arbitrary path.
Verify the self-hosted backend covers the target output before assuming child
processes are optional.
- **The shim cannot host the compiler.** danos's current `runtime`/`posix` is fine for
danos's *own* native programs, but the compiler `import`s *upstream* `std`, which on
a non-target hits the void `system` stub. So the compiler forces the real target
(Phase 0's fork). Do not over-invest in extending the hand-shim for compiler
purposes; put that effort into `runtime.os` + the VFS/FAT operations, which both the
fork *and* a future musl consume.
- **`"w"`/`O_CREAT` does not truncate — a silent-corruption bug on this road.** The FAT
engine's `writeFile` only *grows* `node.size`, so overwriting a shorter file leaves
trailing garbage. Harmless for the boot log today, but for a compiler it means
**corrupt `.o`/cache files that look like nondeterministic compiler bugs.** Land
`truncate` (Phase 2) before the compiler ever writes cache.
- **Exit status is categorical, not numeric.** `process_exit_reason` returns an
`ExitReason` *category*, not a numeric code (`WEXITSTATUS`). Fine while spawn is
stubbed; the day `zig build` or external tools arrive, plan a kernel exit-record
extension — do not let it surprise you.
## How this informs decisions now
Two current decisions fall out of this roadmap:
1. **The `runtime.fs` / `std.Io` question resolves at the vtable seam.** Because 0.16
routes *all* output through the `std.Io` vtable's file-write entry, and stdout/stderr
are just `File`s with well-known handles, build `runtime.fs` (and the console stdout)
as the concrete backing for that entry — not as a bespoke `std.Io.Writer`-only shim.
Decide it once, at the seam, and program stdout, diagnostics, and file writes all
flow through the same danos VFS/console path.
2. **The boot-log `truncate` caveat is now fixed** (Phase 2a). It was the same
`writeFile`-only-grows gap that on the self-hosting road would corrupt build output;
`engine.truncate` + an O_TRUNC open flag now free the old chain so a shorter rewrite
leaves no stale tail, and the boot-log flush opens with it.
## Related
- [vision.md](vision.md) — the north star this serves.
- [syscall.md](syscall.md) — the kernel↔runtime ABI `runtime.os` is built on.
- [sysv.md](sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs.
- [ipc.md](ipc.md) — the IPC the VFS/FAT operations travel over.
- [danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md) — the
filesystem layout the file surface serves.
- [coding-standards.md](coding-standards.md) — danos naming (why the compat spellings
are confined, and now retired).
-13
View File
@@ -1,13 +0,0 @@
//! DanOS's POSIX / C compatibility layer — `unistd`, `stdio`, and (later) the C
//! `errno` / `struct stat` / `extern "C"` surface. This is the *one* place POSIX and
//! C spellings are allowed to appear verbatim (see docs/coding-standards.md): a file
//! under library/posix/ *is* the foreign ABI, so it keeps the ABI's names. Everything
//! it touches on the danos side (the VFS protocol, the runtime) uses danos names,
//! which this layer translates to at the boundary.
//!
//! It is layered strictly *over* the runtime: it calls the runtime's IPC and heap,
//! never the kernel's system calls directly. danos-native applications use the
//! runtime; this exists so *POSIX* software can too.
pub const unistd = @import("unistd.zig");
pub const stdio = @import("stdio.zig");
-115
View File
@@ -1,115 +0,0 @@
//! A small C stdio layer over the POSIX-style file API (unistd.zig). Unbuffered
//! for now — each fread/fwrite is one VFS round trip; an internal buffer (fewer
//! IPC calls) is a later optimisation. Both a Zig-callable API and `extern "C"`
//! symbols are provided, so Zig and future C programs share it.
const std = @import("std");
const unistd = @import("unistd.zig");
const heap = @import("runtime").heap;
pub const SEEK_SET = unistd.SEEK_SET;
pub const SEEK_CURRENT = unistd.SEEK_CURRENT;
pub const SEEK_END = unistd.SEEK_END;
/// A C `FILE`: an fd plus sticky end-of-file / error flags. Allocated on the
/// heap; `fclose` frees it.
pub const FILE = extern struct {
fd: i32,
eof: c_int = 0,
err: c_int = 0,
};
fn flagsFor(mode: []const u8) u32 {
if (mode.len == 0) return 0;
return switch (mode[0]) {
'w', 'a' => unistd.O_CREAT,
else => 0,
};
}
/// Open `path` in `mode` ("r"/"w"/"a", '+' ignored for now). Returns null on error.
pub fn fopen(path: []const u8, mode: []const u8) ?*FILE {
const fd = unistd.open(path, flagsFor(mode));
if (fd < 0) return null;
const f = heap.allocator().create(FILE) catch {
unistd.close(fd);
return null;
};
f.* = .{ .fd = fd };
if (mode.len > 0 and mode[0] == 'a') _ = unistd.lseek(fd, 0, unistd.SEEK_END);
return f;
}
pub fn fclose(f: *FILE) c_int {
unistd.close(f.fd);
heap.allocator().destroy(f);
return 0;
}
/// Read `size*nmemb` bytes; returns the number of whole items read.
pub fn fread(buffer: []u8, size: usize, nmemb: usize, f: *FILE) usize {
const total = size * nmemb;
if (total == 0) return 0;
const n = unistd.read(f.fd, buffer[0..@min(buffer.len, total)]);
if (n <= 0) {
f.eof = 1;
return 0;
}
return @as(usize, @intCast(n)) / size;
}
/// Write `size*nmemb` bytes; returns the number of whole items written.
pub fn fwrite(data: []const u8, size: usize, nmemb: usize, f: *FILE) usize {
const total = @min(data.len, size * nmemb);
if (total == 0) return 0;
const n = unistd.write(f.fd, data[0..total]);
if (n <= 0) {
f.err = 1;
return 0;
}
return @as(usize, @intCast(n)) / size;
}
pub fn fseek(f: *FILE, off: i64, whence: u32) c_int {
f.eof = 0;
return if (unistd.lseek(f.fd, off, whence) < 0) -1 else 0;
}
pub fn ftell(f: *FILE) i64 {
return unistd.lseek(f.fd, 0, unistd.SEEK_CURRENT);
}
pub fn rewind(f: *FILE) void {
_ = fseek(f, 0, SEEK_SET);
}
pub fn feof(f: *FILE) c_int {
return f.eof;
}
pub fn ferror(f: *FILE) c_int {
return f.err;
}
pub fn fputs(s: []const u8, f: *FILE) c_int {
return if (unistd.write(f.fd, s) < 0) -1 else 0;
}
pub fn fputc(c: u8, f: *FILE) c_int {
const b = [_]u8{c};
return if (unistd.write(f.fd, &b) == 1) c else -1;
}
pub fn fgetc(f: *FILE) c_int {
var b: [1]u8 = undefined;
const n = unistd.read(f.fd, &b);
if (n <= 0) {
f.eof = 1;
return -1; // EOF
}
return b[0];
}
// Real `extern "C"` symbols (fopen/fread/fseek/...) — with a C-string signature
// distinct from the Zig slice API above — land with the first C program, wired
// via @export so they don't collide with these Zig names.
-148
View File
@@ -1,148 +0,0 @@
//! POSIX-style file API for user programs — the low level under C stdio. Files
//! are named objects served by the user-space VFS server (system/services/vfs/vfs.zig); each
//! call marshals a request, IPC_Calls the VFS, and unmarshals the reply. The
//! kernel knows nothing of files or fds — the fd table lives here, per process.
const std = @import("std");
const protocol = @import("vfs-protocol");
const ipc = @import("runtime").ipc;
pub const O_CREAT = protocol.create;
pub const SEEK_SET: u32 = 0;
pub const SEEK_CURRENT: u32 = 1;
pub const SEEK_END: u32 = 2;
// Resolve (and cache) the VFS server endpoint, looked up by well-known id.
var vfs_handle: usize = 0;
var vfs_resolved = false;
fn vfs() ?usize {
if (!vfs_resolved) {
vfs_handle = ipc.lookup(.vfs) orelse return null;
vfs_resolved = true;
}
return vfs_handle;
}
const maximum_fds = 32;
const Fd = struct { used: bool = false, node: u64 = 0, offset: u64 = 0 };
var fds = [_]Fd{.{}} ** maximum_fds;
fn allocFd() ?usize {
for (&fds, 0..) |*f, i| {
if (!f.used) {
f.* = .{ .used = true };
return i;
}
}
return null;
}
const Result = struct { reply: protocol.Reply, payload: []u8 };
/// One request/reply round trip: [Request header][send payload] -> VFS ->
/// [Reply header][receive payload]. The receive payload is written into `out`.
fn transact(request: protocol.Request, send: []const u8, out: []u8) ?Result {
const h = vfs() orelse return null;
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const slen = @min(send.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..slen], send[0..slen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(h, message[0 .. protocol.request_size + slen], &rbuf) catch return null;
if (n < protocol.reply_size) return null;
const reply = std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]);
const rpl = @min(n - protocol.reply_size, out.len);
@memcpy(out[0..rpl], rbuf[protocol.reply_size..][0..rpl]);
return .{ .reply = reply, .payload = out[0..rpl] };
}
/// Open (or create, with O_CREAT) `path`; returns an fd or -1.
pub fn open(path: []const u8, flags: u32) i32 {
const fd = allocFd() orelse return -1;
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = flags };
const r = transact(request, path, &.{}) orelse {
fds[fd].used = false;
return -1;
};
if (r.reply.status != 0) {
fds[fd].used = false;
return -1;
}
fds[fd] = .{ .used = true, .node = r.reply.node, .offset = 0 };
return @intCast(fd);
}
fn fdPtr(fd: i32) ?*Fd {
if (fd < 0 or fd >= maximum_fds) return null;
const f = &fds[@intCast(fd)];
return if (f.used) f else null;
}
/// Read up to `buffer.len` bytes at the current offset; returns the count or -1.
pub fn read(fd: i32, buffer: []u8) isize {
const f = fdPtr(fd) orelse return -1;
const want: u32 = @intCast(@min(buffer.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .read, .node = f.node, .offset = f.offset, .len = want, .flags = 0 };
const r = transact(request, &.{}, buffer) orelse return -1;
if (r.reply.status != 0) return -1;
f.offset += r.reply.len;
return @intCast(r.reply.len);
}
/// Write `data` at the current offset; returns the count or -1.
pub fn write(fd: i32, data: []const u8) isize {
const f = fdPtr(fd) orelse return -1;
const want: u32 = @intCast(@min(data.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .write, .node = f.node, .offset = f.offset, .len = want, .flags = 0 };
const r = transact(request, data[0..want], &.{}) orelse return -1;
if (r.reply.status != 0) return -1;
f.offset += r.reply.len;
return @intCast(r.reply.len);
}
/// Reposition the fd's offset. Returns the new offset or -1. (SEEK_END needs the
/// file size, which `stat` provides; handled by fetching it here.)
pub fn lseek(fd: i32, off: i64, whence: u32) i64 {
const f = fdPtr(fd) orelse return -1;
const base: i64 = switch (whence) {
SEEK_SET => 0,
SEEK_CURRENT => @intCast(f.offset),
SEEK_END => blk: {
const request = protocol.Request{ .operation = .status, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
var sbuf: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &sbuf) orelse return -1;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return -1;
const st = std.mem.bytesToValue(protocol.FileStatus, sbuf[0..@sizeOf(protocol.FileStatus)]);
break :blk @intCast(st.size);
},
else => return -1,
};
const pos = base + off;
if (pos < 0) return -1;
f.offset = @intCast(pos);
return pos;
}
/// Stat `path`. Returns 0 or -1.
pub fn stat(path: []const u8, out: *protocol.FileStatus) i32 {
// Open, stat by node, close — simple and enough for now.
const fd = open(path, 0);
if (fd < 0) return -1;
defer close(fd);
const f = fdPtr(fd).?;
const request = protocol.Request{ .operation = .status, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
var sbuf: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &sbuf) orelse return -1;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return -1;
out.* = std.mem.bytesToValue(protocol.FileStatus, sbuf[0..@sizeOf(protocol.FileStatus)]);
return 0;
}
/// Close an fd (best effort — tells the VFS to release the open file).
pub fn close(fd: i32) void {
const f = fdPtr(fd) orelse return;
const request = protocol.Request{ .operation = .close, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
_ = transact(request, &.{}, &.{});
f.used = false;
}
+62
View File
@@ -0,0 +1,62 @@
//! Block-device client: the helper a filesystem uses to read and write a block
//! device (a USB stick, via usb-storage) without hand-rolling the block-protocol
//! IPC. Layered over `ipc` and the shared `block-protocol` wire format, like
//! `runtime.usb` over the transfer protocol.
//!
//! Transfers name a caller-owned DMA buffer by physical address (from
//! `runtime.dma.alloc`), so whole sectors move without crossing the IPC size
//! limit — the same handoff usb-storage uses toward the controller.
const std = @import("std");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
const protocol = @import("block-protocol");
pub const Geometry = struct { block_size: u32, block_count: u64 };
pub const Device = struct {
endpoint: ipc.Handle,
/// The device's block size and total block count.
pub fn geometry(self: Device) ?Geometry {
var request = protocol.Request{ .operation = @intFromEnum(protocol.Operation.geometry), .lba = 0, .count = 0, .physical = 0 };
var reply: [protocol.reply_size]u8 = undefined;
const n = ipc.call(self.endpoint, std.mem.asBytes(&request), &reply) catch return null;
if (n < protocol.reply_size) return null;
const result = std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]);
if (result.status != 0) return null;
return .{ .block_size = result.block_size, .block_count = result.block_count };
}
/// Read `count` blocks starting at `lba` into the DMA buffer at `physical`.
pub fn read(self: Device, lba: u64, count: u32, physical: u64) bool {
return self.transfer(.read, lba, count, physical);
}
/// Write `count` blocks starting at `lba` from the DMA buffer at `physical`.
pub fn write(self: Device, lba: u64, count: u32, physical: u64) bool {
return self.transfer(.write, lba, count, physical);
}
fn transfer(self: Device, operation: protocol.Operation, lba: u64, count: u32, physical: u64) bool {
var request = protocol.Request{ .operation = @intFromEnum(operation), .lba = lba, .count = count, .physical = physical };
var reply: [protocol.reply_size]u8 = undefined;
const n = ipc.call(self.endpoint, std.mem.asBytes(&request), &reply) catch return false;
if (n < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]).status == 0;
}
};
/// Look up the block device, retrying generously while the USB storage chain
/// (controller reset, enumeration, mass-storage bring-up) comes up.
pub fn open() ?Device {
// Patient: the whole USB storage chain (firmware discovery, xHCI reset and
// enumeration, mass-storage bring-up) must complete first, which can take
// tens of seconds under emulation.
var attempts: usize = 0;
while (attempts < 1200) : (attempts += 1) {
if (ipc.lookup(.block)) |handle| return .{ .endpoint = handle };
system.sleep(50);
}
return null;
}
+274
View File
@@ -0,0 +1,274 @@
//! runtime.fs — the danos-native file API. A program opens, reads, writes, and
//! lists files served by the user-space VFS (system/services/vfs), each call
//! marshalling a vfs-protocol request over IPC. This is the danos-native layer
//! danos programs use directly; it is also where the file operations that later
//! become `std.os.danos` are staged (see docs/zig-self-hosting.md). It replaces
//! the old POSIX `unistd` shim — a compatibility spelling danos does not need yet.
//!
//! Handles are *values*, not entries in a global descriptor table: a `File` /
//! `Directory` owns its VFS node id and (for files) a byte offset. So there is no
//! per-process fd limit and no shared table to synchronise — the danos-native
//! shape, unlike the POSIX fd model the old shim emulated.
const std = @import("std");
const ipc = @import("ipc.zig");
const protocol = @import("vfs-protocol");
/// The kind of a filesystem node — re-exported so a caller need not import the
/// wire protocol.
pub const Kind = protocol.NodeKind;
/// A node's metadata (the answer to a status request).
pub const Attributes = struct {
size: u64,
kind: Kind,
/// Modification time — Unix epoch seconds, UTC. 0 if the filesystem has none.
mtime: u64 = 0,
};
// Map a wire `NodeKind` value to the enum, defaulting anything unrecognised to
// `.regular` (the server is trusted, but a value outside the enum would be
// illegal to `@enumFromInt` directly).
fn kindFromWire(value: u32) Kind {
return switch (value) {
@intFromEnum(Kind.directory) => .directory,
@intFromEnum(Kind.character_device) => .character_device,
@intFromEnum(Kind.block_device) => .block_device,
@intFromEnum(Kind.symbolic_link) => .symbolic_link,
@intFromEnum(Kind.fifo) => .fifo,
@intFromEnum(Kind.socket) => .socket,
else => .regular,
};
}
/// How to open a path.
pub const OpenOptions = struct {
/// Create the file if it does not exist.
create: bool = false,
/// Open a directory node (for listing) rather than a file.
directory: bool = false,
/// Truncate an existing file to zero length on open (O_TRUNC) — replace its
/// contents rather than overwriting in place.
truncate: bool = false,
fn wireFlags(self: OpenOptions) u32 {
var f: u32 = 0;
if (self.create) f |= protocol.create;
if (self.directory) f |= protocol.directory;
if (self.truncate) f |= protocol.truncate;
return f;
}
};
// The VFS server endpoint, looked up once by well-known id and cached.
var vfs_handle: ipc.Handle = 0;
var vfs_resolved = false;
fn vfs() ?ipc.Handle {
if (!vfs_resolved) {
vfs_handle = ipc.lookup(.vfs) orelse return null;
vfs_resolved = true;
}
return vfs_handle;
}
const Result = struct { reply: protocol.Reply, payload: []u8 };
// One request/reply round trip: [Request header][send payload] -> VFS ->
// [Reply header][receive payload]. The receive payload lands in `out`.
fn transact(request: protocol.Request, send: []const u8, out: []u8) ?Result {
const h = vfs() orelse return null;
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const slen = @min(send.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..slen], send[0..slen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(h, message[0 .. protocol.request_size + slen], &rbuf) catch return null;
if (n < protocol.reply_size) return null;
const reply = std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]);
const rpl = @min(n - protocol.reply_size, out.len);
@memcpy(out[0..rpl], rbuf[protocol.reply_size..][0..rpl]);
return .{ .reply = reply, .payload = out[0..rpl] };
}
/// An open file: a VFS node plus a byte cursor. Read and write advance the cursor.
pub const File = struct {
node: u64,
offset: u64 = 0,
/// Read up to `buffer.len` bytes at the current offset; returns the count, or
/// null on error.
pub fn read(self: *File, buffer: []u8) ?usize {
const want: u32 = @intCast(@min(buffer.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .read, .node = self.node, .offset = self.offset, .len = want, .flags = 0 };
const r = transact(request, &.{}, buffer) orelse return null;
if (r.reply.status != 0) return null;
self.offset += r.reply.len;
return r.reply.len;
}
/// Write `data` at the current offset; returns the count written. A single
/// call is capped at the VFS payload size, so the return may be short — use
/// `writeAll` to write the whole slice. Null on error.
pub fn write(self: *File, data: []const u8) ?usize {
const want: u32 = @intCast(@min(data.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .write, .node = self.node, .offset = self.offset, .len = want, .flags = 0 };
const r = transact(request, data[0..want], &.{}) orelse return null;
if (r.reply.status != 0) return null;
self.offset += r.reply.len;
return r.reply.len;
}
/// Write all of `data`, looping past the per-call payload cap. Returns the
/// total written, or null if a write failed before any progress.
pub fn writeAll(self: *File, data: []const u8) ?usize {
var written: usize = 0;
while (written < data.len) {
const n = self.write(data[written..]) orelse return if (written == 0) null else written;
if (n == 0) return written; // no forward progress; stop rather than spin
written += n;
}
return written;
}
/// Move the read/write cursor to an absolute byte position.
pub fn seekTo(self: *File, position: u64) void {
self.offset = position;
}
/// This file's metadata.
pub fn attributes(self: *File) ?Attributes {
const request = protocol.Request{ .operation = .status, .node = self.node, .offset = 0, .len = 0, .flags = 0 };
var buffer: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return null;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return null;
const status = std.mem.bytesToValue(protocol.FileStatus, buffer[0..@sizeOf(protocol.FileStatus)]);
return .{ .size = status.size, .kind = kindFromWire(status.kind), .mtime = status.mtime };
}
/// Release the VFS's open handle for this file.
pub fn close(self: *File) void {
const request = protocol.Request{ .operation = .close, .node = self.node, .offset = 0, .len = 0, .flags = 0 };
_ = transact(request, &.{}, &.{});
}
};
/// Open (or create, with `.create`) `path`. Returns the open file, or null.
pub fn open(path: []const u8, options: OpenOptions) ?File {
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = options.wireFlags() };
const r = transact(request, path, &.{}) orelse return null;
if (r.reply.status != 0) return null;
return .{ .node = r.reply.node };
}
/// A path's metadata without keeping it open (open -> status -> close).
pub fn attributes(path: []const u8) ?Attributes {
var file = open(path, .{}) orelse return null;
defer file.close();
return file.attributes();
}
/// Whether `path` resolves — handy as a readiness check (e.g. waiting for a mount
/// to come up before writing to it).
pub fn exists(path: []const u8) bool {
return attributes(path) != null;
}
/// One entry returned by `Directory.next`.
pub const Entry = struct {
kind: Kind = .regular,
size: u64 = 0,
name_buffer: [64]u8 = undefined,
name_len: usize = 0,
pub fn name(self: *const Entry) []const u8 {
return self.name_buffer[0..self.name_len];
}
};
/// An open directory being listed, cursor-advanced by `next`.
pub const Directory = struct {
node: u64,
cursor: u64 = 0,
/// Fill `entry` with the next directory entry; false at end of directory or
/// on error.
pub fn next(self: *Directory, entry: *Entry) bool {
const request = protocol.Request{ .operation = .readdir, .node = self.node, .offset = self.cursor, .len = 0, .flags = 0 };
var buffer: [protocol.message_maximum]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return false;
if (r.reply.status != 0 or r.reply.len == 0) return false; // error or EOF
if (r.payload.len < protocol.directory_entry_size) return false;
const header = std.mem.bytesToValue(protocol.DirectoryEntry, r.payload[0..protocol.directory_entry_size]);
entry.kind = kindFromWire(header.kind);
entry.size = header.size;
const source = r.payload[protocol.directory_entry_size..];
const nlen = @min(@min(@as(usize, header.name_len), source.len), entry.name_buffer.len);
@memcpy(entry.name_buffer[0..nlen], source[0..nlen]);
entry.name_len = nlen;
self.cursor += 1;
return true;
}
/// Release the VFS's open handle for this directory.
pub fn close(self: *Directory) void {
var f = File{ .node = self.node };
f.close();
}
};
/// Open `path` as a directory for listing. Returns null if it isn't one / on error.
pub fn openDirectory(path: []const u8) ?Directory {
const file = open(path, .{ .directory = true }) orelse return null;
return .{ .node = file.node };
}
// A path-based request that returns only a status (mkdir, unlink).
fn pathOperation(operation: protocol.Operation, path: []const u8) bool {
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = 0 };
const r = transact(request, path, &.{}) orelse return false;
return r.reply.status == 0;
}
/// Create a directory at `path` (its parent must already exist). Returns true on
/// success. Only works under a mounted filesystem that supports directories.
pub fn makeDirectory(path: []const u8) bool {
return pathOperation(.mkdir, path);
}
/// Remove the file at `path`. Returns true on success. Directories are refused
/// (a separate directory-removal would have to check emptiness).
pub fn remove(path: []const u8) bool {
return pathOperation(.unlink, path);
}
/// Rename `old_path` to `new_path`. Both must be in the same directory (same-
/// directory, 8.3-name rename only for now). Returns true on success.
pub fn rename(old_path: []const u8, new_path: []const u8) bool {
const total = old_path.len + 1 + new_path.len;
if (total > protocol.maximum_payload) return false;
var payload: [protocol.maximum_payload]u8 = undefined;
@memcpy(payload[0..old_path.len], old_path);
payload[old_path.len] = 0;
@memcpy(payload[old_path.len + 1 ..][0..new_path.len], new_path);
const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 };
const r = transact(request, payload[0..total], &.{}) orelse return false;
return r.reply.status == 0;
}
/// Mount a filesystem backend (its server endpoint) at absolute path `target`;
/// the VFS then routes everything under `target` to that backend. This is the one
/// call that hands the VFS a capability (the backend endpoint). Returns true on
/// success.
pub fn mount(target: []const u8, backend: ipc.Handle) bool {
const h = vfs() orelse return false;
const request = protocol.Request{ .operation = .mount, .node = 0, .offset = 0, .len = @intCast(target.len), .flags = 0 };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const tlen = @min(target.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..tlen], target[0..tlen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const result = ipc.callCap(h, message[0 .. protocol.request_size + tlen], &rbuf, backend) catch return false;
if (result.len < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]).status == 0;
}
+9
View File
@@ -42,6 +42,15 @@ pub const dma = @import("dma.zig");
/// (control / interrupt / bulk transfers). See library/runtime/usb.zig. /// (control / interrupt / bulk transfers). See library/runtime/usb.zig.
pub const usb = @import("usb.zig"); pub const usb = @import("usb.zig");
/// Block-device client: read/write a block device (a USB stick, via
/// usb-storage). See library/runtime/block.zig.
pub const block = @import("block.zig");
/// The danos-native file API (open/read/write/list over the user-space VFS) — the
/// layer danos programs use directly, and where the operations that later become
/// `std.os.danos` are staged. See docs/zig-self-hosting.md.
pub const fs = @import("fs.zig");
/// Re-exported so a user binary can `pub const panic = runtime.panic;`. /// Re-exported so a user binary can `pub const panic = runtime.panic;`.
pub const panic = start.panic; pub const panic = start.panic;
+18
View File
@@ -51,6 +51,24 @@ pub fn clock() u64 {
return @intCast(sc.systemCall0(.clock)); return @intCast(sc.systemCall0(.clock));
} }
/// Wall-clock time in Unix epoch seconds (UTC) — the real date/time, from the RTC.
/// Unlike `clock` (monotonic since boot), this tracks calendar time, so it is what a
/// filesystem stamps as a file's modification time. Formatting it into a calendar
/// date/timezone is user-space policy layered on top.
pub fn wallClock() u64 {
return @intCast(sc.systemCall0(.wall_clock));
}
/// Copy bytes out of the kernel's in-memory diagnostic log — the accumulated
/// stream of everything `write` (and the kernel itself) has emitted — starting at
/// `offset`, into `out`. Returns the number of bytes copied (0 at end of buffer).
/// A program reads the whole log by looping from offset 0, advancing by the return
/// value, until it gets 0. This is how the boot log is persisted to disk on a
/// headless/real machine where serial output is otherwise lost.
pub fn klogRead(offset: usize, out: []u8) usize {
return sc.systemCall3(.klog_read, offset, @intFromPtr(out.ptr), out.len);
}
/// End the process. Never returns. /// End the process. Never returns.
pub fn exit(code: usize) noreturn { pub fn exit(code: usize) noreturn {
_ = sc.systemCall1(.exit, code); _ = sc.systemCall1(.exit, code);
+3
View File
@@ -58,6 +58,8 @@ pub const SystemCall = enum(u64) {
signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on
process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself) process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself)
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
klog_read = 32, // klog_read(offset, ptr, len) -> bytes copied: copy the kernel RAM log buffer out to a user buffer (for persisting the boot log to disk)
wall_clock = 33, // wall_clock() -> Unix epoch seconds (UTC): the RTC wall-clock time, for filesystem timestamps (mtime). Monotonic time is `clock`.
_, _,
}; };
@@ -180,6 +182,7 @@ pub const ServiceId = enum(u32) {
power = 5, // system power: events (button, lid, battery) + shutdown (docs/power.md; domain-named per docs/discovery.md — the acpi service registers it on x86, a PSCI service will on ARM) power = 5, // system power: events (button, lid, battery) + shutdown (docs/power.md; domain-named per docs/discovery.md — the acpi service registers it on x86, a PSCI service will on ARM)
usb_bus = 6, // the xHCI host-controller driver's transfer endpoint; USB class drivers look it up and `callCap`-open their device to get a private per-device transfer channel (docs/driver-model.md) usb_bus = 6, // the xHCI host-controller driver's transfer endpoint; USB class drivers look it up and `callCap`-open their device to get a private per-device transfer channel (docs/driver-model.md)
block = 7, // a block-device driver (USB mass storage today): read/write of fixed-size blocks, the storage a filesystem sits on block = 7, // a block-device driver (USB mass storage today): read/write of fixed-size blocks, the storage a filesystem sits on
fat = 8, // the FAT filesystem server; the VFS mounts it and forwards paths under its mount point (/mnt/usb) to it
_, _,
}; };
+89
View File
@@ -469,6 +469,95 @@ pub fn clockHz() u64 {
return apic.tscHz(); return apic.tscHz();
} }
// --- real-time clock (CMOS) --------------------------------------------------
//
// The battery-backed CMOS clock, read once at boot and thereafter anchored to the
// monotonic clock (see kernel/wall-clock.zig) — so this is never on a hot path and
// needs no lock. Wall-clock *seconds* are mechanism the kernel owns (the hardware's
// value), like the monotonic clock; calendars/timezones are policy layered on top.
fn cmosRead(register: u8) u8 {
io.outb(0x70, register);
return io.inb(0x71);
}
const RtcFields = struct { second: u8, minute: u8, hour: u8, day: u8, month: u8, year: u8 };
fn rtcRaw() RtcFields {
while (cmosRead(0x0A) & 0x80 != 0) {} // wait out any update in progress (status A bit 7)
return .{
.second = cmosRead(0x00),
.minute = cmosRead(0x02),
.hour = cmosRead(0x04),
.day = cmosRead(0x07),
.month = cmosRead(0x08),
.year = cmosRead(0x09),
};
}
fn bcdToBinary(v: u8) u8 {
return (v & 0x0F) + ((v >> 4) * 10);
}
fn isLeapYear(y: u32) bool {
return (y % 4 == 0 and y % 100 != 0) or (y % 400 == 0);
}
/// Read the CMOS real-time clock and convert it to Unix epoch seconds (UTC).
pub fn readRtcUnixSeconds() u64 {
// Read until two consecutive reads agree, so we never latch a half-updated time.
var a = rtcRaw();
while (true) {
const b = rtcRaw();
if (a.second == b.second and a.minute == b.minute and a.hour == b.hour and
a.day == b.day and a.month == b.month and a.year == b.year) break;
a = b;
}
const status_b = cmosRead(0x0B);
const binary_mode = status_b & 0x04 != 0; // else BCD
const hour_24 = status_b & 0x02 != 0; // else 12-hour with a PM bit
var second = a.second;
var minute = a.minute;
var hour_field = a.hour;
var day = a.day;
var month = a.month;
var year = a.year;
if (!binary_mode) {
second = bcdToBinary(second);
minute = bcdToBinary(minute);
hour_field = bcdToBinary(hour_field & 0x7F) | (hour_field & 0x80); // preserve the PM bit
day = bcdToBinary(day);
month = bcdToBinary(month);
year = bcdToBinary(year);
}
var hour: u32 = hour_field & 0x7F;
if (!hour_24) {
const pm = hour_field & 0x80 != 0;
hour %= 12; // 12 AM/PM -> 0
if (pm) hour += 12;
}
// The CMOS year is 0..99; QEMU and modern hardware mean 20xx (there is no
// reliable century register on QEMU). Treat < 70 as 20xx, else 19xx.
const full_year: u32 = if (year < 70) 2000 + @as(u32, year) else 1900 + @as(u32, year);
var days: u64 = 0;
var y: u32 = 1970;
while (y < full_year) : (y += 1) days += if (isLeapYear(y)) 366 else 365;
const month_lengths = [_]u8{ 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31 };
var m: u8 = 1;
while (m < month) : (m += 1) {
days += month_lengths[m - 1];
if (m == 2 and isLeapYear(full_year)) days += 1;
}
days += @as(u64, day) - 1;
return ((days * 24 + hour) * 60 + minute) * 60 + second;
}
/// Whether the CPU guarantees an **invariant** TSC (CPUID 0x80000007 EDX[8] on /// Whether the CPU guarantees an **invariant** TSC (CPUID 0x80000007 EDX[8] on
/// x86; the analogous architectural guarantee elsewhere). When false the TSC is not /// x86; the analogous architectural guarantee elsewhere). When false the TSC is not
/// used as the clocksource. /// used as the clocksource.
+3 -1
View File
@@ -37,7 +37,9 @@ const Task = scheduler.Task;
pub const MESSAGE_MAXIMUM: usize = 256; pub const MESSAGE_MAXIMUM: usize = 256;
pub const maximum_handles = scheduler.ipc_maximum_handles; pub const maximum_handles = scheduler.ipc_maximum_handles;
pub const maximum_services = 8; // The name registry is indexed directly by ServiceId, so this must exceed the
// largest id (currently fat = 8). Sized with headroom for new services.
pub const maximum_services = 16;
/// Errno-style failures, returned as `-value` in the system_call result register. /// Errno-style failures, returned as `-value` in the system_call result register.
pub const EBADF: i64 = 1; // bad handle pub const EBADF: i64 = 1; // bad handle
+9
View File
@@ -5,6 +5,7 @@ const parameters = @import("parameters");
const architecture = @import("architecture"); const architecture = @import("architecture");
const console = @import("console.zig"); const console = @import("console.zig");
const log = @import("log.zig"); const log = @import("log.zig");
const wall_clock = @import("wall-clock.zig");
const pmm = @import("pmm.zig"); const pmm = @import("pmm.zig");
const heap = @import("heap.zig"); const heap = @import("heap.zig");
const scheduler = @import("scheduler.zig"); const scheduler = @import("scheduler.zig");
@@ -62,6 +63,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
architecture.serialInit(); architecture.serialInit();
log.addSink(architecture.serialWrite); log.addSink(architecture.serialWrite);
if (architecture.debugconPresent()) log.addSink(architecture.debugconWrite); if (architecture.debugconPresent()) log.addSink(architecture.debugconWrite);
// Retain the whole stream in a RAM buffer too, so a user program can later
// read it back (klog_read) and persist the boot log to disk — the only way to
// see it on a headless/real machine with no host capturing serial.
log.addSink(log.ramSink);
// The **framebuffer** is deliberately *not* a log sink. It's a separate output // The **framebuffer** is deliberately *not* a log sink. It's a separate output
// surface — a bootstrap text console today, a graphics device driver later — so // surface — a bootstrap text console today, a graphics device driver later — so
@@ -281,6 +286,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
if (!architecture.clockSynchronized()) if (!architecture.clockSynchronized())
log.write("/system/kernel: WARNING: per-core TSCs are not synchronized; monotonic clock moved off the TSC\n"); log.write("/system/kernel: WARNING: per-core TSCs are not synchronized; monotonic clock moved off the TSC\n");
// Anchor wall-clock time: read the RTC once, now the monotonic clock is final.
wall_clock.init();
log.print("/system/kernel: wall clock {d} (Unix epoch seconds, UTC, from the RTC)\n", .{wall_clock.nowSeconds()});
// In a test build (`zig build -Dtest-case=<name>`), run that case and stop. // In a test build (`zig build -Dtest-case=<name>`), run that case and stop.
// Normal builds fall through to the idle halt. // Normal builds fall through to the idle halt.
if (build_options.test_case) |case| { if (build_options.test_case) |case| {
+33
View File
@@ -41,6 +41,39 @@ pub fn write(bytes: []const u8) void {
for (sinks[0..sink_count]) |sink| sink(bytes); for (sinks[0..sink_count]) |sink| sink(bytes);
} }
// --- the RAM sink: a retained copy of the whole diagnostic stream ------------
//
// A fixed in-image buffer that accumulates every logged byte, so a user program
// (`log-flush`, and init at shutdown) can read it back through `klog_read` and
// persist it to a file — the boot log survives on a headless/real machine that
// has no host capturing serial. It is a *sink like any other*: register it with
// `addSink(ramSink)` at boot. No allocation (works pre-heap and in a panic).
//
// It fills linearly and stops when full: the earliest output — the most valuable
// for diagnosing a boot — is kept, and the tail is still on the live serial sink.
// 256 KiB comfortably holds a full boot plus a long run (a boot is ~15 KiB).
const ram_capacity = 256 * 1024;
var ram_buffer: [ram_capacity]u8 = undefined;
var ram_len: usize = 0;
/// The RAM sink. Best-effort and self-guarding like every sink: appends what fits
/// and silently drops the rest once full. (Concurrency matches the other sinks —
/// the dominant writer, debug_write, already holds the kernel lock; a rare torn
/// append on a kernel-internal line is an accepted diagnostic imperfection.)
pub fn ramSink(bytes: []const u8) void {
const n = @min(ram_buffer.len - ram_len, bytes.len);
if (n != 0) {
@memcpy(ram_buffer[ram_len..][0..n], bytes[0..n]);
ram_len += n;
}
}
/// The accumulated log so far — what `klog_read` copies out.
pub fn ramSnapshot() []const u8 {
return ram_buffer[0..ram_len];
}
/// A formatted log line. Truncates past 256 bytes; the buffer is on the stack, so /// A formatted log line. Truncates past 256 bytes; the buffer is on the stack, so
/// this is safe to call from interrupt context and from a panic. /// this is safe to call from interrupt context and from a panic.
pub fn print(comptime fmt: []const u8, args: anytype) void { pub fn print(comptime fmt: []const u8, args: anytype) void {
+41
View File
@@ -34,6 +34,7 @@ const devices_broker = @import("devices-broker.zig");
const irq = @import("irq.zig"); const irq = @import("irq.zig");
const initial_ramdisk = @import("initial-ramdisk"); const initial_ramdisk = @import("initial-ramdisk");
const log = @import("log.zig"); const log = @import("log.zig");
const wall_clock = @import("wall-clock.zig");
const page_size = abi.page_size; const page_size = abi.page_size;
const SystemCall = abi.SystemCall; const SystemCall = abi.SystemCall;
@@ -206,6 +207,8 @@ fn system_call(state: *architecture.CpuState) void {
.signal_bind => systemSignalBind(state), .signal_bind => systemSignalBind(state),
.process_signal => systemProcessSignal(state), .process_signal => systemProcessSignal(state),
.timer_bind => systemTimerBind(state), .timer_bind => systemTimerBind(state),
.klog_read => systemKlogRead(state),
.wall_clock => systemWallClock(state),
_ => fail(state), _ => fail(state),
} }
} }
@@ -974,6 +977,37 @@ fn systemDebugWrite(state: *architecture.CpuState) void {
} }
} }
/// klog_read(offset, ptr, len) -> bytes copied: copy the kernel's in-memory
/// diagnostic log (the RAM sink in log.zig) out to the user buffer at `ptr`,
/// starting at `offset`. Returns the count copied — 0 once `offset` reaches the
/// end — so a program reads the whole log by looping from 0 until it gets 0.
///
/// The mirror of `debug_write`: the same overflow-safe user-half bounds check,
/// but the copy runs kernel -> user. Written under the kernel lock so the source
/// snapshot can't grow underneath the copy. A read-only diagnostic — it exposes
/// only the log the kernel already broadcasts to serial, nothing else.
fn systemKlogRead(state: *architecture.CpuState) void {
const offset = architecture.systemCallArg(state, 0);
const ptr = architecture.systemCallArg(state, 1);
const len = architecture.systemCallArg(state, 2);
// Confine the whole destination span to the user (low) half. `len <=
// user_half_end - ptr` bounds the length without an overflowing add.
if (ptr < user_half_end and len <= user_half_end - ptr) {
const flags = sync.enter();
defer sync.leave(flags);
const snapshot = log.ramSnapshot();
var n: usize = 0;
if (offset < snapshot.len) {
n = @min(len, snapshot.len - offset);
const dest: [*]u8 = @ptrFromInt(ptr);
@memcpy(dest[0..n], snapshot[offset..][0..n]);
}
architecture.setSystemCallResult(state, n);
} else {
fail(state);
}
}
/// mmap(len, prot) -> base: grant `len` bytes (rounded up to whole pages) of /// mmap(len, prot) -> base: grant `len` bytes (rounded up to whole pages) of
/// fresh, zeroed, writable+NX memory in the caller's mmap arena, and return the /// fresh, zeroed, writable+NX memory in the caller's mmap arena, and return the
/// base virtual address. `prot` is accepted but not yet honoured (grants are /// base virtual address. `prot` is accepted but not yet honoured (grants are
@@ -1304,3 +1338,10 @@ pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []c
fn systemClock(state: *architecture.CpuState) void { fn systemClock(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, architecture.nanos()); architecture.setSystemCallResult(state, architecture.nanos());
} }
/// wall_clock() -> Unix epoch seconds (UTC). The RTC value, read at boot and offset
/// by the monotonic clock (wall-clock.zig) — mechanism, not policy: calendars and
/// timezones layer on top in user space. Needed for filesystem timestamps (mtime).
fn systemWallClock(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, wall_clock.nowSeconds());
}
+42
View File
@@ -14,6 +14,7 @@ const boot_handoff = @import("boot-handoff");
const abi = @import("abi"); const abi = @import("abi");
const device_abi = @import("device-abi"); const device_abi = @import("device-abi");
const architecture = @import("architecture"); const architecture = @import("architecture");
const wall_clock = @import("wall-clock.zig");
const devices_broker = @import("devices-broker.zig"); const devices_broker = @import("devices-broker.zig");
const platform = @import("platform"); const platform = @import("platform");
const pmm = @import("pmm.zig"); const pmm = @import("pmm.zig");
@@ -66,6 +67,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
timer(); timer();
} else if (eql(case, "clock")) { } else if (eql(case, "clock")) {
clock(); clock();
} else if (eql(case, "wall-clock")) {
wallClock();
} else if (eql(case, "vmm")) { } else if (eql(case, "vmm")) {
vmm(); vmm();
} else if (eql(case, "heap")) { } else if (eql(case, "heap")) {
@@ -148,6 +151,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
usbHidTest(boot_information); usbHidTest(boot_information);
} else if (eql(case, "usb-storage")) { } else if (eql(case, "usb-storage")) {
usbStorageTest(boot_information); usbStorageTest(boot_information);
} else if (eql(case, "fat-mount")) {
fatMountTest(boot_information);
} else if (eql(case, "device-list")) { } else if (eql(case, "device-list")) {
deviceListTest(boot_information); deviceListTest(boot_information);
} else if (eql(case, "pci-scan")) { } else if (eql(case, "pci-scan")) {
@@ -456,6 +461,17 @@ fn heapTest() void {
/// Verify the calibrated clocks: sane measured frequencies, monotonic uptime that /// Verify the calibrated clocks: sane measured frequencies, monotonic uptime that
/// advances with real ticks, and — the point of the TSC clock — nanosecond /// advances with real ticks, and — the point of the TSC clock — nanosecond
/// resolution far finer than the 1 ms tick, with the unit functions consistent. /// resolution far finer than the 1 ms tick, with the unit functions consistent.
fn wallClock() void {
log("DANOS-TEST-BEGIN: wall-clock\n", .{});
// The RTC was read and anchored at boot (kmain -> wall_clock.init()).
const seconds = wall_clock.nowSeconds();
log(" epoch: {d}\n", .{seconds});
// A plausible current wall-clock: after 2020-01-01 (1577836800) and before 2050
// (2524608000) — catches a broken CMOS read or a wrong epoch conversion.
check("wall clock reads a plausible current epoch", seconds > 1_577_836_800 and seconds < 2_524_608_000);
result();
}
fn clock() void { fn clock() void {
log("DANOS-TEST-BEGIN: clock\n", .{}); log("DANOS-TEST-BEGIN: clock\n", .{});
@@ -1991,6 +2007,32 @@ fn usbStorageTest(boot_information: *const BootInformation) void {
bootServiceTreeTest(boot_information, "usb-storage"); bootServiceTreeTest(boot_information, "usb-storage");
} }
/// The FAT mount chain: boot the full tree (init spawns the fat server, which
/// brings up the USB storage chain, mounts the FAT volume, and mounts itself into
/// the VFS at /mnt/usb), then spawn a fat-test client that lists and reads through
/// the mount. The harness attaches a usb-storage device; the expect regex requires
/// the fat mount and the client's success.
fn fatMountTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: fat-mount\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over init and the initial_ramdisk", false);
result();
return;
}
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(ramdisk);
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const init_ok = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned (boots the tree, incl. the fat server)", init_ok);
check("fat-test client spawned", spawnNamed(rd, "fat-test"));
result();
}
fn bootServiceTreeTest(boot_information: *const BootInformation, comptime label: []const u8) void { fn bootServiceTreeTest(boot_information: *const BootInformation, comptime label: []const u8) void {
log("DANOS-TEST-BEGIN: " ++ label ++ "\n", .{}); log("DANOS-TEST-BEGIN: " ++ label ++ "\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) { if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
+26
View File
@@ -0,0 +1,26 @@
//! Wall-clock time: the CMOS real-time clock read once at boot and anchored to the
//! monotonic clock, so a query is a cheap arithmetic offset — no per-call CMOS poll,
//! no lock, no SMP hazard on the shared 0x70/0x71 ports.
//!
//! Wall-clock *seconds* are mechanism the kernel owns, exactly like the monotonic
//! clock ([[time-architecture]]): reading the hardware's value is not policy.
//! Calendars, timezones, and formatting layer on top in user space. It exists so the
//! filesystem can stamp real timestamps (mtime) — see docs/zig-self-hosting.md.
const architecture = @import("architecture");
var boot_unix_seconds: u64 = 0;
var boot_nanos: u64 = 0;
/// Read the RTC once and anchor it to the monotonic clock. Call at boot, after the
/// monotonic clock is calibrated.
pub fn init() void {
boot_unix_seconds = architecture.readRtcUnixSeconds();
boot_nanos = architecture.nanos();
}
/// The current wall-clock time in Unix epoch seconds (UTC): the boot RTC value plus
/// the monotonic time elapsed since. Zero until `init` runs.
pub fn nowSeconds() u64 {
return boot_unix_seconds + (architecture.nanos() -% boot_nanos) / 1_000_000_000;
}
File diff suppressed because it is too large Load Diff
+108
View File
@@ -0,0 +1,108 @@
//! system/services/fat/fat-test — a client that proves the FAT mount end to end:
//! it waits for the fat server to mount the USB volume at /mnt/usb, lists the
//! root directory through the VFS (which routes /mnt/usb to the fat backend), and
//! reads a known file off it. Shipped in the initial_ramdisk; the `fat-mount`
//! kernel test spawns it alongside init.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main(init: runtime.process.Init) void {
_ = init;
// Wait for /mnt/usb to be mounted — the fat server races us at boot (it must
// bring up the whole USB storage chain first).
var opened: ?fs.Directory = null;
var tries: u32 = 0;
while (opened == null and tries < 1400) : (tries += 1) {
opened = fs.openDirectory("/mnt/usb");
if (opened == null) runtime.system.sleep(50);
}
var dir = opened orelse {
_ = runtime.system.write("fat-test: /mnt/usb never became available\n");
return;
};
var count: u32 = 0;
var entry: fs.Entry = .{};
while (dir.next(&entry)) {
writeLine("fat-test: entry '{s}' kind={d} size={d}\n", .{ entry.name(), @intFromEnum(entry.kind), entry.size });
count += 1;
if (count > 32) break;
}
dir.close();
writeLine("fat-test: listed {d} entries\n", .{count});
// Read a known file off the boot volume through the mount (best effort): the
// kernel image is an ELF, so its first bytes are the ELF magic.
if (fs.open("/mnt/usb/system/kernel", .{})) |opened_file| {
var file = opened_file;
var magic: [4]u8 = undefined;
const n = file.read(&magic) orelse 0;
file.close();
if (n == 4 and magic[0] == 0x7F and magic[1] == 'E' and magic[2] == 'L' and magic[3] == 'F') {
_ = runtime.system.write("fat-test: read /mnt/usb/system/kernel ELF magic ok\n");
} else {
writeLine("fat-test: /mnt/usb/system/kernel read {d} bytes (not ELF magic)\n", .{n});
}
}
// Exercise directory + file mutation through the mount: mkdir, create a file
// inside it, read it back, then remove it — proof mkdir/unlink reach the engine.
if (fs.makeDirectory("/mnt/usb/TESTDIR")) {
var wrote = false;
if (fs.open("/mnt/usb/TESTDIR/HELLO.TXT", .{ .create = true, .truncate = true })) |created| {
var f = created;
wrote = (f.writeAll("mutation-ok") orelse 0) == "mutation-ok".len;
f.close();
}
// The created file carries a real modification time (stamped from the RTC).
var mtime_ok = false;
if (fs.attributes("/mnt/usb/TESTDIR/HELLO.TXT")) |attrs| {
writeLine("fat-test: mtime {d}\n", .{attrs.mtime});
mtime_ok = attrs.mtime > 1_577_836_800; // after 2020-01-01
}
if (mtime_ok) _ = runtime.system.write("fat-test: mtime ok\n");
// Rename it, then read from the new name and confirm the old name is gone.
const renamed = fs.rename("/mnt/usb/TESTDIR/HELLO.TXT", "/mnt/usb/TESTDIR/RENAMED.TXT");
const old_gone = !fs.exists("/mnt/usb/TESTDIR/HELLO.TXT");
if (renamed and old_gone) _ = runtime.system.write("fat-test: rename ok\n");
var readback = false;
if (fs.open("/mnt/usb/TESTDIR/RENAMED.TXT", .{})) |reopened| {
var f = reopened;
var buf: [16]u8 = undefined;
const got = f.read(&buf) orelse 0;
f.close();
readback = std.mem.eql(u8, buf[0..got], "mutation-ok");
}
const removed = fs.remove("/mnt/usb/TESTDIR/RENAMED.TXT");
const gone = !fs.exists("/mnt/usb/TESTDIR/RENAMED.TXT");
if (wrote and mtime_ok and renamed and old_gone and readback and removed and gone) {
_ = runtime.system.write("fat-test: mutations ok\n");
} else {
writeLine("fat-test: mutations FAILED (wrote={} mtime={} renamed={} oldgone={} read={} removed={} gone={})\n", .{ wrote, mtime_ok, renamed, old_gone, readback, removed, gone });
}
} else {
_ = runtime.system.write("fat-test: mkdir /mnt/usb/TESTDIR failed\n");
}
if (count > 0) {
while (true) {
_ = runtime.system.write("fat-test: ok\n");
runtime.system.sleep(1000);
}
}
_ = runtime.system.write("fat-test: root listing was empty\n");
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+236
View File
@@ -0,0 +1,236 @@
//! system/services/fat — the FAT filesystem server. Spawned as a boot service, it
//! opens the block device (a USB stick via usb-storage) under `.block`, mounts the
//! FAT filesystem on it (the pure engine in engine.zig), and mounts itself into
//! the VFS at /mnt/usb. From then on the VFS forwards every open/read/write/
//! status/readdir/close under /mnt/usb to this server, which serves the same
//! vfs-protocol as a backend — turning block reads into file reads.
//!
//! The block data path never crosses IPC: a DMA bounce buffer is handed to the
//! block driver by physical address, and the engine copies sectors in and out of
//! it.
const std = @import("std");
const runtime = @import("runtime");
const engine = @import("engine.zig");
const on_disk = @import("on-disk.zig");
const protocol = runtime.vfs_protocol;
const dma = runtime.dma;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
const mount_point = "/mnt/usb";
// The engine's BlockDevice, backed by the `.block` driver plus a DMA bounce
// buffer the driver reads/writes by physical address.
const IpcBlock = struct {
device: runtime.block.Device,
bounce: dma.Region,
fn readBlock(context: *anyopaque, lba: u64, buffer: []u8) bool {
const self: *IpcBlock = @ptrCast(@alignCast(context));
if (!self.device.read(lba, 1, self.bounce.physical)) return false;
const source: [*]const u8 = @ptrFromInt(self.bounce.virtual);
@memcpy(buffer[0..512], source[0..512]);
return true;
}
fn writeBlock(context: *anyopaque, lba: u64, buffer: []const u8) bool {
const self: *IpcBlock = @ptrCast(@alignCast(context));
const destination: [*]u8 = @ptrFromInt(self.bounce.virtual);
@memcpy(destination[0..512], buffer[0..512]);
return self.device.write(lba, 1, self.bounce.physical);
}
};
var ipc_block: IpcBlock = undefined;
var filesystem: engine.FileSystem = undefined;
// Open handles the VFS holds against this backend: each maps a node id to a
// resolved engine node.
const OpenNode = struct { used: bool = false, node: engine.Node = undefined, owner: u32 = 0 };
var open_nodes = [_]OpenNode{.{}} ** 32;
fn allocOpen() ?usize {
for (&open_nodes, 0..) |*o, i| {
if (!o.used) return i;
}
return null;
}
fn openAt(id: u64) ?*OpenNode {
if (id >= open_nodes.len) return null;
const o = &open_nodes[@intCast(id)];
return if (o.used) o else null;
}
fn writeReply(out: []u8, reply: protocol.Reply, payload: []const u8) usize {
@memcpy(out[0..protocol.reply_size], std.mem.asBytes(&reply));
const n = @min(payload.len, out.len - protocol.reply_size);
@memcpy(out[protocol.reply_size..][0..n], payload[0..n]);
return protocol.reply_size + n;
}
fn fail(out: []u8) usize {
return writeReply(out, .{ .status = -1 }, &.{});
}
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = runtime.system.write("/system/services/fat: starting, waiting for a block device\n");
const device = runtime.block.open() orelse {
_ = runtime.system.write("/system/services/fat: no block device (no storage attached)\n");
return false; // clean exit: nothing to serve
};
const geometry = device.geometry() orelse {
_ = runtime.system.write("/system/services/fat: block geometry unavailable\n");
return false;
};
ipc_block = .{ .device = device, .bounce = dma.alloc(4096, dma.coherent) orelse return false };
const block_device = engine.BlockDevice{
.context = &ipc_block,
.block_size = geometry.block_size,
.block_count = geometry.block_count,
.readBlockFn = IpcBlock.readBlock,
.writeBlockFn = IpcBlock.writeBlock,
};
filesystem = engine.FileSystem.mount(block_device) orelse {
_ = runtime.system.write("/system/services/fat: not a FAT filesystem\n");
return false;
};
writeLine("/system/services/fat: mounted FAT ({s}, {d} clusters, partition lba {d})\n", .{ @tagName(filesystem.geometry.fat_type), filesystem.geometry.cluster_count, filesystem.base_lba });
// Mount ourselves into the VFS namespace at /mnt/usb (retry while the VFS
// comes up). From here the VFS routes /mnt/usb/... to this server.
var tries: u32 = 0;
while (tries < 100) : (tries += 1) {
if (runtime.fs.mount(mount_point, endpoint)) {
writeLine("/system/services/fat: mounted {s}\n", .{mount_point});
return true;
}
runtime.system.sleep(50);
}
_ = runtime.system.write("/system/services/fat: could not mount into the VFS\n");
return true; // still serve directly, even if the namespace mount didn't take
}
const ParentLeaf = struct { parent: []const u8, leaf: []const u8 };
// Split a path into its parent directory and final component: "/a/b" -> ("/a",
// "b"); "/b" -> ("/", "b"); "b" -> ("/", "b").
fn splitParent(path: []const u8) ParentLeaf {
const slash = std.mem.lastIndexOfScalar(u8, path, '/');
return .{
.parent = if (slash) |s| (if (s == 0) "/" else path[0..s]) else "/",
.leaf = if (slash) |s| path[s + 1 ..] else path,
};
}
fn handleOpen(out: []u8, path: []const u8, flags: u32) usize {
var node = filesystem.resolve(path);
if (node == null and flags & protocol.create != 0) {
const split = splitParent(path);
const parent = filesystem.resolve(split.parent) orelse return fail(out);
node = filesystem.createFile(parent, split.leaf);
}
var resolved = node orelse return fail(out);
// O_TRUNC: replace an existing file's contents rather than overwriting in place
// (frees the old chain, so a shorter rewrite leaves no stale tail).
if (flags & protocol.truncate != 0 and !resolved.is_directory) {
filesystem.truncate(&resolved);
}
const index = allocOpen() orelse return fail(out);
open_nodes[index] = .{ .used = true, .node = resolved };
return writeReply(out, .{ .status = 0, .node = index }, &.{});
}
fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = capability;
_ = sender;
if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..];
// Stamp create/write with the current wall-clock time (mtime). Cheap, and it
// keeps the engine pure (it takes the time as data, not a syscall).
filesystem.current_time_epoch = runtime.system.wallClock();
switch (request.operation) {
.open => return handleOpen(out, payload[0..@min(payload.len, request.len)], request.flags),
.read => {
const o = openAt(request.node) orelse return fail(out);
var buffer: [protocol.maximum_payload]u8 = undefined;
const want = @min(@as(usize, request.len), buffer.len);
const n = filesystem.readFile(o.node, @intCast(request.offset), buffer[0..want]);
return writeReply(out, .{ .status = 0, .len = @intCast(n) }, buffer[0..n]);
},
.write => {
const o = openAt(request.node) orelse return fail(out);
const data = payload[0..@min(payload.len, request.len)];
const n = filesystem.writeFile(&o.node, @intCast(request.offset), data);
return writeReply(out, .{ .status = 0, .len = @intCast(n) }, &.{});
},
.status => {
const o = openAt(request.node) orelse return fail(out);
const kind: protocol.NodeKind = if (o.node.is_directory) .directory else .regular;
const status = protocol.FileStatus{ .size = o.node.size, .kind = @intFromEnum(kind), .mtime = o.node.mtime };
return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&status));
},
.readdir => {
const o = openAt(request.node) orelse return fail(out);
if (!o.node.is_directory) return writeReply(out, .{ .status = 0, .len = 0 }, &.{});
const listing = filesystem.listEntry(o.node, @intCast(request.offset)) orelse return writeReply(out, .{ .status = 0, .len = 0 }, &.{});
const kind: protocol.NodeKind = if (listing.is_directory) .directory else .regular;
const header = protocol.DirectoryEntry{ .kind = @intFromEnum(kind), .name_len = @intCast(listing.name_len), .size = listing.size };
var buffer: [protocol.maximum_payload]u8 = undefined;
@memcpy(buffer[0..protocol.directory_entry_size], std.mem.asBytes(&header));
const nlen = @min(listing.name_len, buffer.len - protocol.directory_entry_size);
@memcpy(buffer[protocol.directory_entry_size..][0..nlen], listing.name_buffer[0..nlen]);
const total = protocol.directory_entry_size + nlen;
return writeReply(out, .{ .status = 0, .len = @intCast(total) }, buffer[0..total]);
},
.close => {
if (openAt(request.node)) |o| o.used = false;
return writeReply(out, .{ .status = 0 }, &.{});
},
.mkdir => {
const split = splitParent(payload[0..@min(payload.len, request.len)]);
const parent = filesystem.resolve(split.parent) orelse return fail(out);
if (filesystem.createDirectory(parent, split.leaf) == null) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
.unlink => {
const split = splitParent(payload[0..@min(payload.len, request.len)]);
const parent = filesystem.resolve(split.parent) orelse return fail(out);
if (!filesystem.removeFile(parent, split.leaf)) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
.rename => {
const both = payload[0..@min(payload.len, request.len)];
const sep = std.mem.indexOfScalar(u8, both, 0) orelse return fail(out);
const old_split = splitParent(both[0..sep]);
const new_split = splitParent(both[sep + 1 ..]);
// Same-directory rename only.
if (!std.mem.eql(u8, old_split.parent, new_split.parent)) return fail(out);
const parent = filesystem.resolve(old_split.parent) orelse return fail(out);
if (!filesystem.rename(parent, old_split.leaf, new_split.leaf)) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
// A backend is never itself a mount target.
.mount, .unmount => return fail(out),
}
}
pub fn main() void {
runtime.service.run(protocol.message_maximum, .{
.service = .fat,
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+310
View File
@@ -0,0 +1,310 @@
//! The on-disk layout of a FAT filesystem — the boot sector / BIOS Parameter
//! Block, directory entries, long-file-name entries, and the FAT32 FSInfo — as
//! `align(1)` extern structs that bit-cast straight out of a 512-byte sector
//! (multi-byte fields are little-endian, like usb-abi.zig). Pure data, plus the
//! cluster-count FAT-type detection. Host-testable.
const std = @import("std");
/// The BIOS Parameter Block, common to FAT12/16/32 (offset 0..36 of the boot
/// sector). The extended part that follows differs by FAT type.
pub const BiosParameterBlock = extern struct {
jump: [3]u8,
oem_name: [8]u8,
bytes_per_sector: u16 align(1),
sectors_per_cluster: u8,
reserved_sector_count: u16 align(1),
fat_count: u8,
root_entry_count: u16 align(1),
total_sectors_16: u16 align(1),
media: u8,
fat_size_16: u16 align(1),
sectors_per_track: u16 align(1),
head_count: u16 align(1),
hidden_sectors: u32 align(1),
total_sectors_32: u32 align(1),
};
/// The FAT12/16 extended boot record (offset 36).
pub const ExtendedBootRecord16 = extern struct {
drive_number: u8,
reserved: u8,
boot_signature: u8,
volume_id: u32 align(1),
volume_label: [11]u8,
filesystem_type: [8]u8,
};
/// The FAT32 extended boot record (offset 36).
pub const ExtendedBootRecord32 = extern struct {
fat_size_32: u32 align(1),
extended_flags: u16 align(1),
filesystem_version: u16 align(1),
root_cluster: u32 align(1),
filesystem_information_sector: u16 align(1),
backup_boot_sector: u16 align(1),
reserved: [12]u8,
drive_number: u8,
reserved1: u8,
boot_signature: u8,
volume_id: u32 align(1),
volume_label: [11]u8,
filesystem_type: [8]u8,
};
/// A 32-byte directory entry (8.3 short name form).
pub const DirectoryEntry = extern struct {
name: [11]u8, // 8 name + 3 extension, space-padded
attributes: u8,
reserved_nt: u8,
creation_time_tenth: u8,
creation_time: u16 align(1),
creation_date: u16 align(1),
last_access_date: u16 align(1),
first_cluster_high: u16 align(1),
write_time: u16 align(1),
write_date: u16 align(1),
first_cluster_low: u16 align(1),
file_size: u32 align(1),
pub fn firstCluster(self: DirectoryEntry) u32 {
return (@as(u32, self.first_cluster_high) << 16) | self.first_cluster_low;
}
pub fn setFirstCluster(self: *DirectoryEntry, cluster: u32) void {
self.first_cluster_low = @truncate(cluster);
self.first_cluster_high = @truncate(cluster >> 16);
}
pub fn isFree(self: DirectoryEntry) bool {
return self.name[0] == 0x00 or self.name[0] == 0xE5;
}
pub fn isEnd(self: DirectoryEntry) bool {
return self.name[0] == 0x00;
}
pub fn isDirectory(self: DirectoryEntry) bool {
return self.attributes & attribute_directory != 0;
}
pub fn isLongName(self: DirectoryEntry) bool {
return self.attributes & attribute_long_name_mask == attribute_long_name;
}
pub fn isVolumeLabel(self: DirectoryEntry) bool {
return self.attributes & attribute_volume_id != 0 and !self.isLongName();
}
};
/// A 32-byte long-file-name entry (attributes == 0x0F). A sequence of these
/// precedes the 8.3 entry they name, each carrying 13 UTF-16 code units.
pub const LongNameEntry = extern struct {
order: u8,
name1: [5]u16 align(1),
attributes: u8,
kind: u8,
checksum: u8,
name2: [6]u16 align(1),
first_cluster_low: u16 align(1),
name3: [2]u16 align(1),
};
/// The FAT32 FSInfo sector (usually sector 1): advisory free-cluster bookkeeping.
pub const FileSystemInformation = extern struct {
lead_signature: u32 align(1), // 0x41615252
reserved1: [480]u8,
struct_signature: u32 align(1), // 0x61417272
free_count: u32 align(1),
next_free: u32 align(1),
reserved2: [12]u8,
trail_signature: u32 align(1), // 0xAA550000
};
// Directory-entry attribute bits.
pub const attribute_read_only: u8 = 0x01;
pub const attribute_hidden: u8 = 0x02;
pub const attribute_system: u8 = 0x04;
pub const attribute_volume_id: u8 = 0x08;
pub const attribute_directory: u8 = 0x10;
pub const attribute_archive: u8 = 0x20;
pub const attribute_long_name: u8 = 0x0F; // read_only|hidden|system|volume_id
pub const attribute_long_name_mask: u8 = 0x3F;
// FSInfo signatures.
pub const fsinfo_lead_signature: u32 = 0x41615252;
pub const fsinfo_struct_signature: u32 = 0x61417272;
pub const fsinfo_trail_signature: u32 = 0xAA550000;
/// End-of-chain markers (a cluster value >= these ends a chain).
pub const end_of_chain_12: u32 = 0xFF8;
pub const end_of_chain_16: u32 = 0xFFF8;
pub const end_of_chain_32: u32 = 0x0FFFFFF8;
pub const bad_cluster_32: u32 = 0x0FFFFFF7;
pub const free_cluster: u32 = 0;
pub const boot_signature_offset: usize = 510; // 0x55 0xAA at the end of the boot sector
pub const FatType = enum { fat12, fat16, fat32 };
/// The geometry derived from the BPB, plus the FAT type (by the Microsoft
/// cluster-count rule: <4085 FAT12, <65525 FAT16, else FAT32).
pub const Geometry = struct {
fat_type: FatType,
bytes_per_sector: u32,
sectors_per_cluster: u32,
reserved_sector_count: u32,
fat_count: u32,
fat_size_sectors: u32, // per FAT
root_entry_count: u32, // FAT12/16
root_cluster: u32, // FAT32
first_data_sector: u32,
total_sectors: u32,
cluster_count: u32,
fsinfo_sector: u32, // FAT32
};
/// Derive the geometry (and FAT type) from a boot sector's first 512 bytes.
/// Returns null if the sector is not a plausible FAT boot sector.
pub fn geometryOf(sector: []const u8) ?Geometry {
if (sector.len < 512) return null;
if (sector[boot_signature_offset] != 0x55 or sector[boot_signature_offset + 1] != 0xAA) return null;
const bpb = std.mem.bytesToValue(BiosParameterBlock, sector[0..@sizeOf(BiosParameterBlock)]);
if (bpb.bytes_per_sector == 0 or bpb.sectors_per_cluster == 0 or bpb.fat_count == 0) return null;
const fat_size_16: u32 = bpb.fat_size_16;
var fat_size: u32 = fat_size_16;
var root_cluster: u32 = 0;
var fsinfo_sector: u32 = 0;
if (fat_size_16 == 0) {
const ebr = std.mem.bytesToValue(ExtendedBootRecord32, sector[36 .. 36 + @sizeOf(ExtendedBootRecord32)]);
fat_size = ebr.fat_size_32;
root_cluster = ebr.root_cluster;
fsinfo_sector = ebr.filesystem_information_sector;
}
const total_sectors: u32 = if (bpb.total_sectors_16 != 0) bpb.total_sectors_16 else bpb.total_sectors_32;
const root_dir_sectors = (@as(u32, bpb.root_entry_count) * 32 + bpb.bytes_per_sector - 1) / bpb.bytes_per_sector;
const first_data_sector = bpb.reserved_sector_count + bpb.fat_count * fat_size + root_dir_sectors;
if (total_sectors < first_data_sector) return null;
const data_sectors = total_sectors - first_data_sector;
const cluster_count = data_sectors / bpb.sectors_per_cluster;
const fat_type: FatType = if (cluster_count < 4085) .fat12 else if (cluster_count < 65525) .fat16 else .fat32;
return .{
.fat_type = fat_type,
.bytes_per_sector = bpb.bytes_per_sector,
.sectors_per_cluster = bpb.sectors_per_cluster,
.reserved_sector_count = bpb.reserved_sector_count,
.fat_count = bpb.fat_count,
.fat_size_sectors = fat_size,
.root_entry_count = bpb.root_entry_count,
.root_cluster = root_cluster,
.first_data_sector = first_data_sector,
.total_sectors = total_sectors,
.cluster_count = cluster_count,
.fsinfo_sector = fsinfo_sector,
};
}
// --- DOS date/time <-> Unix epoch --------------------------------------------
//
// FAT stamps a file's modification time as two 16-bit DOS fields. There is no
// timezone, so danos treats them as UTC. `date`: year-1980(7)|month(4)|day(5);
// `time`: hour(5)|minute(6)|(second/2)(5).
fn isLeapYear(year: u32) bool {
return (year % 4 == 0 and year % 100 != 0) or (year % 400 == 0);
}
const days_in_month = [_]u8{ 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31 };
/// Convert a FAT date+time to Unix epoch seconds (UTC). Returns 0 for an unset
/// (zero) date.
pub fn fatToEpoch(date: u16, time: u16) u64 {
if (date == 0) return 0;
const day: u32 = date & 0x1F;
const month: u32 = (date >> 5) & 0x0F;
const year: u32 = 1980 + (date >> 9);
if (month < 1 or month > 12 or day < 1) return 0;
const second: u32 = @as(u32, time & 0x1F) * 2;
const minute: u32 = (time >> 5) & 0x3F;
const hour: u32 = (time >> 11) & 0x1F;
var days: u64 = 0;
var y: u32 = 1970;
while (y < year) : (y += 1) days += if (isLeapYear(y)) 366 else 365;
var m: u32 = 1;
while (m < month) : (m += 1) {
days += days_in_month[m - 1];
if (m == 2 and isLeapYear(year)) days += 1;
}
days += day - 1;
return ((days * 24 + hour) * 60 + minute) * 60 + second;
}
pub const FatDateTime = struct { date: u16, time: u16 };
/// Convert Unix epoch seconds (UTC) to a FAT date+time. Returns {0,0} for epoch 0 or
/// any time before 1980 (which DOS cannot represent).
pub fn epochToFatDateTime(epoch: u64) FatDateTime {
if (epoch == 0) return .{ .date = 0, .time = 0 };
var remaining = epoch;
const second: u32 = @intCast(remaining % 60);
remaining /= 60;
const minute: u32 = @intCast(remaining % 60);
remaining /= 60;
const hour: u32 = @intCast(remaining % 24);
remaining /= 24;
var days: u32 = @intCast(remaining); // whole days since 1970-01-01
var year: u32 = 1970;
while (true) {
const y_days: u32 = if (isLeapYear(year)) 366 else 365;
if (days < y_days) break;
days -= y_days;
year += 1;
}
if (year < 1980) return .{ .date = 0, .time = 0 };
var month: u32 = 1;
while (true) {
var m_days: u32 = days_in_month[month - 1];
if (month == 2 and isLeapYear(year)) m_days += 1;
if (days < m_days) break;
days -= m_days;
month += 1;
}
const day = days + 1;
return .{
.date = @intCast(((year - 1980) << 9) | (month << 5) | day),
.time = @intCast((hour << 11) | (minute << 5) | (second / 2)),
};
}
test "FAT date/time <-> Unix epoch round trip" {
// Even-second UTC times (FAT stores seconds/2, so even seconds round-trip exactly).
for ([_]u64{ 1_577_836_800, 1_700_000_000, 1_262_304_000, 1_783_971_244 }) |epoch| {
const fat = epochToFatDateTime(epoch);
try std.testing.expectEqual(epoch, fatToEpoch(fat.date, fat.time));
}
// Absolute check: 1577836800 is 2020-01-01 00:00:00 UTC.
const y2020 = epochToFatDateTime(1_577_836_800);
try std.testing.expectEqual(@as(u16, 2020), 1980 + (y2020.date >> 9));
try std.testing.expectEqual(@as(u16, 1), (y2020.date >> 5) & 0x0F); // month
try std.testing.expectEqual(@as(u16, 1), y2020.date & 0x1F); // day
// 0 is "unset" both ways.
try std.testing.expectEqual(@as(u64, 0), fatToEpoch(0, 0));
try std.testing.expectEqual(@as(u16, 0), epochToFatDateTime(0).date);
}
test "on-disk struct sizes match the specification" {
try std.testing.expectEqual(@as(usize, 36), @sizeOf(BiosParameterBlock));
try std.testing.expectEqual(@as(usize, 26), @sizeOf(ExtendedBootRecord16));
try std.testing.expectEqual(@as(usize, 54), @sizeOf(ExtendedBootRecord32));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(DirectoryEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(LongNameEntry));
try std.testing.expectEqual(@as(usize, 512), @sizeOf(FileSystemInformation));
}
test "directory entry cluster split/join" {
var entry = std.mem.zeroes(DirectoryEntry);
entry.setFirstCluster(0x01234567);
try std.testing.expectEqual(@as(u16, 0x4567), entry.first_cluster_low);
try std.testing.expectEqual(@as(u16, 0x0123), entry.first_cluster_high);
try std.testing.expectEqual(@as(u32, 0x01234567), entry.firstCluster());
}
+42 -4
View File
@@ -21,11 +21,16 @@ const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const power = runtime.power_protocol; const power = runtime.power_protocol;
/// Where the kernel boot log is persisted on the USB FAT volume — an 8.3 name at
/// the mount root (see system/services/log-flush). init writes it at shutdown;
/// the log-flush one-shot writes it once at boot.
const log_path = "/mnt/usb/DANOS.LOG";
/// The system services init brings up at boot, in order. This is init's policy — the /// The system services init brings up at boot, in order. This is init's policy — the
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent /// microkernel keeps such choices in user space, not the kernel. Drivers are absent
/// on purpose: the device manager owns those. (A future init reads this from a /// on purpose: the device manager owns those. (A future init reads this from a
/// manifest under /system/services instead of a hardcoded list.) /// manifest under /system/services instead of a hardcoded list.)
const boot_services = [_][]const u8{ "vfs", "input", "device-manager" }; const boot_services = [_][]const u8{ "vfs", "input", "device-manager", "fat" };
var children: [boot_services.len]u32 = .{0} ** boot_services.len; var children: [boot_services.len]u32 = .{0} ** boot_services.len;
var child_count: usize = 0; var child_count: usize = 0;
@@ -65,6 +70,14 @@ pub fn main() void {
} }
} }
// Once the storage stack is up, a one-shot copies the boot log to the USB
// volume (/mnt/usb/DANOS.LOG) so it can be read on another machine — the only
// way to see it on a headless/real board with no host capturing serial. Fire
// and forget: it polls for the mount itself, and is deliberately NOT one of
// init's supervised children (a transient one-shot must not be stopped-and-
// waited-for during shutdown).
_ = runtime.system.spawn("log-flush");
// Subscribe to power events (retry: the power service registers well after // Subscribe to power events (retry: the power service registers well after
// init starts). Best-effort — without it, a `terminate` signal still // init starts). Best-effort — without it, a `terminate` signal still
// triggers the same shutdown path. // triggers the same shutdown path.
@@ -115,11 +128,36 @@ fn subscribePower() void {
_ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {}; _ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {};
} }
/// The stop sequence: terminate each child in reverse spawn order (vfs last — /// Copy the whole kernel log to /mnt/usb/DANOS.LOG (the same file log-flush
/// other services may flush through it), waiting up to a deadline for each to /// writes at boot), so a poweroff captures the fullest log. Best-effort: if the
/// exit before killing it, then ask the power service to enter S5. /// USB volume is not mounted, the open fails and it does nothing. Must run while
/// the storage services are still alive (see shutDown).
fn flushKernelLog() void {
// Truncate on open so this fuller flush replaces the boot-time one cleanly.
var file = runtime.fs.open(log_path, .{ .create = true, .truncate = true }) orelse return; // no USB volume
defer file.close();
var chunk: [4096]u8 = undefined;
var offset: usize = 0;
while (true) {
const got = runtime.system.klogRead(offset, &chunk);
if (got == 0) break; // reached the end of the accumulated log
if (file.writeAll(chunk[0..got]) == null) break; // storage went away
offset += got;
}
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, "/system/services/init: flushed log to {s} ({d} bytes)\n", .{ log_path, offset }) catch "");
}
/// The stop sequence: persist the log while storage is still up, then terminate
/// each child in reverse spawn order (vfs last — other services may flush through
/// it), waiting up to a deadline for each to exit before killing it, then ask the
/// power service to enter S5.
fn shutDown() void { fn shutDown() void {
_ = runtime.system.write("/system/services/init: shutting down\n"); _ = runtime.system.write("/system/services/init: shutting down\n");
// Persist the fullest log to the USB volume BEFORE tearing anything down: the
// reverse-order stop loop below kills the fat server (children[3]) first, so
// /mnt/usb must be written while it is still mounted.
flushKernelLog();
var i = child_count; var i = child_count;
while (i > 0) { while (i > 0) {
i -= 1; i -= 1;
+67
View File
@@ -0,0 +1,67 @@
//! system/services/log-flush — a one-shot that copies the kernel's in-memory
//! diagnostic log to a file on the mounted USB FAT volume, so the boot log
//! survives to be read on another machine. On a headless or real board there is
//! no host capturing serial, so without this the log is lost at power-off; this
//! is the on-disk equivalent of QEMU's `-serial file:`.
//!
//! It reads the whole kernel log back through `klog_read` (the RAM sink in
//! system/kernel/log.zig) and writes it to /mnt/usb/DANOS.LOG. The name is 8.3
//! (FAT short-name rule: base <= 8, extension <= 3) and lives at the mount root
//! (there is no mkdir on the FAT path yet). init spawns this once the boot
//! services are up; init itself repeats the flush at shutdown for a fuller log.
//!
//! If no USB volume is mounted — no stick, or the initial-ramdisk sweep that
//! spawns every bundled binary bare with no VFS — it waits briefly, then exits
//! silently, deranging no other test's output.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
const log_path = "/mnt/usb/DANOS.LOG";
/// Copy the whole kernel log to the open file, looping klog_read -> write until
/// the log is exhausted. Returns the number of bytes written.
fn drainKernelLog(file: *fs.File) usize {
var chunk: [4096]u8 = undefined;
var offset: usize = 0;
while (true) {
const got = runtime.system.klogRead(offset, &chunk);
if (got == 0) break; // reached the end of the accumulated log
if (file.writeAll(chunk[0..got]) == null) break; // storage went away
offset += got;
}
return offset;
}
pub fn main() void {
// Wait for the fat server to mount /mnt/usb (it must bring up the whole USB
// storage chain first, so it races us at boot). Bounded: if the mount never
// appears — no volume, or the no-VFS ramdisk sweep — give up silently.
var ready = false;
var tries: u32 = 0;
while (tries < 1400) : (tries += 1) {
if (fs.openDirectory("/mnt/usb")) |directory| {
var dir = directory;
dir.close();
ready = true;
break;
}
runtime.system.sleep(50);
}
if (!ready) return; // /mnt/usb never became available — nothing to persist to
// Truncate on open: each flush replaces the file, so a shorter log on a later
// boot of the same stick leaves no stale tail from a previous, longer one.
var file = fs.open(log_path, .{ .create = true, .truncate = true }) orelse return;
const written = drainKernelLog(&file);
file.close();
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, "log-flush: wrote {d} bytes to {s}\n", .{ written, log_path }) catch return);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+39
View File
@@ -0,0 +1,39 @@
//! Pure path utilities for the VFS mount router — no IPC, no state, so they are
//! host-testable in isolation. The router uses these to decide whether an opened
//! path lies under a mount point and, if so, what it looks like relative to that
//! mount.
const std = @import("std");
/// If `path` lies under `mount_prefix` — equal to it, or the prefix followed by a
/// path separator — return the path relative to the mount ("/" for an exact
/// match, otherwise the tail beginning with '/'). Returns null when `path` is not
/// under the mount, so a prefix like "/mnt/usb" never captures "/mnt/usbextra".
pub fn underMount(path: []const u8, mount_prefix: []const u8) ?[]const u8 {
if (path.len < mount_prefix.len) return null;
if (!std.mem.eql(u8, path[0..mount_prefix.len], mount_prefix)) return null;
if (path.len == mount_prefix.len) return "/";
if (path[mount_prefix.len] != '/') return null;
return path[mount_prefix.len..];
}
/// Whether `path` is absolute (rooted at '/'). Bare names — what the flat ramfs
/// uses — are relative and never route through a mount.
pub fn isAbsolute(path: []const u8) bool {
return path.len > 0 and path[0] == '/';
}
test "underMount matches only at path boundaries" {
try std.testing.expectEqualStrings("/", underMount("/mnt/usb", "/mnt/usb").?);
try std.testing.expectEqualStrings("/system/kernel", underMount("/mnt/usb/system/kernel", "/mnt/usb").?);
try std.testing.expect(underMount("/mnt/usbextra", "/mnt/usb") == null); // not a boundary
try std.testing.expect(underMount("/mnt", "/mnt/usb") == null); // shorter than the prefix
try std.testing.expect(underMount("/other", "/mnt/usb") == null);
try std.testing.expect(underMount("greeting", "/mnt/usb") == null); // a bare name
}
test "isAbsolute distinguishes paths from bare names" {
try std.testing.expect(isAbsolute("/mnt/usb"));
try std.testing.expect(!isAbsolute("greeting"));
try std.testing.expect(!isAbsolute(""));
}
+60 -5
View File
@@ -4,12 +4,11 @@
//! header followed by an inline payload (read bytes, or a FileStatus). Everything fits //! header followed by an inline payload (read bytes, or a FileStatus). Everything fits
//! in one IPC message (<= ipc MESSAGE_MAXIMUM = 256 bytes). //! in one IPC message (<= ipc MESSAGE_MAXIMUM = 256 bytes).
//! //!
//! This is a danos-native contract, so it uses danos names throughout — the POSIX //! This is a danos-native contract, so it uses danos names throughout. The client
//! spellings (`stat`, `O_CREAT`, ...) live only in the POSIX layer //! side is `runtime.fs` (library/runtime/fs.zig), which programs use directly.
//! (library/posix/unistd.zig), which translates to these.
//! //!
//! This is user-space only — the kernel knows nothing of files or paths; it only moves the bytes. //! This is user-space only — the kernel knows nothing of files or paths; it only moves the bytes.
//! Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server). //! Shared by library/runtime/fs.zig (client) and system/services/vfs/vfs.zig (server).
pub const Operation = enum(u32) { pub const Operation = enum(u32) {
open, // open(path) -> node id open, // open(path) -> node id
@@ -17,8 +16,43 @@ pub const Operation = enum(u32) {
read, // read(node, offset, len) -> bytes read, // read(node, offset, len) -> bytes
write, // write(node, offset, bytes) -> count write, // write(node, offset, bytes) -> count
status, // status(node) -> FileStatus status, // status(node) -> FileStatus
// Appended for the mount router (M5). Values stay stable, so existing clients
// and the flat-ramfs tests are unaffected.
readdir, // readdir(dir_node, cursor=offset) -> one DirectoryEntry (len==0 => EOF)
mount, // mount(prefix payload, capability = backend endpoint)
unmount, // unmount(prefix payload)
// Appended for filesystem mutation (Phase 2). Path-based (the path is the
// payload); a mounted backend handles them, the flat ramfs refuses them.
mkdir, // mkdir(path payload) -> status
unlink, // unlink(path payload) -> status
// rename: the payload is the old path, a single 0x00 separator, then the new
// path. Same-directory rename only (the router requires both under one mount).
rename, // rename(old\0new payload) -> status
}; };
/// The type of a filesystem node, aligned to the FSH file-type table
/// (docs/danos-file-system-hierarchy-FSH.md). Fills `FileStatus.kind` and
/// `DirectoryEntry.kind`; `regular = 0` keeps the historical hardcoded value.
pub const NodeKind = enum(u32) {
regular = 0,
directory = 1,
character_device = 2,
block_device = 3,
symbolic_link = 4,
fifo = 5,
socket = 6,
};
/// One directory entry, returned by `readdir`: a fixed header followed inline in
/// the reply payload by `name_len` bytes of name. A zero-length reply is EOF.
pub const DirectoryEntry = extern struct {
kind: u32, // a NodeKind
name_len: u32,
size: u64,
};
pub const directory_entry_size: usize = @sizeOf(DirectoryEntry);
/// Request header. `node` is the server-side open-file id (from a prior open); /// Request header. `node` is the server-side open-file id (from a prior open);
/// for `open` the path is the payload and `len` is its length. `offset`/`len` /// for `open` the path is the payload and `len` is its length. `offset`/`len`
/// carry the read/write position and count. /// carry the read/write position and count.
@@ -47,6 +81,9 @@ pub const FileStatus = extern struct {
size: u64, size: u64,
kind: u32, kind: u32,
_padding: u32 = 0, _padding: u32 = 0,
/// Modification time — Unix epoch seconds, UTC. 0 if the backend has none (the
/// flat ramfs). Filled from the FAT directory entry's write date/time.
mtime: u64 = 0,
}; };
pub const message_maximum: usize = 256; pub const message_maximum: usize = 256;
@@ -55,5 +92,23 @@ pub const reply_size: usize = @sizeOf(Reply);
/// Largest inline payload that still fits one IPC message alongside a header. /// Largest inline payload that still fits one IPC message alongside a header.
pub const maximum_payload: usize = message_maximum - request_size; pub const maximum_payload: usize = message_maximum - request_size;
/// Open flags (danos-native; the POSIX layer maps `O_CREAT` onto `create`). /// Open flags (danos-native; `runtime.fs.OpenOptions` maps its booleans onto these).
pub const create: u32 = 1; pub const create: u32 = 1;
/// Open a directory (for readdir) rather than a file. A mounted backend uses
/// this to open a directory node; the flat ramfs ignores it.
pub const directory: u32 = 2;
/// Truncate the file to zero length on open (O_TRUNC): replace its contents rather
/// than overwriting in place, so a shorter new file leaves no stale tail. A mounted
/// backend frees the old cluster chain; the flat ramfs ignores it.
pub const truncate: u32 = 4;
test "protocol struct sizes and node kinds" {
const std = @import("std");
try std.testing.expectEqual(@as(u32, 0), @intFromEnum(NodeKind.regular));
try std.testing.expectEqual(@as(u32, 1), @intFromEnum(NodeKind.directory));
try std.testing.expectEqual(@as(usize, 16), @sizeOf(DirectoryEntry));
// The appended operations keep the original values.
try std.testing.expectEqual(@as(u32, 0), @intFromEnum(Operation.open));
try std.testing.expectEqual(@as(u32, 4), @intFromEnum(Operation.status));
try std.testing.expectEqual(@as(u32, 5), @intFromEnum(Operation.readdir));
}
+18 -18
View File
@@ -1,26 +1,26 @@
//! /system/services/vfs/vfs-test — a client that proves the VFS round trip end to end: open a //! /system/services/vfs/vfs-test — a client that proves the VFS round trip end to end: open a
//! file through the `runtime` file API, write to it, seek back, read it, and compare. //! file through the `runtime.fs` file API, write to it, seek back, read it, and compare.
//! On success it heartbeats "vfstest: ok" so the kernel test can observe it; //! On success it heartbeats "vfstest: ok" so the kernel test can observe it;
//! on failure it reports what went wrong. Shipped in the initial_ramdisk alongside vfs. //! on failure it reports what went wrong. Shipped in the initial_ramdisk alongside vfs.
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const fs = runtime.fs;
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
const u = @import("posix").unistd;
const payload = "hello-vfs"; const payload = "hello-vfs";
// The "park" role (the vfs-client-death test): open a file, then hold the // The "park" role (the vfs-client-death test): open a file, then hold the
// handle forever without closing — the kill and the VFS's release-on-death // handle forever without closing — the kill and the VFS's release-on-death
// are the point. // are the point.
if (init.arguments.count > 1) { if (init.arguments.count > 1) {
var fd: i32 = -1; var parked: ?fs.File = null;
var tries: u32 = 0; var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) { while (parked == null and tries < 200) : (tries += 1) {
fd = u.open("parked", u.O_CREAT); parked = fs.open("parked", .{ .create = true });
if (fd < 0) runtime.system.sleep(20); if (parked == null) runtime.system.sleep(20);
} }
if (fd < 0) { if (parked == null) {
_ = runtime.system.write("vfstest: park open failed\n"); _ = runtime.system.write("vfstest: park open failed\n");
return; return;
} }
@@ -31,28 +31,28 @@ pub fn main(init: runtime.process.Init) void {
} }
// The VFS server may not have registered yet — retry open until it's up. // The VFS server may not have registered yet — retry open until it's up.
var fd: i32 = -1; var opened: ?fs.File = null;
var tries: u32 = 0; var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) { while (opened == null and tries < 200) : (tries += 1) {
fd = u.open("greeting", u.O_CREAT); opened = fs.open("greeting", .{ .create = true });
if (fd < 0) runtime.system.sleep(20); if (opened == null) runtime.system.sleep(20);
} }
if (fd < 0) { var greeting = opened orelse {
_ = runtime.system.write("vfstest: open failed\n"); _ = runtime.system.write("vfstest: open failed\n");
return; return;
} };
if (u.write(fd, payload) != @as(isize, payload.len)) { if ((greeting.write(payload) orelse 0) != payload.len) {
_ = runtime.system.write("vfstest: write failed\n"); _ = runtime.system.write("vfstest: write failed\n");
return; return;
} }
_ = u.lseek(fd, 0, u.SEEK_SET); greeting.seekTo(0);
var buffer: [32]u8 = undefined; var buffer: [32]u8 = undefined;
const n = u.read(fd, &buffer); const n = greeting.read(&buffer) orelse 0;
u.close(fd); greeting.close();
if (n == @as(isize, payload.len) and std.mem.eql(u8, buffer[0..@intCast(n)], payload)) { if (n == payload.len and std.mem.eql(u8, buffer[0..n], payload)) {
while (true) { while (true) {
_ = runtime.system.write("vfstest: ok\n"); _ = runtime.system.write("vfstest: ok\n");
runtime.system.sleep(1000); runtime.system.sleep(1000);
+242 -16
View File
@@ -3,14 +3,24 @@
//! file API marshals open/read/write/stat/close into calls to this server's //! file API marshals open/read/write/stat/close into calls to this server's
//! endpoint, published under the well-known `vfs` service id). //! endpoint, published under the well-known `vfs` service id).
//! //!
//! For now the namespace is a small in-memory ramfs (opening a name creates it): //! Two namespaces meet here (M5):
//! enough to prove the whole path — client file API -> IPC -> server dispatch -> //! - a small in-memory **ramfs** — opening a bare name creates it — enough to
//! reply. Device nodes backed by user-space drivers (/device) layer on top in M10, //! prove the round trip and to back the existing tests;
//! where `open` on a /device name forwards to the owning driver's endpoint. //! - **mounted filesystems**: a mount table maps an absolute path prefix (e.g.
//! `/mnt/usb`) to a backend server's endpoint. An open of a path under a mount
//! is *forwarded* to that backend (which speaks this same protocol), and every
//! later read/write/status/readdir/close on the resulting handle is relayed to
//! it. The VFS is the router; a filesystem (FAT) is the backend.
//!
//! A path routes through a mount only when it is absolute and lies under a mount
//! prefix; bare names always resolve in the flat ramfs — the backward-compat
//! contract the `vfs` / `vfs-client-death` tests rely on.
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const protocol = runtime.vfs_protocol; const protocol = runtime.vfs_protocol;
const path = @import("path.zig");
const ipc = runtime.ipc;
const Node = struct { const Node = struct {
used: bool = false, used: bool = false,
@@ -22,15 +32,29 @@ const Node = struct {
const OpenFile = struct { const OpenFile = struct {
used: bool = false, used: bool = false,
// For a local handle: an index into `nodes`. For a forwarding handle: the
// node id the backend returned. (usize == u64 here, so it holds either.)
node: usize = 0, node: usize = 0,
// Non-null for a handle that forwards to a mounted backend.
backend: ?ipc.Handle = null,
// The client (task id — an IPC badge is one) that opened this handle. What // The client (task id — an IPC badge is one) that opened this handle. What
// release-on-death sweeps by: a service must never depend on its clients // release-on-death sweeps by: a service must never depend on its clients
// cleaning up after themselves (docs/process-lifecycle.md). // cleaning up after themselves (docs/process-lifecycle.md).
owner: u32 = 0, owner: u32 = 0,
}; };
// One mounted filesystem: an absolute path prefix and the backend endpoint that
// serves everything under it.
const Mount = struct {
used: bool = false,
prefix: [64]u8 = undefined,
prefix_len: usize = 0,
backend: ipc.Handle = 0,
};
var nodes = [_]Node{.{}} ** 8; var nodes = [_]Node{.{}} ** 8;
var opens = [_]OpenFile{.{}} ** 16; var opens = [_]OpenFile{.{}} ** 16;
var mounts = [_]Mount{.{}} ** 8;
fn findNode(name: []const u8) ?usize { fn findNode(name: []const u8) ?usize {
for (&nodes, 0..) |*n, i| { for (&nodes, 0..) |*n, i| {
@@ -57,6 +81,26 @@ fn openAt(id: u64) ?*OpenFile {
return if (o.used) o else null; return if (o.used) o else null;
} }
/// The mount whose prefix most specifically contains `name`, and the path
/// relative to it. Only absolute paths route; bare names never match.
const MountMatch = struct { backend: ipc.Handle, relative: []const u8 };
fn longestMount(name: []const u8) ?MountMatch {
if (!path.isAbsolute(name)) return null;
var best: ?MountMatch = null;
var best_len: usize = 0;
for (&mounts) |*m| {
if (!m.used) continue;
const prefix = m.prefix[0..m.prefix_len];
if (path.underMount(name, prefix)) |relative| {
if (best == null or prefix.len >= best_len) {
best_len = prefix.len;
best = .{ .backend = m.backend, .relative = relative };
}
}
}
return best;
}
/// Serialise a reply header + payload into `out`; returns the total length. /// Serialise a reply header + payload into `out`; returns the total length.
fn writeReply(out: []u8, reply: protocol.Reply, payload: []const u8) usize { fn writeReply(out: []u8, reply: protocol.Reply, payload: []const u8) usize {
@memcpy(out[0..protocol.reply_size], std.mem.asBytes(&reply)); @memcpy(out[0..protocol.reply_size], std.mem.asBytes(&reply));
@@ -76,13 +120,134 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return); _ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
} }
// --- mount routing ----------------------------------------------------------
/// Forward an open under a mount to its backend and, on success, allocate a local
/// forwarding handle that remembers the backend's node id.
fn forwardOpen(out: []u8, backend: ipc.Handle, relative: []const u8, flags: u32, sender: u32) usize {
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = flags };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const rel = relative[0..@min(relative.len, protocol.maximum_payload)];
@memcpy(message[protocol.request_size..][0..rel.len], rel);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + rel.len], &reply) catch return fail(out);
if (n < protocol.reply_size) return fail(out);
const backend_reply = std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]);
if (backend_reply.status != 0) return writeReply(out, .{ .status = backend_reply.status }, &.{});
for (&opens, 0..) |*o, i| {
if (!o.used) {
o.* = .{ .used = true, .node = @intCast(backend_reply.node), .backend = backend, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{});
}
}
return fail(out);
}
/// Relay a read/write/status/readdir/close on a forwarding handle to the backend
/// (the node already rewritten to the backend's id) and copy its reply out.
fn forwardRequest(out: []u8, backend: ipc.Handle, request: protocol.Request, payload: []const u8) usize {
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const plen = @min(payload.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..plen], payload[0..plen]);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + plen], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Forward a path-based operation (mkdir, unlink) under a mount to its backend and
/// relay the reply. No handle is created — these operate by path and return only a
/// status.
fn forwardPath(out: []u8, backend: ipc.Handle, operation: protocol.Operation, relative: []const u8) usize {
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = 0 };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const rel = relative[0..@min(relative.len, protocol.maximum_payload)];
@memcpy(message[protocol.request_size..][0..rel.len], rel);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + rel.len], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Forward a rename to its backend: the payload is the mount-relative old path, a
/// 0x00 separator, then the mount-relative new path. Relays the backend's reply.
fn forwardRename(out: []u8, backend: ipc.Handle, old_relative: []const u8, new_relative: []const u8) usize {
const total = old_relative.len + 1 + new_relative.len;
var message: [protocol.message_maximum]u8 = undefined;
if (protocol.request_size + total > message.len) return fail(out);
const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 };
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
var p = protocol.request_size;
@memcpy(message[p..][0..old_relative.len], old_relative);
p += old_relative.len;
message[p] = 0;
p += 1;
@memcpy(message[p..][0..new_relative.len], new_relative);
p += new_relative.len;
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0..p], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Best-effort close of a backend node (used when a dead client's forwarding
/// handles are swept — the backend must not leak the vfs's opens).
fn forwardClose(backend: ipc.Handle, backend_node: u64) void {
const request = protocol.Request{ .operation = .close, .node = backend_node, .offset = 0, .len = 0, .flags = 0 };
var reply: [protocol.message_maximum]u8 = undefined;
_ = ipc.call(backend, std.mem.asBytes(&request), &reply) catch {};
}
fn doMount(out: []u8, prefix: []const u8, backend: ipc.Handle) usize {
for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefix[0..m.prefix_len], prefix)) {
m.backend = backend;
writeLine("/system/services/vfs: remounted {s}\n", .{prefix});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
for (&mounts) |*m| {
if (!m.used) {
const l = @min(prefix.len, m.prefix.len);
m.used = true;
@memcpy(m.prefix[0..l], prefix[0..l]);
m.prefix_len = l;
m.backend = backend;
writeLine("/system/services/vfs: mounted {s}\n", .{prefix[0..l]});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
return fail(out);
}
fn doUnmount(out: []u8, prefix: []const u8) usize {
for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefix[0..m.prefix_len], prefix)) {
m.used = false;
writeLine("/system/services/vfs: unmounted {s}\n", .{prefix});
return writeReply(out, .{ .status = 0 }, &.{});
}
}
return fail(out);
}
/// Release every open handle `client` held — called on that client's published /// Release every open handle `client` held — called on that client's published
/// exit event. The nodes (the files) stay: ramfs contents outlive their writers, /// exit event. Forwarding handles also tell their backend to release; local
/// only the dead client's handles go. /// nodes (the ramfs files) stay, since ramfs contents outlive their writers.
fn releaseClientHandles(client: u32) void { fn releaseClientHandles(client: u32) void {
var released: u32 = 0; var released: u32 = 0;
for (&opens) |*o| { for (&opens) |*o| {
if (o.used and o.owner == client) { if (o.used and o.owner == client) {
if (o.backend) |backend| forwardClose(backend, o.node);
o.used = false; o.used = false;
released += 1; released += 1;
} }
@@ -91,19 +256,32 @@ fn releaseClientHandles(client: u32) void {
} }
/// Handle one request from `sender`; write the reply into `out`, return its length. /// Handle one request from `sender`; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize { fn handle(message: []const u8, out: []u8, sender: u32, capability: ?ipc.Handle) usize {
_ = capability;
if (message.len < protocol.request_size) return fail(out); if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]); const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..]; const payload = message[protocol.request_size..];
switch (request.operation) { switch (request.operation) {
.mount => {
const prefix = payload[0..@min(payload.len, request.len)];
const backend = capability orelse return fail(out);
return doMount(out, prefix, backend);
},
.unmount => {
const prefix = payload[0..@min(payload.len, request.len)];
return doUnmount(out, prefix);
},
.open => { .open => {
const name = payload[0..@min(payload.len, request.len)]; const name = payload[0..@min(payload.len, request.len)];
if (longestMount(name)) |m| return forwardOpen(out, m.backend, m.relative, request.flags, sender);
// An absolute path with no matching mount is simply not found — only
// bare names live in the flat ramfs. (Else /mnt/usb would be silently
// created as a flat file when its filesystem is not yet mounted.)
if (path.isAbsolute(name)) return fail(out);
const ni = findNode(name) orelse createNode(name) orelse return fail(out); const ni = findNode(name) orelse createNode(name) orelse return fail(out);
for (&opens, 0..) |*o, i| { for (&opens, 0..) |*o, i| {
if (!o.used) { if (!o.used) {
o.* = .{ .used = true, .node = ni, .owner = sender }; o.* = .{ .used = true, .node = ni, .backend = null, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{}); return writeReply(out, .{ .status = 0, .node = i }, &.{});
} }
} }
@@ -111,7 +289,12 @@ fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.
}, },
.read => { .read => {
const of = openAt(request.node) orelse return fail(out); const of = openAt(request.node) orelse return fail(out);
const nd = &nodes[of.node]; if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const nd = &nodes[@intCast(of.node)];
const off: usize = @intCast(request.offset); const off: usize = @intCast(request.offset);
if (off >= nd.size) return writeReply(out, .{ .status = 0, .len = 0 }, &.{}); // EOF if (off >= nd.size) return writeReply(out, .{ .status = 0, .len = 0 }, &.{}); // EOF
const n = @min(@min(nd.size - off, request.len), protocol.maximum_payload); const n = @min(@min(nd.size - off, request.len), protocol.maximum_payload);
@@ -119,7 +302,12 @@ fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.
}, },
.write => { .write => {
const of = openAt(request.node) orelse return fail(out); const of = openAt(request.node) orelse return fail(out);
const nd = &nodes[of.node]; if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const nd = &nodes[@intCast(of.node)];
const off: usize = @intCast(request.offset); const off: usize = @intCast(request.offset);
if (off > nd.data.len) return fail(out); if (off > nd.data.len) return fail(out);
const n = @min(@min(payload.len, request.len), nd.data.len - off); const n = @min(@min(payload.len, request.len), nd.data.len - off);
@@ -129,20 +317,58 @@ fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.
}, },
.status => { .status => {
const of = openAt(request.node) orelse return fail(out); const of = openAt(request.node) orelse return fail(out);
const st = protocol.FileStatus{ .size = nodes[of.node].size, .kind = 0 }; if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
const st = protocol.FileStatus{ .size = nodes[@intCast(of.node)].size, .kind = @intFromEnum(protocol.NodeKind.regular) };
return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&st)); return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&st));
}, },
.readdir => {
const of = openAt(request.node) orelse return fail(out);
if (of.backend) |backend| {
var forwarded = request;
forwarded.node = of.node;
return forwardRequest(out, backend, forwarded, payload);
}
// The flat ramfs has no directories: report EOF.
return writeReply(out, .{ .status = 0, .len = 0 }, &.{});
},
.close => { .close => {
if (request.node < opens.len) opens[@intCast(request.node)].used = false; const of = openAt(request.node);
if (of) |o| {
if (o.backend) |backend| forwardClose(backend, o.node);
o.used = false;
}
return writeReply(out, .{ .status = 0 }, &.{}); return writeReply(out, .{ .status = 0 }, &.{});
}, },
.mkdir, .unlink => {
const name = payload[0..@min(payload.len, request.len)];
if (longestMount(name)) |m| return forwardPath(out, m.backend, request.operation, m.relative);
// Only a mounted backend has real directories; the flat ramfs cannot
// create or remove them (and a bare-name path is not a mount target).
return fail(out);
},
.rename => {
const both = payload[0..@min(payload.len, request.len)];
const sep = std.mem.indexOfScalar(u8, both, 0) orelse return fail(out);
const old_path = both[0..sep];
const new_path = both[sep + 1 ..];
const mo = longestMount(old_path) orelse return fail(out);
const mn = longestMount(new_path) orelse return fail(out);
// Both paths must live under the same mount — cross-filesystem rename is
// not supported.
if (mo.backend != mn.backend) return fail(out);
return forwardRename(out, mo.backend, mo.relative, mn.relative);
},
} }
} }
/// Startup, under the harness: subscribe to the published exit events — when a /// Startup, under the harness: subscribe to the published exit events — when a
/// client dies holding open handles, the exit notification is how the VFS learns /// client dies holding open handles, the exit notification is how the VFS learns
/// to release them (docs/process-lifecycle.md). /// to release them (docs/process-lifecycle.md).
fn initialise(endpoint: runtime.ipc.Handle) bool { fn initialise(endpoint: ipc.Handle) bool {
if (!runtime.process.subscribeExits(endpoint)) { if (!runtime.process.subscribeExits(endpoint)) {
_ = runtime.system.write("/system/services/vfs: exit subscription failed\n"); _ = runtime.system.write("/system/services/vfs: exit subscription failed\n");
} }
@@ -152,8 +378,8 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
/// A non-signal notification: the only kind the VFS subscribes to is exit events. /// A non-signal notification: the only kind the VFS subscribes to is exit events.
fn onNotification(badge: u64) void { fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_exit_bit != 0) { if (badge & ipc.notify_exit_bit != 0) {
releaseClientHandles(@intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit))); releaseClientHandles(@intCast(badge & ~(ipc.notify_badge_bit | ipc.notify_exit_bit)));
} }
} }
+84 -27
View File
@@ -67,7 +67,14 @@ ARCHES = {
"-machine", "q35", "-m", "128M", "-machine", "q35", "-m", "128M",
"-drive", f"if=pflash,format=raw,readonly=on,file={a['ovmf_code']}", "-drive", f"if=pflash,format=raw,readonly=on,file={a['ovmf_code']}",
"-drive", f"if=pflash,format=raw,file={vars_fd}", "-drive", f"if=pflash,format=raw,file={vars_fd}",
"-drive", f"format=raw,file=fat:rw:{boot_volume}", # Boot off a FAT USB device: the boot volume is a mass-storage device on
# the xHCI bus (usb-kbd/usb-mouse ride the same controller). `boot_volume`
# is the FAT image the build produces. bootindex=0 steers OVMF to it.
"-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0",
"-drive", f"if=none,id=bootusb,format=raw,file={boot_volume}",
"-device", "usb-storage,bus=xhci.0,drive=bootusb,removable=on,bootindex=0",
"-net", "none", "-net", "none",
"-vga", "none", "-device", "VGA,edid=on,xres=1280,yres=720", "-vga", "none", "-device", "VGA,edid=on,xres=1280,yres=720",
"-display", "none", "-display", "none",
@@ -101,6 +108,11 @@ CASES = [
{"name": "clock", {"name": "clock",
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# Wall-clock: the CMOS RTC read at boot gives a plausible current epoch (the
# foundation for filesystem mtime).
{"name": "wall-clock",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "vmm", {"name": "vmm",
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -275,9 +287,8 @@ CASES = [
{"name": "usb-report", {"name": "usb-report",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci", # The xHCI bus + usb-kbd/usb-mouse come from the default boot config now
"-device", "usb-kbd,bus=xhci.0", # (every case boots off a usb-storage device on that bus).
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-manager: child added[\s\S]*" "expect": r"device-manager: child added[\s\S]*"
r"device-manager: child added[\s\S]*" r"device-manager: child added[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*" r"device-manager: test mode: killing the reporter[\s\S]*"
@@ -292,24 +303,62 @@ CASES = [
{"name": "usb-hid", {"name": "usb-hid",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci", # usb-kbd/usb-mouse ride the default boot xHCI bus (see qemu_args).
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"(?=[\s\S]*usb-hid/keyboard: ok)(?=[\s\S]*usb-hid/mouse: ok)", "expect": r"(?=[\s\S]*usb-hid/keyboard: ok)(?=[\s\S]*usb-hid/mouse: ok)",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# USB mass storage end to end: attach a usb-storage device (a FAT volume via # USB mass storage end to end: the boot usb-storage device (the FAT32 image,
# QEMU's VVFAT, so it has a real boot sector), boot the full tree, and let the # which has a real 0x55AA boot sector) is enough — the manager spawns
# manager spawn usb-storage, which opens the device, runs the Bulk-Only / # usb-storage, which opens the device, runs the Bulk-Only / SCSI bring-up,
# SCSI bring-up, reads its capacity, and reads block 0 (the 0x55AA boot sig) — # reads its capacity, and reads block 0 (the 0x55AA boot sig). Proof of the
# proof of the bulk transfer path + BOT + SCSI end to end. # bulk transfer path + BOT + SCSI end to end.
{"name": "usb-storage", {"name": "usb-storage",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-drive", "if=none,id=stick,format=raw,file=fat:rw:" + os.path.join(REPO, "zig-out"),
"-device", "usb-storage,drive=stick,bus=xhci.0"],
"expect": r"usb-storage: ready[\s\S]*usb-storage: block 0 signature 0x55aa", "expect": r"usb-storage: ready[\s\S]*usb-storage: block 0 signature 0x55aa",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# FAT mount end to end: the fat server mounts the boot usb-storage device (the
# FAT32 image) into the VFS at /mnt/usb. A fat-test client then lists and reads
# through the mount — proof of the whole stack: block device -> FAT parse ->
# VFS routing -> file read.
{"name": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat: mounted /mnt/usb[\s\S]*fat-test: ok",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the
# fat-test client, after listing, makes a directory, writes+reads a file inside
# it, then removes the file, exercising the whole VFS -> fat mutation path.
{"name": "fat-mutations",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: mutations ok",
"fail": r"fat-test: mutations FAILED|fat-test: mkdir .* failed|DANOS-TEST-RESULT: FAIL"},
# Phase 2c: rename through the mount — fat-test renames the file it created
# before removing it, and confirms the old name is gone.
{"name": "fat-rename",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: rename ok",
"fail": r"fat-test: mutations FAILED|DANOS-TEST-RESULT: FAIL"},
# Phase 2d: filesystem timestamps — a freshly-created file's mtime is a real
# current wall-clock time (stamped from the RTC), read back through stat.
{"name": "fat-mtime",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: mtime ok",
"fail": r"fat-test: mutations FAILED|DANOS-TEST-RESULT: FAIL"},
# Boot-from-USB smoke: the whole system now boots off the FAT32 image on a
# usb-storage device (OVMF -> \EFI\BOOT\BOOTX64.efi -> kernel), so the kernel
# reaching its PASS marker at all proves the USB boot path end to end. Reuses
# the smoke kernel build; the value is the explicit, named regression guard.
{"name": "usb-boot",
"build_case": "smoke",
"qmp_after": {"delay": 2, "command": "query-status"},
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.1: the ring-3 AML parse (the acpi service maps the blobs and parses # M20.1: the ring-3 AML parse (the acpi service maps the blobs and parses
# them) finds exactly the Device count the kernel's own parse produced. # them) finds exactly the Device count the kernel's own parse produced.
{"name": "acpi-parse", {"name": "acpi-parse",
@@ -349,15 +398,26 @@ CASES = [
r"init: shutting down[\s\S]*" r"init: shutting down[\s\S]*"
r"power: entering S5", r"power: entering S5",
"fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"}, "fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"},
# M8: the boot log is persisted to the USB FAT volume. Reuses the orderly-
# shutdown build (full tree + power button): init spawns log-flush at boot,
# which copies the kernel log to /mnt/usb/DANOS.LOG once /mnt/usb is mounted
# (first marker); then the power button drives init's own pre-teardown flush
# (second marker), proving both triggers write the file while storage is up.
{"name": "log-flush",
"build_case": "orderly-shutdown",
"smp": 4,
"timeout": 150,
"qmp_after": {"delay": 8, "command": "system_powerdown"},
"expect": r"log-flush: wrote \d+ bytes to /mnt/usb/DANOS\.LOG[\s\S]*"
r"init: flushed log to /mnt/usb/DANOS\.LOG[\s\S]*"
r"power: entering S5",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers + # M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers +
# reports its _HID devices — the two PS/2 nodes must appear with resources # reports its _HID devices — the two PS/2 nodes must appear with resources
# (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/discovery.md). # (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/discovery.md).
{"name": "acpi-report", {"name": "acpi-report",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"acpi: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*" "expect": r"acpi: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*"
r"acpi: reported PNP0F13 \(device \d+, 1 resources\)", r"acpi: reported PNP0F13 \(device \d+, 1 resources\)",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -374,9 +434,6 @@ CASES = [
{"name": "device-list", {"name": "device-list",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-list: \d+ devices[\s\S]*" "expect": r"device-list: \d+ devices[\s\S]*"
r"device-list: subscribed[\s\S]*" r"device-list: subscribed[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*" r"device-manager: test mode: killing the reporter[\s\S]*"
@@ -389,9 +446,6 @@ CASES = [
{"name": "driver-restart", {"name": "driver-restart",
"smp": 4, "smp": 4,
"timeout": 150, "timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"usb-xhci-bus: hello acknowledged[\s\S]*" "expect": r"usb-xhci-bus: hello acknowledged[\s\S]*"
r"device-manager: restarting crash-test[\s\S]*" r"device-manager: restarting crash-test[\s\S]*"
r"device-manager: crash-test is failing repeatedly", r"device-manager: crash-test is failing repeatedly",
@@ -497,12 +551,15 @@ def qmp_send(path, command):
def run_case(arch, case): def run_case(arch, case):
err = build(arch, case["name"]) # A case's kernel build defaults to its name; `build_case` decouples the two
# so a case can reuse another's kernel (e.g. usb-boot reuses smoke's).
err = build(arch, case.get("build_case", case["name"]))
if err: if err:
return False, "build failed:\n" + err return False, "build failed:\n" + err
# zig-out is the FHS boot volume; hand it to the guest as-is (see qemu_args). # The bootable FAT32 USB image the build produced (tools/make-fat-image.py),
boot_volume = os.path.join(REPO, "zig-out") # presented to the guest as a usb-storage device (see qemu_args).
boot_volume = os.path.join(REPO, "zig-out", "danos-usb.img")
vars_fd = os.path.join(WORK, "vars.fd") vars_fd = os.path.join(WORK, "vars.fd")
shutil.copy(arch["ovmf_vars"], vars_fd) shutil.copy(arch["ovmf_vars"], vars_fd)
serial = os.path.join(WORK, "serial.log") serial = os.path.join(WORK, "serial.log")
+362
View File
@@ -0,0 +1,362 @@
#!/usr/bin/env python3
"""Format a real FAT32 image from a set of host files — the danos boot volume.
Mirrors tools/make-initial-ramdisk.py in spirit: pure Python 3 standard library,
no external tools (no mkfs.fat / mtools). It writes a valid FAT32 filesystem — a
boot sector + BPB, an FSInfo sector, a backup boot sector, two FATs, and a
directory tree of clusters — so UEFI/OVMF boots \\EFI\\BOOT\\BOOTX64.efi off it
and the danos FAT driver mounts the same image.
make-fat-image.py <out.img> <size-MiB> [<dest-path> <host-file>]...
make-fat-image.py --verify <out.img>
Each <dest-path> is a forward-slash path inside the image (e.g.
"EFI/BOOT/BOOTX64.efi"); intermediate directories are created. Names that do not
fit 8.3 get a mangled short name plus long-file-name (LFN) entries.
"""
import struct
import sys
SECTOR = 512
SECTORS_PER_CLUSTER = 1 # 512-byte clusters keep the cluster count high for FAT32
RESERVED_SECTORS = 32
NUM_FATS = 2
CLUSTER_BYTES = SECTOR * SECTORS_PER_CLUSTER
END_OF_CHAIN = 0x0FFFFFFF
BAD_CLUSTER = 0x0FFFFFF7
ATTR_ARCHIVE = 0x20
ATTR_DIRECTORY = 0x10
ATTR_LONG_NAME = 0x0F
VALID_83 = set("ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789$%'-_@~!(){}^#& ")
def fat32_geometry(total_sectors):
"""Solve for the FAT size (sectors per FAT) and cluster count that fit."""
fat_size = 1
while True:
data_sectors = total_sectors - RESERVED_SECTORS - NUM_FATS * fat_size
cluster_count = data_sectors // SECTORS_PER_CLUSTER
needed = ((cluster_count + 2) * 4 + SECTOR - 1) // SECTOR
if needed <= fat_size:
return fat_size, cluster_count
fat_size = needed
class Fat32Image:
def __init__(self, total_sectors):
self.total_sectors = total_sectors
self.fat_size, self.cluster_count = fat32_geometry(total_sectors)
if self.cluster_count < 65525:
sys.exit(f"error: image too small for FAT32 ({self.cluster_count} clusters "
f"< 65525); use a larger size")
self.first_data_sector = RESERVED_SECTORS + NUM_FATS * self.fat_size
# The FAT, in memory: entry 0 media, entry 1 EOC, entry 2 the root dir.
self.fat = [0] * (self.cluster_count + 2)
self.fat[0] = 0x0FFFFFF8
self.fat[1] = END_OF_CHAIN
self.fat[2] = END_OF_CHAIN
self.next_free = 3
self.cluster_data = {} # cluster number -> bytes (one cluster's worth)
def alloc(self):
cluster = self.next_free
if cluster >= self.cluster_count + 2:
sys.exit("error: image out of clusters")
self.next_free += 1
self.fat[cluster] = END_OF_CHAIN
return cluster
def store_chain(self, content):
"""Allocate a cluster chain holding `content` and return its first cluster."""
length = max(1, (len(content) + CLUSTER_BYTES - 1) // CLUSTER_BYTES)
clusters = [self.alloc() for _ in range(length)]
for i in range(length - 1):
self.fat[clusters[i]] = clusters[i + 1]
for i, cluster in enumerate(clusters):
chunk = content[i * CLUSTER_BYTES:(i + 1) * CLUSTER_BYTES]
self.cluster_data[cluster] = chunk + b"\x00" * (CLUSTER_BYTES - len(chunk))
return clusters[0]
def store_directory(self, first_cluster, entries):
"""Write directory `entries` (bytes) into `first_cluster`, extending the chain."""
length = max(1, (len(entries) + CLUSTER_BYTES - 1) // CLUSTER_BYTES)
clusters = [first_cluster]
for _ in range(length - 1):
clusters.append(self.alloc())
for i in range(len(clusters) - 1):
self.fat[clusters[i]] = clusters[i + 1]
for i, cluster in enumerate(clusters):
chunk = entries[i * CLUSTER_BYTES:(i + 1) * CLUSTER_BYTES]
self.cluster_data[cluster] = chunk + b"\x00" * (CLUSTER_BYTES - len(chunk))
def cluster_sector(self, cluster):
return self.first_data_sector + (cluster - 2) * SECTORS_PER_CLUSTER
def serialize(self):
image = bytearray(self.total_sectors * SECTOR)
image[0:SECTOR] = self.boot_sector()
image[SECTOR:2 * SECTOR] = self.fsinfo_sector()
image[6 * SECTOR:7 * SECTOR] = self.boot_sector() # backup boot sector
# Both FATs.
fat_bytes = b"".join(struct.pack("<I", entry & 0x0FFFFFFF) for entry in self.fat)
fat_bytes += b"\x00" * (self.fat_size * SECTOR - len(fat_bytes))
for copy in range(NUM_FATS):
base = (RESERVED_SECTORS + copy * self.fat_size) * SECTOR
image[base:base + len(fat_bytes)] = fat_bytes
# The data region (clusters).
for cluster, data in self.cluster_data.items():
base = self.cluster_sector(cluster) * SECTOR
image[base:base + len(data)] = data
return bytes(image)
def boot_sector(self):
sector = bytearray(SECTOR)
# BPB.
struct.pack_into(
"<3s8sHBHBHHBHHHII", sector, 0,
b"\xEB\x58\x90", # jump
b"MSWIN4.1", # OEM name (widest firmware compatibility)
SECTOR, # bytes per sector
SECTORS_PER_CLUSTER, # sectors per cluster
RESERVED_SECTORS, # reserved sector count
NUM_FATS, # number of FATs
0, # root entry count (0 for FAT32)
0, # total sectors 16 (0 -> use 32)
0xF8, # media descriptor
0, # FAT size 16 (0 for FAT32)
32, # sectors per track
2, # heads
0, # hidden sectors
self.total_sectors, # total sectors 32
)
# FAT32 extended BPB (offset 36).
struct.pack_into(
"<IHHIHH12sBBBI11s8s", sector, 36,
self.fat_size, # FAT size 32
0, # extended flags
0, # filesystem version
2, # root cluster
1, # FSInfo sector
6, # backup boot sector
b"\x00" * 12, # reserved
0x80, # drive number
0, # reserved
0x29, # extended boot signature
0x12345678, # volume id
b"DANOS ", # volume label
b"FAT32 ", # filesystem type
)
sector[510] = 0x55
sector[511] = 0xAA
return bytes(sector)
def fsinfo_sector(self):
sector = bytearray(SECTOR)
struct.pack_into("<I", sector, 0, 0x41615252) # lead signature
struct.pack_into("<I", sector, 484, 0x61417272) # struct signature
free = self.cluster_count - (self.next_free - 2)
struct.pack_into("<I", sector, 488, free) # free count
struct.pack_into("<I", sector, 492, self.next_free) # next free hint
struct.pack_into("<I", sector, 508, 0xAA550000) # trail signature
return bytes(sector)
def lfn_checksum(short_name):
checksum = 0
for byte in short_name:
checksum = (((checksum & 1) << 7) + (checksum >> 1) + byte) & 0xFF
return checksum
def short_name_for(name, used):
"""Return (raw 11-byte 8.3 name, needs_lfn)."""
if "." in name and not name.startswith("."):
base, ext = name.rsplit(".", 1)
else:
base, ext = name, ""
upper_base, upper_ext = base.upper(), ext.upper()
# A name fits 8.3 if it is short enough and uses valid characters; a lowercase
# name is simply stored uppercased (FAT is case-insensitive, so the bootloader
# and the danos driver still find it). Only genuinely non-8.3 names (too long,
# e.g. initial-ramdisk.img) get a mangled short name plus LFN entries.
fits = (1 <= len(base) <= 8 and len(ext) <= 3
and all(c in VALID_83 for c in upper_base + upper_ext))
if fits:
return (upper_base.ljust(8) + upper_ext.ljust(3)).encode("ascii"), False
# Mangle to STEM~N.EXT.
stem = "".join(c for c in upper_base if c in VALID_83 and c != " ")[:6] or "FILE"
index = 1
while True:
candidate = f"{stem}~{index}".ljust(8)[:8] + upper_ext.ljust(3)[:3]
raw = candidate.encode("ascii")
if raw not in used:
used.add(raw)
return raw, True
index += 1
def lfn_entries(name, short_raw):
checksum = lfn_checksum(short_raw)
units = list(name.encode("utf-16-le"))
pairs = [bytes(units[i:i + 2]) for i in range(0, len(units), 2)]
pairs.append(b"\x00\x00") # null terminator
while len(pairs) % 13 != 0:
pairs.append(b"\xff\xff")
count = len(pairs) // 13
out = bytearray()
for sequence in range(count, 0, -1): # stored last-logical-first
piece = pairs[(sequence - 1) * 13:sequence * 13]
entry = bytearray(32)
entry[0] = sequence | (0x40 if sequence == count else 0)
for i in range(5):
entry[1 + i * 2:1 + i * 2 + 2] = piece[i]
entry[11] = ATTR_LONG_NAME
entry[12] = 0
entry[13] = checksum
for i in range(6):
entry[14 + i * 2:14 + i * 2 + 2] = piece[5 + i]
entry[26:28] = b"\x00\x00"
for i in range(2):
entry[28 + i * 2:28 + i * 2 + 2] = piece[11 + i]
out += entry
return bytes(out)
def short_entry(raw11, attributes, cluster, size):
return struct.pack(
"<11sBBBHHHHHHHI",
raw11, attributes, 0, 0, 0, 0, 0,
(cluster >> 16) & 0xFFFF, 0, 0, cluster & 0xFFFF, size,
)
def write_directory(image, cluster, children, parent_cluster, is_root):
"""Recursively lay out a directory: allocate child clusters, build entries."""
entries = bytearray()
if not is_root:
entries += short_entry(b". ", ATTR_DIRECTORY, cluster, 0)
parent = 0 if parent_cluster == 2 else parent_cluster
entries += short_entry(b".. ", ATTR_DIRECTORY, parent, 0)
used_short_names = set()
for name, child in children.items():
raw, needs_lfn = short_name_for(name, used_short_names)
used_short_names.add(raw)
if child["type"] == "dir":
child_cluster = image.alloc()
if needs_lfn:
entries += lfn_entries(name, raw)
entries += short_entry(raw, ATTR_DIRECTORY, child_cluster, 0)
write_directory(image, child_cluster, child["children"], cluster, False)
else:
data = child["data"]
first = image.store_chain(data) if data else 0
if needs_lfn:
entries += lfn_entries(name, raw)
entries += short_entry(raw, ATTR_ARCHIVE, first, len(data))
image.store_directory(cluster, bytes(entries))
def build_tree(pairs):
root = {}
for dest, host in pairs:
with open(host, "rb") as handle:
data = handle.read()
parts = [p for p in dest.replace("\\", "/").split("/") if p]
node = root
for part in parts[:-1]:
node = node.setdefault(part, {"type": "dir", "children": {}})["children"]
node[parts[-1]] = {"type": "file", "data": data}
return root
def build(out_path, size_mib, pairs):
total_sectors = size_mib * 1024 * 1024 // SECTOR
image = Fat32Image(total_sectors)
tree = build_tree(pairs)
write_directory(image, 2, tree, 0, True)
with open(out_path, "wb") as handle:
handle.write(image.serialize())
print(f"make-fat-image: wrote {out_path} "
f"({size_mib} MiB FAT32, {image.cluster_count} clusters)")
def verify(path):
with open(path, "rb") as handle:
data = handle.read()
if len(data) < SECTOR or data[510] != 0x55 or data[511] != 0xAA:
sys.exit("verify: missing 0x55AA boot signature")
bytes_per_sector, sectors_per_cluster = struct.unpack_from("<HB", data, 11)
reserved, num_fats = struct.unpack_from("<H", data, 14)[0], data[16]
fat_size_32, root_cluster = struct.unpack_from("<I", data, 36)[0], struct.unpack_from("<I", data, 44)[0]
total_sectors = struct.unpack_from("<I", data, 32)[0]
if bytes_per_sector != SECTOR or sectors_per_cluster == 0 or num_fats == 0 or fat_size_32 == 0:
sys.exit("verify: implausible BPB")
first_data = reserved + num_fats * fat_size_32
cluster_count = (total_sectors - first_data) // sectors_per_cluster
if cluster_count < 65525:
sys.exit(f"verify: not FAT32 ({cluster_count} clusters)")
# Resolve EFI/BOOT/BOOTX64.efi through the directory tree to prove it is present.
if not _resolve(data, ["EFI", "BOOT", "BOOTX64.EFI"], root_cluster,
reserved, num_fats, fat_size_32, first_data, sectors_per_cluster):
sys.exit("verify: EFI/BOOT/BOOTX64.efi not found")
print(f"verify: {path} is FAT32 ({cluster_count} clusters); EFI/BOOT/BOOTX64.efi present")
def _read_fat(data, cluster, reserved):
offset = reserved * SECTOR + cluster * 4
return struct.unpack_from("<I", data, offset)[0] & 0x0FFFFFFF
def _resolve(data, parts, cluster, reserved, num_fats, fat_size, first_data, spc):
for part in parts:
cluster = _find(data, cluster, part, reserved, first_data, spc)
if cluster is None:
return False
return True
def _find(data, dir_cluster, name, reserved, first_data, spc):
target = name.upper()
cluster = dir_cluster
guard = 0
while cluster >= 2 and cluster < BAD_CLUSTER and guard < 100000:
sector = first_data + (cluster - 2) * spc
for s in range(spc):
base = (sector + s) * SECTOR
for i in range(SECTOR // 32):
entry = data[base + i * 32:base + i * 32 + 32]
if entry[0] == 0x00:
return None
if entry[0] == 0xE5 or (entry[11] & ATTR_LONG_NAME) == ATTR_LONG_NAME:
continue
raw = entry[0:11]
short = (raw[0:8].rstrip().decode("latin1") +
("." + raw[8:11].rstrip().decode("latin1") if raw[8:11].strip() else "")).upper()
if short == target:
return ((entry[20] | (entry[21] << 8)) << 16) | (entry[26] | (entry[27] << 8))
cluster = _read_fat(data, cluster, reserved)
guard += 1
return None
def main(argv):
if len(argv) == 3 and argv[1] == "--verify":
verify(argv[2])
return 0
if len(argv) < 3 or (len(argv) - 3) % 2 != 0:
sys.exit("usage: make-fat-image.py <out.img> <size-MiB> [<dest> <host>]...\n"
" make-fat-image.py --verify <out.img>")
out_path = argv[1]
size_mib = int(argv[2])
rest = argv[3:]
pairs = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)]
build(out_path, size_mib, pairs)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))