Author SHA1 Message Date
Daniel Samson 67702fa250 Phase 2d (ii): filesystem modification time (mtime)
Completes Phase 2: the FAT filesystem now stamps and reports a real modification
time, built on the Phase 2d(i) kernel wall-clock. This is the last stat field the
compiler's build cache needs to reason about (source vs cached output).

- on-disk.zig: fatToEpoch / epochToFatDateTime convert between the two 16-bit DOS
  date/time fields and Unix epoch seconds (UTC — FAT has no timezone). Host-tested
  round-trip + an absolute check (1577836800 == 2020-01-01).
- engine: a settable current_time_epoch that create/write stamp into the entry's
  write (and creation) date/time; Node/Listing gained an mtime decoded from those
  fields on read. Host test: a create stamps the mtime, read back through resolve
  and listEntry.
- vfs protocol FileStatus + runtime.fs.Attributes gained an mtime field; the fat
  server sets current_time_epoch from runtime.system.wallClock() per request and
  returns mtime from stat. The flat ramfs reports 0 (it has no timestamps).
- fat-test reads the created file's mtime through stat and checks it is a real
  current time, behind a new `fat-mtime` QEMU case.

Verified against the host: the guest stamped mtime 1783971676 while the host clock
was 1783971680 (boot+test lag) — the file's mtime is real current time. zig build,
zig build test (the epoch<->DOS conversions + the engine mtime test),
zig build check-fat-image, and a sequential QEMU sweep — fat-mount, fat-mutations,
fat-rename, fat-mtime, vfs, vfs-client-death, log-flush, orderly-shutdown,
initial-ramdisk, smoke, wall-clock, usb-storage — all green. mode/inode remain.
2026-07-13 20:43:40 +01:00
Daniel Samson 8a38540312 Phase 2d (i): a kernel wall-clock from the CMOS RTC
Adds real (calendar) time, the foundation for filesystem timestamps. Monotonic
time (`clock`) says how long since boot; this says what time it actually is.

- cpu.zig (x86_64): readRtcUnixSeconds() reads the CMOS real-time clock (ports
  0x70/0x71) — waits out an update-in-progress, reads twice until stable, handles
  BCD-vs-binary and 12-vs-24-hour per status register B — and converts to Unix
  epoch seconds (UTC).
- kernel/wall-clock.zig: reads the RTC once at boot and anchors it to the monotonic
  clock, so a query is a cheap arithmetic offset — no per-call CMOS poll, no lock,
  no SMP hazard on the shared ports. kmain calls init() once the monotonic clock is
  final and logs the epoch.
- wall_clock() syscall (33) -> Unix epoch seconds, wrapped by runtime.system
  .wallClock(). Wall-clock *seconds* are mechanism the kernel owns like the
  monotonic clock; calendars/timezones are user-space policy (the stale comment on
  systemClock that called wall-clock a "user-space service" is updated in spirit by
  the new handler's doc).
- A `wall-clock` kernel test asserts the boot RTC read is a plausible current epoch.

Verified against the host: the guest read epoch 1783971244 while `date -u +%s` gave
1783971245 (one second of boot lag) — the CMOS read + epoch conversion are correct to
the second. zig build, zig build test, and smoke/clock/init are green.
2026-07-13 20:35:03 +01:00
Daniel Samson 54635eecf5 Phase 2c: rename, wired through the VFS to runtime.fs
Completes the Phase 2 FAT mutation set (truncate, mkdir, unlink, rename).

- engine: rename(dir, old_name, new_name) rewrites an existing entry's 8.3 name in
  place within the same directory. Refuses a missing source, a non-8.3 target, or a
  name that already exists; drops any long-name entries on the old file (it takes
  its new 8.3 name), LFN-aware like removeFile. Cross-directory and long-name-
  preserving rename are noted limitations. Host-tested (rename keeps contents;
  collision, non-8.3, and missing-source are refused).
- vfs protocol: a `rename` operation whose payload is old-path, a 0x00 separator,
  then new-path.
- VFS router: a forwardRename helper + a `.rename` case that requires both paths
  under the same mount (cross-filesystem rename is refused) and forwards the
  mount-relative old+new.
- fat server: a `.rename` handler that requires the same parent directory and calls
  engine.rename.
- runtime.fs: rename(old_path, new_path).
- fat-test now renames the file it created (before removing it) and asserts the old
  name is gone, behind a new `fat-rename` QEMU case.

Verified: zig build, zig build test (the engine rename unit test), zig build
check-fat-image, and a sequential QEMU sweep — fat-mount, fat-mutations, fat-rename,
vfs, vfs-client-death, log-flush, orderly-shutdown, initial-ramdisk, smoke — green.
2026-07-13 20:22:32 +01:00
Daniel Samson 184d90c2c6 Phase 2b: wire mkdir + unlink through the VFS to runtime.fs
The engine gained mkdir/unlink in Phase 2a; this exposes them as first-class
filesystem operations so programs can use them.

- vfs protocol: two new path-based operations, mkdir and unlink (appended, so
  existing opcodes/offsets are unchanged).
- VFS router: a forwardPath helper relays a path-based op under a mount to its
  backend; the mkdir/unlink cases forward to the mounted filesystem (the flat
  ramfs refuses them — it has no directories).
- fat server: mkdir -> engine.createDirectory, unlink -> engine.removeFile, each
  resolving the parent via a shared splitParent helper (also used by open-create).
- runtime.fs: makeDirectory(path) and remove(path).
- fat-test now exercises the whole path — mkdir /mnt/usb/TESTDIR, create + write +
  read a file inside it, then remove it — behind a new `fat-mutations` QEMU case.

Verified: zig build, zig build test, zig build check-fat-image, and a sequential
QEMU sweep — fat-mount, fat-mutations (mkdir/write/read/unlink through the mount),
vfs, vfs-client-death, log-flush, orderly-shutdown, initial-ramdisk, smoke — green.
2026-07-13 20:12:07 +01:00
Daniel Samson a32eed877d Phase 2a: FAT engine mutations + O_TRUNC (fix the overwrite corruption)
The FAT engine could create, read, write, and grow files, but never free clusters
or make directories — so overwriting a shorter file left a stale tail (a real bug:
corrupt boot-log re-flushes, and later corrupt compiler cache/.o files). This adds
the mutation half of the engine, with the corruption fix wired all the way through.

Engine (system/services/fat/engine.zig), all host-tested:
- freeChain: return a cluster chain to the pool (bounded against a corrupt cycle) —
  the shared primitive under truncate and remove.
- truncate: free the chain and zero the entry's size/first-cluster (O_TRUNC).
- createDirectory (mkdir): allocate + initialise a cluster with "." and ".." and add
  the directory entry to the parent.
- removeFile (unlink): free the chain and mark the 8.3 entry plus any preceding
  long-name entries deleted, so a reused slot can't inherit an orphaned long name.
  Refuses directories.
- createFile now shares a common addEntry helper with createDirectory.

O_TRUNC wired end to end: a truncate open-flag (vfs protocol) that the router already
forwards; runtime.fs.OpenOptions.truncate; and fat's handleOpen calls engine.truncate
on an existing file. The boot-log flush (log-flush + init) now opens with truncate, so
a shorter log on a later boot of the same stick leaves no stale tail — closing the
caveat from the boot-log work.

mkdir/unlink are engine-complete and host-tested but not yet exposed as VFS
operations / runtime.fs methods (they are new path-based ops needing router cases);
rename and the richer stat (mtime/mode, blocked on wall-clock) remain. See
docs/zig-self-hosting.md (Phase 2).

Verified: zig build, zig build test (8 engine host tests, incl. truncate, the
overwrite-no-stale-tail regression, remove, and mkdir), zig build check-fat-image, and
a sequential QEMU sweep — fat-mount, log-flush (DANOS.LOG read back at 11804 bytes,
clean, with the truncate-based final flush), vfs, vfs-client-death, usb-storage,
orderly-shutdown, initial-ramdisk, smoke — all green.
2026-07-13 20:02:01 +01:00
Daniel Samson 347a041d85 Phase 1a: add the native runtime.fs, retire the posix shim
The first step of the Zig self-hosting roadmap (docs/zig-self-hosting.md): give
danos programs a danos-native file API and remove the premature POSIX compatibility
shim. This also resolves the earlier misplacement of a full-write helper into the
compat layer — that behaviour now lives natively in runtime.fs.File.writeAll.

- library/runtime/fs.zig: the danos-native file client over the VFS (open/read/
  write/writeAll/seekTo/attributes/close, directory listing, mount). Handles are
  *values* — a File/Directory owns its VFS node id and byte offset — so there is no
  per-process fd table or descriptor limit, unlike the POSIX fd model the shim
  emulated. This is where the operations that later become std.os.danos are staged.
- Retire library/posix/ (unistd, stdio): only five call sites used it, all file
  operations, all migrated to runtime.fs — fat (mount), the vfs-test and fat-test
  clients, and init/log-flush (the boot-log flush). stdio was already dead.
- build.zig: drop the posix module, its addUserBinary parameter, the per-binary
  import, and the ~26 call-site arguments.
- Docs: the VFS protocol's client is now runtime.fs; the docs index and
  coding-standards note posix is retired and the foreign-ABI naming exception now
  applies to the future std.os.danos seam; the process-lifecycle note points the
  future musl layer at that same seam rather than the deleted directory.

Deferred by design (see the roadmap): the C-ABI runtime.os errno seam is built at
fork time (its shape must match std/os/danos.zig); truncate/mkdir/rename are
Phase 2; stdio-byte fds and cwd are later slices.

Verified: zig build, zig build test, zig build check-fat-image, and a sequential
QEMU sweep — vfs, vfs-client-death (the park/hold-handle path), fat-mount, log-flush,
orderly-shutdown, initial-ramdisk (log-flush silent in the bare sweep), smoke, init,
usb-storage, device-manager — all green.
2026-07-13 19:43:56 +01:00
Daniel Samson 53e42837e0 Docs: add the Zig self-hosting roadmap
A forward-looking design note on making danos a real Zig target
(-target x86_64-danos) and eventually running the compiler on it, focused on
the standard-library surface (not the editor/terminal).

The core realisation: Zig 0.16 (post-writergate) collapses an OS port to ONE
seam — std.fs is gone, everything routes through the std.Io vtable, and
std.posix is generic over a single per-OS `system` module (std.os.<tag>). So the
port is "write std.os.danos once" and the whole fs/process/Io tower lights up,
rather than reimplementing the namespaces.

Records the decisions this shapes now: build runtime.os (the seam, promoted into
a forked std/os/danos.zig later) plus a thin runtime.fs; retire the premature
library/posix shim (only 5 unistd call sites); do NOT hand-mirror the high-level
std namespaces; do NOT emulate the Linux ABI; defer musl. Covers the host/target/
self-host roles and the four-part compiler fork, a coverage table of what danos
has vs the gaps (mkdir/unlink/rename/truncate, richer stat, wall-clock, env, cwd,
entropy, stdio bytes), a phased plan (target -> read-side+retire-posix ->
fs-mutation+stat -> single-threaded self-linked compiler), and the risks
(fork rebase treadmill, -fsingle-threaded and -fno-llvm/-fno-lld being
load-bearing, the "w"-does-not-truncate corruption bug). Linked from the docs index.
2026-07-13 19:26:51 +01:00
Daniel Samson f52c591f5e Persist the kernel boot log to the USB FAT volume
On a headless or real board nothing captures serial, so the boot log — the whole
diagnostic stream — is lost at power-off. This retains it in the kernel and copies
it to the boot USB volume as /mnt/usb/DANOS.LOG, the on-disk equivalent of QEMU's
`-serial file:`. Pull the stick, read DANOS.LOG on another machine.

How it fits together:
- Kernel RAM sink (log.zig): a fixed 256 KiB in-image buffer registered as a log
  sink in kmain, right after serial. Because userspace debug_write funnels through
  log.write, it captures the entire stream — kernel lines and every service's
  output — from the first line. Fills linearly and stops when full (earliest boot
  output, the most valuable, is kept); no allocation, so it is panic-safe.
- klog_read syscall (32): copies that buffer out to a user buffer, the mirror of
  debug_write — same overflow-safe user-half bounds check, kernel -> user copy,
  under the kernel lock so the snapshot can't grow mid-copy. Wrapped by
  runtime.system.klogRead.
- log-flush (new one-shot, in the initial-ramdisk): waits for the fat server to
  mount /mnt/usb, then copies the whole log to /mnt/usb/DANOS.LOG. init spawns it
  once the boot services are up (fire-and-forget; it polls the mount itself). If
  no volume is mounted — no stick, or the no-VFS ramdisk sweep — it exits silently.
- init shutdown flush: init repeats the copy inline at the top of shutDown(),
  BEFORE it tears down the storage services (the fat server is stopped first), so
  a clean poweroff captures the fullest log while /mnt/usb is still writable.
- unistd.writeAll: loops write() past the 224-byte VFS payload cap; both flush
  paths use it.

The filename is 8.3 (DANOS.LOG) at the mount root — the FAT short-name rule, and
there is no mkdir on the FAT path yet. Extend-only writes mean the two same-session
flushes never leave stale bytes (the shutdown log is a superset of the boot log);
a shorter log on a later boot of the same stick can leave a stale tail — a noted,
cosmetic limitation, not worth pulling O_TRUNC into the FAT write path for now.

Verified end to end under QEMU: a new `log-flush` case (reusing the orderly-shutdown
build) asserts both markers then S5, and DANOS.LOG is read back out of the image
afterwards (11804 bytes, containing the kernel init line, the FAT mount line, and
the boot flush marker). Regression stays green: zig build, zig build test, zig
build check-fat-image, and a sequential QEMU sweep — smoke, init, initial-ramdisk
(log-flush silent in the bare sweep), orderly-shutdown, fat-mount, usb-storage,
usb-hid, vfs, input, device-manager, process, signals, dma, fault-pf.
2026-07-13 18:18:34 +01:00
28 changed files with 1771 additions and 460 deletions
+26 -38
View File
@@ -58,7 +58,6 @@ fn addUserBinary(
b: *std.Build, b: *std.Build,
target: std.Build.ResolvedTarget, target: std.Build.ResolvedTarget,
runtime_module: *std.Build.Module, runtime_module: *std.Build.Module,
posix_module: *std.Build.Module,
mmio_module: *std.Build.Module, mmio_module: *std.Build.Module,
xkeyboard_config_module: *std.Build.Module, xkeyboard_config_module: *std.Build.Module,
acpi_ids_module: *std.Build.Module, acpi_ids_module: *std.Build.Module,
@@ -78,9 +77,6 @@ fn addUserBinary(
.stack_protector = false, .stack_protector = false,
.imports = &.{ .imports = &.{
.{ .name = "runtime", .module = runtime_module }, .{ .name = "runtime", .module = runtime_module },
// POSIX/C compatibility layer, available to any program that wants it
// (danos-native code uses `runtime` directly). See library/posix/.
.{ .name = "posix", .module = posix_module },
// Typed volatile MMIO + memory barriers, for drivers. See library/mmio/. // Typed volatile MMIO + memory barriers, for drivers. See library/mmio/.
.{ .name = "mmio", .module = mmio_module }, .{ .name = "mmio", .module = mmio_module },
// Keyboard layouts (keycode + modifiers -> keysym/character), available // Keyboard layouts (keycode + modifiers -> keysym/character), available
@@ -280,18 +276,6 @@ pub fn build(b: *std.Build) void {
}, },
}); });
// The POSIX / C compatibility layer, a separate library layered strictly over the
// runtime (it calls the runtime's IPC/heap, never system calls directly). This is
// the one place POSIX/C spellings are allowed verbatim — see docs/coding-standards.md
// and library/posix/posix.zig.
const posix_module = b.addModule("posix", .{
.root_source_file = b.path("library/posix/posix.zig"),
.imports = &.{
.{ .name = "runtime", .module = runtime_module },
.{ .name = "vfs-protocol", .module = vfs_protocol_module },
},
});
// The initial_ramdisk container format, shared by the kernel (unpacks it) and the // The initial_ramdisk container format, shared by the kernel (unpacks it) and the
// build-time packer tools/make-initial-ramdisk.py (produces it). No dependencies. // build-time packer tools/make-initial-ramdisk.py (produces it). No dependencies.
const initial_ramdisk_module = b.addModule("initial-ramdisk", .{ const initial_ramdisk_module = b.addModule("initial-ramdisk", .{
@@ -363,7 +347,7 @@ pub fn build(b: *std.Build) void {
// Built by the shared user-binary recipe (see addUserBinary): freestanding, // Built by the shared user-binary recipe (see addUserBinary): freestanding,
// linked into the kernel's user region against the `runtime` runtime library, and // linked into the kernel's user region against the `runtime` runtime library, and
// started in ring 3 by the kernel's user-ELF loader. // started in ring 3 by the kernel's user-ELF loader.
const init_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "init", "system/services/init/init.zig"); const init_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "init", "system/services/init/init.zig");
const init_install = b.addInstallArtifact(init_exe, .{ .dest_dir = .{ .override = .{ .custom = "system/services" } } }); const init_install = b.addInstallArtifact(init_exe, .{ .dest_dir = .{ .override = .{ .custom = "system/services" } } });
b.getInstallStep().dependOn(&init_install.step); b.getInstallStep().dependOn(&init_install.step);
@@ -371,12 +355,12 @@ pub fn build(b: *std.Build) void {
// Each is built by the same user-binary recipe, then packed into one image by // Each is built by the same user-binary recipe, then packed into one image by
// the host-side make-initial-ramdisk tool. The bootloader ferries the image to the kernel, // the host-side make-initial-ramdisk tool. The bootloader ferries the image to the kernel,
// which unpacks it and spawns each program (system/initial-ramdisk.zig). // which unpacks it and spawns each program (system/initial-ramdisk.zig).
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig"); const vfs_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig"); const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig"); const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig");
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig"); const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig"); const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig"); const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
// The xHCI bus driver builds chapter-9 requests and decodes descriptors from // The xHCI bus driver builds chapter-9 requests and decodes descriptors from
// usb-abi, and reports each interface's (class,subclass,protocol) identity via // usb-abi, and reports each interface's (class,subclass,protocol) identity via
// usb-ids.packTriple. // usb-ids.packTriple.
@@ -386,26 +370,26 @@ pub fn build(b: *std.Build) void {
// The USB HID class drivers: keyboard and mouse. They own no hardware — each // The USB HID class drivers: keyboard and mouse. They own no hardware — each
// opens its device through runtime.usb (the transfer protocol) and publishes to // opens its device through runtime.usb (the transfer protocol) and publishes to
// the input service. They build chapter-9 class requests from usb-abi. // the input service. They build chapter-9 class requests from usb-abi.
const usb_hid_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-keyboard", "system/drivers/usb-hid/keyboard.zig"); const usb_hid_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-keyboard", "system/drivers/usb-hid/keyboard.zig");
usb_hid_keyboard_exe.root_module.addImport("usb-abi", usb_abi_module); usb_hid_keyboard_exe.root_module.addImport("usb-abi", usb_abi_module);
const usb_hid_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-mouse", "system/drivers/usb-hid/mouse.zig"); const usb_hid_mouse_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-hid-mouse", "system/drivers/usb-hid/mouse.zig");
usb_hid_mouse_exe.root_module.addImport("usb-abi", usb_abi_module); usb_hid_mouse_exe.root_module.addImport("usb-abi", usb_abi_module);
// The USB mass-storage class driver: opens its device via runtime.usb, drives it // The USB mass-storage class driver: opens its device via runtime.usb, drives it
// with Bulk-Only Transport + SCSI, and serves the block protocol under `.block`. // with Bulk-Only Transport + SCSI, and serves the block protocol under `.block`.
const usb_storage_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-storage", "system/drivers/usb-storage/usb-storage.zig"); const usb_storage_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-storage", "system/drivers/usb-storage/usb-storage.zig");
usb_storage_exe.root_module.addImport("block-protocol", block_protocol_module); usb_storage_exe.root_module.addImport("block-protocol", block_protocol_module);
// The FAT filesystem server: mounts the block device and serves it into the VFS // The FAT filesystem server: mounts the block device and serves it into the VFS
// at /mnt/usb. Its engine (engine.zig / on-disk.zig) is imported relatively. // at /mnt/usb. Its engine (engine.zig / on-disk.zig) is imported relatively.
const fat_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat", "system/services/fat/fat.zig"); const fat_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat", "system/services/fat/fat.zig");
const fat_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat-test", "system/services/fat/fat-test.zig"); const fat_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat-test", "system/services/fat/fat-test.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig"); const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
// The PCI bus driver decodes each function's class triple to human names in its // The PCI bus driver decodes each function's class triple to human names in its
// boot log (class/subclass/prog-IF), so pull in the shared pci-class reference. // boot log (class/subclass/prog-IF), so pull in the shared pci-class reference.
pci_bus_exe.root_module.addImport("pci-class", pci_class_module); pci_bus_exe.root_module.addImport("pci-class", pci_class_module);
// A test fixture, not a real driver: hellos to the device manager, then faults — // A test fixture, not a real driver: hellos to the device manager, then faults —
// what the driver-restart scenario drives the crash-loop cap with. // what the driver-restart scenario drives the crash-loop cap with.
const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig"); const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig");
const device_list_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig"); const device_list_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig");
// The discovery service: one swappable process per firmware // The discovery service: one swappable process per firmware
// (docs/discovery.md), bundled under the neutral ramdisk name // (docs/discovery.md), bundled under the neutral ramdisk name
// "discovery" so the device manager never learns which firmware it is on. // "discovery" so the device manager never learns which firmware it is on.
@@ -419,9 +403,9 @@ pub fn build(b: *std.Build) void {
.acpi => "system/services/acpi/acpi.zig", .acpi => "system/services/acpi/acpi.zig",
.fdt => "system/services/fdt/fdt.zig", .fdt => "system/services/fdt/fdt.zig",
}; };
const discovery_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source); const discovery_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source);
if (discovery == .acpi) discovery_exe.root_module.addImport("aml", aml_module); if (discovery == .acpi) discovery_exe.root_module.addImport("aml", aml_module);
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig"); const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// Names the xHCI PCI class triple from the shared taxonomy instead of a bare 0x0C0330. // Names the xHCI PCI class triple from the shared taxonomy instead of a bare 0x0C0330.
device_manager_exe.root_module.addImport("pci-class", pci_class_module); device_manager_exe.root_module.addImport("pci-class", pci_class_module);
// The manager matches reported USB interfaces by their (class,subclass,protocol) // The manager matches reported USB interfaces by their (class,subclass,protocol)
@@ -429,11 +413,12 @@ pub fn build(b: *std.Build) void {
device_manager_exe.root_module.addImport("usb-ids", usb_ids_module); device_manager_exe.root_module.addImport("usb-ids", usb_ids_module);
// The input service and its exercisers: the fan-out server, a hardware-free synthetic // The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md. // source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
const input_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig"); const input_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig");
const input_source_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-source", "system/services/input-source/input-source.zig"); const input_source_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-source", "system/services/input-source/input-source.zig");
const input_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig"); const input_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig");
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig"); const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig");
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig"); const process_test_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig");
const log_flush_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "log-flush", "system/services/log-flush/log-flush.zig");
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool // Pack the user binaries into the initial_ramdisk image with the host-side Python tool
// (the container format is trivial, and Python sidesteps std API churn). Args: // (the container format is trivial, and Python sidesteps std API churn). Args:
@@ -483,6 +468,8 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(args_echo_exe.getEmittedBin()); mk_run.addFileArg(args_echo_exe.getEmittedBin());
mk_run.addArg("process-test"); mk_run.addArg("process-test");
mk_run.addFileArg(process_test_exe.getEmittedBin()); mk_run.addFileArg(process_test_exe.getEmittedBin());
mk_run.addArg("log-flush");
mk_run.addFileArg(log_flush_exe.getEmittedBin());
// Also install the packed binaries to their FHS homes, so zig-out is a true image // Also install the packed binaries to their FHS homes, so zig-out is a true image
// of the filesystem — even though at boot they arrive inside the initial-ramdisk. // of the filesystem — even though at boot they arrive inside the initial-ramdisk.
@@ -498,6 +485,7 @@ pub fn build(b: *std.Build) void {
.{ usb_hid_mouse_exe, "system/drivers" }, .{ usb_hid_mouse_exe, "system/drivers" },
.{ usb_storage_exe, "system/drivers" }, .{ usb_storage_exe, "system/drivers" },
.{ fat_exe, "system/services" }, .{ fat_exe, "system/services" },
.{ log_flush_exe, "system/services" },
}) |entry| { }) |entry| {
const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } }); const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } });
b.getInstallStep().dependOn(&step.step); b.getInstallStep().dependOn(&step.step);
+17 -11
View File
@@ -82,6 +82,12 @@ Start with the north star:
- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on - **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on
fault isolation + live restart — the reincarnation-server + capability model that fault isolation + live restart — the reincarnation-server + capability model that
makes "if I break it, I can restart it" real. danos's core motivation. makes "if I break it, I can restart it" real. danos's core motivation.
- **[zig-self-hosting.md](zig-self-hosting.md) — running Zig on danos.** A design note
(not built yet) on making danos a real Zig target (`-target x86_64-danos`) and
eventually running the compiler on it. The key realisation: Zig 0.16 reduces an OS
port to **one seam** (`std.os.danos`), so we build `runtime.os` (→ that seam) plus a
thin `runtime.fs`, retire the `posix` shim, and follow a phased path to
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
Cutting across all of these: Cutting across all of these:
@@ -191,21 +197,22 @@ system/ → /system danos's own internals (the self-representation)
services/ init/ vfs/ device-manager/ system servers → /system/services (vfs/ holds services/ init/ vfs/ device-manager/ system servers → /system/services (vfs/ holds
vfs.zig, vfs-test.zig, protocol.zig) vfs.zig, vfs-test.zig, protocol.zig)
library/ → /lib libraries, one sub-directory each library/ → /lib libraries, one sub-directory each
runtime/ the danos-native runtime — the stable application ABI runtime/ the danos-native runtime + file API (fs) — the stable application ABI
posix/ POSIX/C compatibility, layered over runtime
boot/ → /boot the loaders boot/ → /boot the loaders
tools/ test/ host-side build + QEMU test harness tools/ test/ host-side build + QEMU test harness
``` ```
A sub-project exposes its **public interface as a module**: `system/services/vfs/` owns A sub-project exposes its **public interface as a module**: `system/services/vfs/` owns
the VFS wire protocol (`protocol.zig`, the `vfs-protocol` module), which the POSIX the VFS wire protocol (`protocol.zig`, the `vfs-protocol` module), which the runtime's
layer imports by name. `usb`/`block` drivers will expose their protocols the same way. file API (`runtime.fs`) imports by name. `usb`/`block` drivers expose their protocols the
same way.
`library/posix/` is special: it is the **one place** POSIX/C spellings are allowed There is **no POSIX/C compatibility layer today**: danos programs do file I/O through the
verbatim (`stat`, `O_CREAT`, `fopen`, `errno`). Everywhere else follows the danos danos-native `runtime.fs` (open/read/write/list over the VFS). A hand-rolled POSIX shim
naming rule with no exception — see [coding-standards.md](coding-standards.md). The (`library/posix/`) was retired as premature — the real POSIX/C surface will come later
POSIX layer calls the runtime, never the kernel's system calls directly, so it never from the `std.os.danos` seam (and, eventually, musl) when danos becomes a Zig target (see
appears in the private-ABI path. [zig-self-hosting.md](zig-self-hosting.md)). When it does, the foreign-ABI naming
exception in [coding-standards.md](coding-standards.md) applies to that seam.
## Source map ## Source map
@@ -229,8 +236,7 @@ appears in the private-ABI path.
| Framebuffer text console (mirrors to serial) | `system/kernel/console.zig` | | Framebuffer text console (mirrors to serial) | `system/kernel/console.zig` |
| In-kernel test cases | `system/kernel/tests.zig` | | In-kernel test cases | `system/kernel/tests.zig` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/IO-APIC/timer, serial, linker script) | `system/kernel/architecture/x86_64/` | | Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/IO-APIC/timer, serial, linker script) | `system/kernel/architecture/x86_64/` |
| danos-native runtime (`runtime`): syscall wrappers, heap, IPC, device access — the stable application ABI | `library/runtime/` | | danos-native runtime (`runtime`): syscall wrappers, heap, IPC, device access, the file API (`fs`) — the stable application ABI | `library/runtime/` |
| POSIX/C compatibility (`posix`): unistd, stdio — the one place POSIX names are allowed | `library/posix/` |
| System services (init, the VFS server + `protocol`, the device-manager) | `system/services/` | | System services (init, the VFS server + `protocol`, the device-manager) | `system/services/` |
| Device drivers, one sub-project each (`pci-bus`, `ps2-bus`, `usb-xhci-bus` bus drivers) | `system/drivers/` | | Device drivers, one sub-project each (`pci-bus`, `ps2-bus`, `usb-xhci-bus` bus drivers) | `system/drivers/` |
| Build + `run-x86-64` (QEMU/OVMF) | `build.zig` | | Build + `run-x86-64` (QEMU/OVMF) | `build.zig` |
+12 -9
View File
@@ -65,15 +65,18 @@ Three, and only three.
`errno`, `O_CREAT`. We don't get to rename `fwrite` to `fileWrite` — it wouldn't be `errno`, `O_CREAT`. We don't get to rename `fwrite` to `fileWrite` — it wouldn't be
`fwrite` any more. `fwrite` any more.
**This exception is scoped to one place: `library/posix/`.** A file under **This exception is scoped to a file that *is* a foreign ABI, and nothing else.**
`library/posix/` *is* the foreign ABI, so it keeps the ABI's spellings — that is the danos has no such file today: the old `library/posix/` compatibility shim was retired
whole rule for that directory. **Everywhere else, Zig/danos naming applies with no once its callers moved to the danos-native `runtime.fs`, since a hand-rolled POSIX
POSIX exception**, so there is nothing to get wrong: if you're not in layer is premature until danos actually needs it (see
`library/posix/`, expand it. A concept POSIX also has gets a danos name outside that [zig-self-hosting.md](zig-self-hosting.md)). The exception will apply again to the
layer — the VFS wire protocol carries a `FileStatus`, not a `Stat`, and a `create` `std.os.danos` seam when danos becomes a real Zig target — that module *is* the C-ABI
flag, not `O_CREAT`; `library/posix/` is what maps `stat`→`status` and `system` interface, so it keeps `open`/`read`/`errno`/`O_CREAT`. **Everywhere else,
`O_CREAT`→`create` at the boundary. (The `syscall` *wrappers* elsewhere are not an Zig/danos naming applies with no exception**: a concept POSIX also has gets a danos
exception to this — they wrap the private danos ABI, so they use danos names.) name — the VFS wire protocol carries a `FileStatus`, not a `Stat`, and a `create`
flag, not `O_CREAT`; the boundary is where `stat`→`status` and `O_CREAT`→`create` get
mapped. (The `syscall` *wrappers* elsewhere are not an exception — they wrap the
private danos ABI, so they use danos names.)
2. **Zig idioms are spelled the way Zig spells them.** Three names are the language's, 2. **Zig idioms are spelled the way Zig spells them.** Three names are the language's,
not ours, and are left alone: not ours, and are left alone:
+6 -6
View File
@@ -17,12 +17,12 @@ was a mistake) without inheriting the mechanism, the API, or the names. The nami
rule is danos's own and it is strict: plain words that communicate intent rule is danos's own and it is strict: plain words that communicate intent
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks (`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for (`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
a concept that already has one. Literal POSIX arrives later and lives elsewhere: a a concept that already has one. Literal POSIX arrives later and lives elsewhere: the
**musl-based C layer** (growing out of library/posix) that wires C programs to the `std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C
danos runtime — musl's syscall surface retargeted at danos system calls and IPC layer** on the same native surface (see [zig-self-hosting.md](zig-self-hosting.md)) —
protocols (files onto the VFS protocol, `sigaction`/`wait` onto this lifecycle, musl's syscall surface retargeted at danos system calls and IPC protocols (files onto
sockets onto whatever networking becomes). Ported programs see POSIX; the system the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever
underneath never does. networking becomes). Ported programs see POSIX; the system underneath never does.
## Why a standard vocabulary ## Why a standard vocabulary
+349
View File
@@ -0,0 +1,349 @@
# Running Zig on danos: the self-hosting roadmap
A design note (not built yet) on the path to making danos a **real Zig target** — a
target you can name (`-target x86_64-danos`) and, eventually, run the Zig compiler
itself on. It is forward-looking, like [vision.md](vision.md): it sets a direction
and the decisions that follow from it, so the code we write now bends toward it
instead of away.
This note deliberately does **not** cover a text editor or terminal. Those are
easier (single-process, I/O-bound) and fall out of the early phases here almost for
free; the hard, shaping problem is the standard-library surface, so that is what
this roadmap is about.
The analysis behind it was done against **Zig 0.16** (the pinned toolchain). Zig's
standard library moves between releases — especially the parts described here — so
treat upstream references as "the shape in 0.16.x," and expect to re-check them on a
toolchain bump.
## The win condition
danos runs the Zig compiler when a bare
```
zig build-exe hello.zig
```
completes **on danos** and produces a runnable danos binary. Note the milestone is
`build-exe`, not `zig build`: the `zig build` runner spawns child processes (the
build steps), which needs a whole process-control surface danos does not have yet.
A single `build-exe` needs none of that (see Phase 3). Reaching `build-exe` is
"self-hosting"; reaching `zig build` is a later, separate lift.
### Non-goals
- **No Linux syscall/ABI emulation.** danos will not implement the Linux `syscall`
interface so that stock `x86_64-linux` binaries run. That is a permanent
compatibility treadmill and it inverts the microkernel design — explicitly out.
- **No musl port yet.** A musl libc port is a reasonable *later* effort (it unlocks
the C ecosystem), but it is not on the critical path to Zig-on-danos, and it is
deferred. The roadmap below is arranged so the work still pays off if musl ever
happens (see "The same surface, twice").
- **Editor/terminal are out of scope for this note** (they are downstream of Phase 1).
**On FFI.** Foreign-function interop splits the same way as the doors below. Zig-level
and C-ABI-*exposing* FFI (`extern`, `callconv(.c)`, C-ABI structs) work on a real target
immediately — and the `std.os.danos` seam is C-ABI-shaped by construction, so it is
FFI-friendly from the start. *Consuming* C libraries (`@cImport`, linking archives) is
the part that needs a libc + headers, i.e. the deferred musl door. So an eventual FFI
need reinforces keeping that door open; it does not change the plan.
## The realization that shapes everything: 0.16 gives us *one* seam
The instinct "to target Zig we'd have to reimplement all the `std` namespaces" was
how older Zig worked. Zig 0.16 (post-"writergate") is far kinder:
- **`std.fs` is essentially gone.** It is now path helpers plus deprecated aliases;
there is no `std.fs.File`, `std.fs.Dir`, or `std.fs.cwd()`. File and directory
work goes through **`std.Io`** — a single runtime **vtable** (`Io.zig`) of
function pointers handed to `main` as `std.process.Init.io`. `std.Io.File` and
`std.Io.Dir` are thin forwarders to that vtable. `Io.zig` and the `fs` shim carry
**zero** per-OS branches.
- **`std.posix` is one generic body** parameterised over a single `system` module.
With no libc, `system` resolves **per target OS**: `.linux => std.os.linux`,
`.plan9 => std.os.plan9`, and so on. The generic `std.posix.read`/`write`/`open`
bodies are just `system.read(...)` plus an errno switch — *identical for every
OS*. The only variable is what `system` binds to.
- **`std.os.<tag>`** (e.g. `std/os/linux.zig`) is therefore the real porting seam: a
low-level, C-ABI-shaped module of `read/write/open/close/lseek/mmap/clock/exit/…`
plus an `errno` enum and the constant tables (`O_*`, `CLOCK_*`, `S_*`).
Put together: **to port danos we write `std.os.danos` once** — the ~30-operation
seam — and the whole `std.posix` / `std.fs` / `std.Io` tower above it lights up
generically, because none of it branches on the OS. That is a dramatically smaller
and more contained target than "reimplement the namespaces."
## Three doors, and why we take the first
| Door | What it is | Verdict |
|------|-----------|---------|
| **1. Implement the std seam** (`std.os.danos`) | Write the ~30-op `system` module over danos's native ABI + VFS; the generic std tower lights up. | **Take this.** The only door that touches neither C nor the Linux ABI. |
| **2. Port musl** | Port musl libc to danos, link Zig against it. | Defer. Good later for the *C* ecosystem; barely helps *Zig* (std only uses libc on the libc-linked path). |
| **3. Emulate the Linux ABI** | Implement Linux syscalls so stock linux binaries run. | Reject. Bottomless compatibility treadmill; against the design. |
### The same surface, twice
Doors 1 and 2 are the **same native surface at different layers**. `std.posix.read`
is `system.read(...)` + an errno switch *regardless of OS* — the only question is
whether `system` is **`std.os.danos` (Zig)** or **musl (C)**. Either way, the set of
danos-facing operations you must implement is the *same* ~30 ops, all bottoming out
in danos's native syscalls + the VFS/FAT server.
So the runtime work below is **not throwaway** if musl ever happens: you are building
the danos-native implementations of that surface either way. Door 1 just packages
them as Zig; a future musl re-uses the identical kernel/VFS operations underneath. The
two symmetries worth keeping in mind: doors 1 and 2 converge at the **top** (identical
POSIX surface); doors 2 and 3 converge at the **bottom** (unmodified musl needs the
Linux syscall ABI). Door 1 is the only one that avoids both C and Linux.
### A fork is table stakes — for any door
`std.Target.Os.Tag` is a **closed enum** baked into the compiler binary *and* into
the `std` linked with every program; `-target x86_64-danos` resolves through it. So
adding `danos` as a name requires patching and rebuilding the compiler — even the
musl door needs this. "Fork Zig" is therefore not an extra cost unique to door 1; it
is the price of admission for *any* real target. What door 1 adds on top is small and
localised (below).
## The architecture decision: `runtime.os` + `runtime.fs`, and retire `posix`
danos already has the right split ([the private-ABI boundary](../README.md)): the
kernel exposes a minimal syscall ABI ([syscall.md](syscall.md)); the **`runtime`**
library is the stable, danos-native application ABI. What this roadmap adds:
- **`runtime.os` — the seam.** A C-ABI-shaped module of the ~30 operations
(`read/write/open/close/lseek/mmap/munmap/clock/exit/…`) + an errno enum + the
constant tables, each backed by danos's native syscalls and the VFS. **Structure it
to mirror `std/os/linux.zig`.** This is the load-bearing, *non-throwaway* artifact:
when we fork Zig, `runtime.os` is copy-pasted (near-verbatim) into `std.os.danos`.
- **`runtime.fs` — the thin native file API** danos programs use *today*, layered
over `runtime.os`. It is also the concrete backing for the `std.Io` vtable's
file-write entry once we're a real target, which is why program stdout, diagnostics,
and file writes should all be *decided once at that seam* rather than as bespoke
per-call helpers (see "How this informs decisions now").
**Do not hand-mirror the high-level std namespaces.** `std.fs`/`std.Io`/`std.process`
are generic and OS-agnostic; once `std.os.danos` exists and we fork, upstream *gives*
them to danos for free. Hand-writing `runtime.std.fs` to imitate them would be
redundant the day the fork works, and it would chase a moving target (0.16's `std.Io`
is large and still shifting). Build the seam well; take the tower for free.
**Why not a library called `std`?** Because `@import("std")` resolves to the
compiler-provided standard library; a user module named `std` would *shadow* it for
anything that imports it that way. That is the real reason the seam lives *inside* a
forked std as `std/os/danos.zig`, not as a `runtime.std` library — and why danos's end
state (`@import("std")` just working, and knowing danos) is the most natively Zig it can
be. `runtime.os` is only the interim staging ground: developed against the stock
toolchain so Phase 1 need not wait on the fork, then promoted near-verbatim into the
fork's `std/os/danos.zig`.
### Retire `library/posix`
The `posix` compatibility layer (`unistd`, `stdio`) was the right instinct too early.
Its whole value is POSIX *spellings* for POSIX software — and danos has no POSIX
software; every current caller is danos-native code that could use `runtime.fs`
directly. The real POSIX story arrives later and from elsewhere (musl, or upstream
`std`'s own posix over `std.os.danos`), which supersedes a hand-rolled shim. So it is
premature abstraction that adds a "which layer do I use?" fork with no payoff yet.
Its footprint is tiny: **five** call sites, all `unistd` file operations —
`system/services/fat/fat.zig` (`mount`), the `vfs-test` and `fat-test` clients, and
(from the boot-log work) `init.zig` and `log-flush.zig`. `stdio.zig` is dead — nothing
imports it. The plan: build `runtime.fs`, migrate those five to it, delete
`library/posix/`, and drop the `posix` module from `build.zig`'s `addUserBinary`.
## Where danos stands: coverage vs. the gaps
What the seam needs, and what danos already provides:
| std need | danos today | Gap |
|----------|-------------|-----|
| open / read / write / close / lseek | VFS (via the current `unistd`, → `runtime.fs`) | none — repackage |
| directory read (`getdents`) | VFS `readdir` | none — repackage |
| mmap / munmap | native syscalls ([abi.zig](../system/abi.zig)) | none |
| page allocator | over `mmap`, via `root.os.heap.page_allocator` override | ~30-line hook |
| monotonic clock | `clock` syscall | none |
| args / argv | SysV entry stack ([sysv.md](sysv.md)), `runtime.process.Init` | none |
| stdout / stderr | `debug_write` today | wire fd 1/2 to a console **byte** stream |
| mkdir / unlink / rename / truncate | done — engine + VFS + `runtime.fs` (Phase 2) | — |
| stat fields | `{size, kind, mtime}` | **mode / inode** still missing (cache validity) |
| wall-clock / realtime | done — `wall_clock` syscall (CMOS RTC, Phase 2d) | — |
| **environment variables** | `Init` has no env field | missing (can start empty) |
| **cwd / chdir** | paths are absolute or bare | missing (no cwd anchor) |
| **entropy / random** | — | missing (needed behind `vtable.random`) |
| process spawn + exit status | `system_spawn` starts a *named ramdisk binary*; `ExitReason` is a *category* | no exec-of-path, no numeric `WEXITSTATUS` |
| threads | one thread per process | avoided via `-fsingle-threaded` (below) |
| symlinks | `NodeKind` has the tag; unimplemented | low priority |
The clustering is clear: reads and memory are basically done; the real work is
**filesystem mutation + richer stat + wall-clock**, and a few small seam pieces
(page-allocator hook, stdio bytes, entropy). Process spawning and threads are
side-stepped entirely for a single `build-exe`.
## The roadmap
### Phase 0 — Make `danos` a real target
**Host, target, self-host — keep the three roles straight.** The *host* is where the
compiler runs (your mac + linux dev machines); the *target* is what it emits (`danos`);
and eventually danos becomes a host too (self-hosting — the win condition). So the move
is: fork the compiler, build it **for** your dev hosts, and teach it to **cross-compile
to** danos. You already do this — danos is cross-compiled `freestanding` from your dev
host today; Phase 0 swaps that `freestanding` target for a real `x86_64-danos` one, which
is what unlocks the native `std`.
**Why a compiler fork, not just a `--zig-lib-dir` override.** `std.Target.Os.Tag` is a
*closed enum compiled into the compiler binary*, so `-target x86_64-danos` will not even
parse unless the compiler itself knows the tag. Overriding the std lib directory alone
cannot add a target — and there is no libc-only shortcut (a future musl needs the same
patch). The only alternative, staying on `freestanding` + hand-shims, is exactly the
non-native feel we are leaving: `@import("std")` there is stubbed, not real.
**The fork.** Clone `ziglang/zig` at the pinned 0.16 tag; build it with a stock
same-version `zig` (`zig build` in the tree — a standard, LLVM-pulling, roughly one-time
build); point danos's `build.zig`/CI at the resulting binary. Four localised patches:
- add `danos` to `std.Target.Os.Tag`, in the "no version range" group alongside
plan9/serenity;
- add `danos` to the freestanding/other **no-op `_start` list** in `std`'s `start.zig`,
so std does *not* emit its own System-V `_start` — danos keeps owning the entry shim
and `Init`/argv construction it already builds ([sysv.md](sysv.md));
- wire the `system` selector `.danos => std.os.danos` in `std.posix`;
- add `std/os/danos.zig` — **the seam itself**, promoted near-verbatim from the
`runtime.os` developed first in Phase 1 (against the stock toolchain, so the fork is
not a prerequisite for starting).
This is the fork treadmill we accept once. Keep the patch set tiny and `else`-friendly,
pin to one 0.16.x, and rebase on point releases.
### Phase 1 — `runtime.os` read-side + allocator + stdio + cwd; retire `posix`
Author `runtime.os` (→ `std.os.danos`): the `errno` enum, the constant tables, and
the C-convention `read / write / open / openat / close / lseek / mmap / munmap /
exit`, each returning result-or-`-errno`. Most backing already exists (VFS + native
mmap + clock).
- Provide `page_allocator` via `root.os.heap.page_allocator` (a thin override over
danos `mmap`). This sits **outside** the `std.Io` vtable, so it is wired separately.
- Wire fd 0/1/2 to a console **byte** stream (today output only reaches `debug_write`;
input is structured `InputEvent` IPC — a byte tty is a new, small thing in both
directions).
- Add a `getcwd`/`chdir` anchor so `std.fs.cwd()`-style resolution has something to
resolve against.
- Build `runtime.fs` over `runtime.os`; migrate the five `posix` callers to it; delete
`library/posix/` and drop its build module.
After Phase 1, the surface an editor or terminal needs (open/read/write/close/lseek/
readdir/isatty/args/exit) exists. Those are downstream and out of scope here.
### Phase 2 — Filesystem mutation + real stat (the compiler's cache tower)
danos's biggest genuine gap, and the correctness-critical one:
- Add **mkdir / unlink / rename / truncate** to *both* the VFS wire protocol
([protocol.zig](../system/services/vfs/protocol.zig)) and the FAT engine
([engine.zig](../system/services/fat/engine.zig)), then expose them via `runtime.os`.
- Extend `stat` beyond `{size, kind}` to carry **mtime + inode + mode** — `std`'s file
stat needs them for build-cache validity — which in turn needs **wall-clock** time
(danos is monotonic-only today; an RTC/time service is the dependency).
Because `std.fs`/`std.Io` have no per-OS branches, finishing this in `runtime.os`
lights up the whole file tower for the compiler at once. Environment can stay an empty
map until the kernel populates a non-empty `envp`.
**Status — Phase 2 complete.** `truncate` (O_TRUNC, closing the boot-log stale-tail
bug), `mkdir`, `unlink`, and `rename` are all wired through the FAT engine, the VFS
protocol + router, and `runtime.fs` (`makeDirectory` / `remove` / `rename`) —
host-tested and QEMU-tested (`fat-mutations` + `fat-rename` make a directory, write+read
a file in it, rename it, then remove it through the mount). `removeFile` and `rename`
are LFN-aware; `rename` is same-directory + 8.3 (cross-directory and long-name-
preserving rename are noted limitations). Wall-clock is now a kernel syscall
(`wall_clock`, a CMOS-RTC read anchored to the monotonic clock), and the FAT engine
stamps and reports **mtime** — `stat` / `runtime.fs.Attributes` carry a real
modification time (the `fat-mtime` case reads it back within seconds of the host clock).
The remaining `stat` fields, `mode`/`inode`, are deferred (not needed until the
compiler's cache layer wants them). **Everything past here is gated on Phase 0 (the
fork):** the `runtime.os` seam, `cwd`, stdio-as-fds, and the compiler bring-up.
### Phase 3 — Single-threaded, self-linked compiler bring-up
Build the compiler with **two load-bearing flags**:
- **`-fsingle-threaded`** removes `std.Thread` entirely — `Thread.spawn` is a hard
compile error under it, and `std.Io`'s threaded backend runs inline. danos being
one-thread-per-process is therefore **not** a blocker. Parallel codegen is a
throughput optimisation, not a correctness requirement.
- **`-fno-llvm -fno-lld`** keeps codegen and linking **in-process** (the self-hosted
x86-64 backend + self-linker), so a single `build-exe` **never forks a child**. That
is what lets us defer the entire spawn/exec/wait surface.
Then supply the few remaining seam pieces: `now` (wrap the danos clock), an entropy
source behind `vtable.random` (`randomSecure` can alias it initially — low volume, for
temp-file names and hashmap seeds), and the Phase-2 mkdir/rename/unlink for cache dir
trees and atomic temp-then-rename output.
**Explicitly deferred** (not on the `build-exe` path): child-process spawn/exec (only
`zig build` and external tools need it), `std.Thread`, `fsync` (FAT is write-through
today), symlinks, and musl.
## Risks and gotchas
- **The std-fork rebase treadmill is the main ongoing cost.** A new OS tag touches the
same broad file set plan9/serenity touch (hundreds of `native_os` sites, plus
"unsupported OS" `@compileError` dead-ends a new tag must be routed around), and the
entire `std.Io` layer is new in 0.16 and still moving. Stay pinned to one 0.16.x,
keep additions localised and `else`-friendly. Watch the closed-enum gotcha: adding
`danos` to `Os.Tag` can break existing *exhaustive* switches that lack an `else`, so
expect to touch switch sites beyond the ones you implement.
- **Single-threaded is load-bearing.** The "no `std.Thread`" simplification rests
entirely on `-fsingle-threaded`. If a dependency or flag flips threading back on, you
inherit an unescapable compile error (no root-hook exists) — the only outs are a full
thread-impl fork or linking libc for pthreads. Keep `single_threaded` asserted end to
end.
- **In-process linking is load-bearing.** Reaching the compiler without fork/exec
depends on `-fno-llvm -fno-lld`. The moment you shell out to LLD/`ld`, you need the
full `spawn`/`wait` surface — the hardest microkernel piece — and danos's
`system_spawn` only starts a *named ramdisk binary*, not exec of an arbitrary path.
Verify the self-hosted backend covers the target output before assuming child
processes are optional.
- **The shim cannot host the compiler.** danos's current `runtime`/`posix` is fine for
danos's *own* native programs, but the compiler `import`s *upstream* `std`, which on
a non-target hits the void `system` stub. So the compiler forces the real target
(Phase 0's fork). Do not over-invest in extending the hand-shim for compiler
purposes; put that effort into `runtime.os` + the VFS/FAT operations, which both the
fork *and* a future musl consume.
- **`"w"`/`O_CREAT` does not truncate — a silent-corruption bug on this road.** The FAT
engine's `writeFile` only *grows* `node.size`, so overwriting a shorter file leaves
trailing garbage. Harmless for the boot log today, but for a compiler it means
**corrupt `.o`/cache files that look like nondeterministic compiler bugs.** Land
`truncate` (Phase 2) before the compiler ever writes cache.
- **Exit status is categorical, not numeric.** `process_exit_reason` returns an
`ExitReason` *category*, not a numeric code (`WEXITSTATUS`). Fine while spawn is
stubbed; the day `zig build` or external tools arrive, plan a kernel exit-record
extension — do not let it surprise you.
## How this informs decisions now
Two current decisions fall out of this roadmap:
1. **The `runtime.fs` / `std.Io` question resolves at the vtable seam.** Because 0.16
routes *all* output through the `std.Io` vtable's file-write entry, and stdout/stderr
are just `File`s with well-known handles, build `runtime.fs` (and the console stdout)
as the concrete backing for that entry — not as a bespoke `std.Io.Writer`-only shim.
Decide it once, at the seam, and program stdout, diagnostics, and file writes all
flow through the same danos VFS/console path.
2. **The boot-log `truncate` caveat is now fixed** (Phase 2a). It was the same
`writeFile`-only-grows gap that on the self-hosting road would corrupt build output;
`engine.truncate` + an O_TRUNC open flag now free the old chain so a shorter rewrite
leaves no stale tail, and the boot-log flush opens with it.
## Related
- [vision.md](vision.md) — the north star this serves.
- [syscall.md](syscall.md) — the kernel↔runtime ABI `runtime.os` is built on.
- [sysv.md](sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs.
- [ipc.md](ipc.md) — the IPC the VFS/FAT operations travel over.
- [danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md) — the
filesystem layout the file surface serves.
- [coding-standards.md](coding-standards.md) — danos naming (why the compat spellings
are confined, and now retired).
-13
View File
@@ -1,13 +0,0 @@
//! DanOS's POSIX / C compatibility layer — `unistd`, `stdio`, and (later) the C
//! `errno` / `struct stat` / `extern "C"` surface. This is the *one* place POSIX and
//! C spellings are allowed to appear verbatim (see docs/coding-standards.md): a file
//! under library/posix/ *is* the foreign ABI, so it keeps the ABI's names. Everything
//! it touches on the danos side (the VFS protocol, the runtime) uses danos names,
//! which this layer translates to at the boundary.
//!
//! It is layered strictly *over* the runtime: it calls the runtime's IPC and heap,
//! never the kernel's system calls directly. danos-native applications use the
//! runtime; this exists so *POSIX* software can too.
pub const unistd = @import("unistd.zig");
pub const stdio = @import("stdio.zig");
-115
View File
@@ -1,115 +0,0 @@
//! A small C stdio layer over the POSIX-style file API (unistd.zig). Unbuffered
//! for now — each fread/fwrite is one VFS round trip; an internal buffer (fewer
//! IPC calls) is a later optimisation. Both a Zig-callable API and `extern "C"`
//! symbols are provided, so Zig and future C programs share it.
const std = @import("std");
const unistd = @import("unistd.zig");
const heap = @import("runtime").heap;
pub const SEEK_SET = unistd.SEEK_SET;
pub const SEEK_CURRENT = unistd.SEEK_CURRENT;
pub const SEEK_END = unistd.SEEK_END;
/// A C `FILE`: an fd plus sticky end-of-file / error flags. Allocated on the
/// heap; `fclose` frees it.
pub const FILE = extern struct {
fd: i32,
eof: c_int = 0,
err: c_int = 0,
};
fn flagsFor(mode: []const u8) u32 {
if (mode.len == 0) return 0;
return switch (mode[0]) {
'w', 'a' => unistd.O_CREAT,
else => 0,
};
}
/// Open `path` in `mode` ("r"/"w"/"a", '+' ignored for now). Returns null on error.
pub fn fopen(path: []const u8, mode: []const u8) ?*FILE {
const fd = unistd.open(path, flagsFor(mode));
if (fd < 0) return null;
const f = heap.allocator().create(FILE) catch {
unistd.close(fd);
return null;
};
f.* = .{ .fd = fd };
if (mode.len > 0 and mode[0] == 'a') _ = unistd.lseek(fd, 0, unistd.SEEK_END);
return f;
}
pub fn fclose(f: *FILE) c_int {
unistd.close(f.fd);
heap.allocator().destroy(f);
return 0;
}
/// Read `size*nmemb` bytes; returns the number of whole items read.
pub fn fread(buffer: []u8, size: usize, nmemb: usize, f: *FILE) usize {
const total = size * nmemb;
if (total == 0) return 0;
const n = unistd.read(f.fd, buffer[0..@min(buffer.len, total)]);
if (n <= 0) {
f.eof = 1;
return 0;
}
return @as(usize, @intCast(n)) / size;
}
/// Write `size*nmemb` bytes; returns the number of whole items written.
pub fn fwrite(data: []const u8, size: usize, nmemb: usize, f: *FILE) usize {
const total = @min(data.len, size * nmemb);
if (total == 0) return 0;
const n = unistd.write(f.fd, data[0..total]);
if (n <= 0) {
f.err = 1;
return 0;
}
return @as(usize, @intCast(n)) / size;
}
pub fn fseek(f: *FILE, off: i64, whence: u32) c_int {
f.eof = 0;
return if (unistd.lseek(f.fd, off, whence) < 0) -1 else 0;
}
pub fn ftell(f: *FILE) i64 {
return unistd.lseek(f.fd, 0, unistd.SEEK_CURRENT);
}
pub fn rewind(f: *FILE) void {
_ = fseek(f, 0, SEEK_SET);
}
pub fn feof(f: *FILE) c_int {
return f.eof;
}
pub fn ferror(f: *FILE) c_int {
return f.err;
}
pub fn fputs(s: []const u8, f: *FILE) c_int {
return if (unistd.write(f.fd, s) < 0) -1 else 0;
}
pub fn fputc(c: u8, f: *FILE) c_int {
const b = [_]u8{c};
return if (unistd.write(f.fd, &b) == 1) c else -1;
}
pub fn fgetc(f: *FILE) c_int {
var b: [1]u8 = undefined;
const n = unistd.read(f.fd, &b);
if (n <= 0) {
f.eof = 1;
return -1; // EOF
}
return b[0];
}
// Real `extern "C"` symbols (fopen/fread/fseek/...) — with a C-string signature
// distinct from the Zig slice API above — land with the first C program, wired
// via @export so they don't collide with these Zig names.
-206
View File
@@ -1,206 +0,0 @@
//! POSIX-style file API for user programs — the low level under C stdio. Files
//! are named objects served by the user-space VFS server (system/services/vfs/vfs.zig); each
//! call marshals a request, IPC_Calls the VFS, and unmarshals the reply. The
//! kernel knows nothing of files or fds — the fd table lives here, per process.
const std = @import("std");
const protocol = @import("vfs-protocol");
const ipc = @import("runtime").ipc;
pub const O_CREAT = protocol.create;
pub const SEEK_SET: u32 = 0;
pub const SEEK_CURRENT: u32 = 1;
pub const SEEK_END: u32 = 2;
// Resolve (and cache) the VFS server endpoint, looked up by well-known id.
var vfs_handle: usize = 0;
var vfs_resolved = false;
fn vfs() ?usize {
if (!vfs_resolved) {
vfs_handle = ipc.lookup(.vfs) orelse return null;
vfs_resolved = true;
}
return vfs_handle;
}
const maximum_fds = 32;
const Fd = struct { used: bool = false, node: u64 = 0, offset: u64 = 0 };
var fds = [_]Fd{.{}} ** maximum_fds;
fn allocFd() ?usize {
for (&fds, 0..) |*f, i| {
if (!f.used) {
f.* = .{ .used = true };
return i;
}
}
return null;
}
const Result = struct { reply: protocol.Reply, payload: []u8 };
/// One request/reply round trip: [Request header][send payload] -> VFS ->
/// [Reply header][receive payload]. The receive payload is written into `out`.
fn transact(request: protocol.Request, send: []const u8, out: []u8) ?Result {
const h = vfs() orelse return null;
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const slen = @min(send.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..slen], send[0..slen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(h, message[0 .. protocol.request_size + slen], &rbuf) catch return null;
if (n < protocol.reply_size) return null;
const reply = std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]);
const rpl = @min(n - protocol.reply_size, out.len);
@memcpy(out[0..rpl], rbuf[protocol.reply_size..][0..rpl]);
return .{ .reply = reply, .payload = out[0..rpl] };
}
/// Open (or create, with O_CREAT) `path`; returns an fd or -1.
pub fn open(path: []const u8, flags: u32) i32 {
const fd = allocFd() orelse return -1;
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = flags };
const r = transact(request, path, &.{}) orelse {
fds[fd].used = false;
return -1;
};
if (r.reply.status != 0) {
fds[fd].used = false;
return -1;
}
fds[fd] = .{ .used = true, .node = r.reply.node, .offset = 0 };
return @intCast(fd);
}
fn fdPtr(fd: i32) ?*Fd {
if (fd < 0 or fd >= maximum_fds) return null;
const f = &fds[@intCast(fd)];
return if (f.used) f else null;
}
/// Read up to `buffer.len` bytes at the current offset; returns the count or -1.
pub fn read(fd: i32, buffer: []u8) isize {
const f = fdPtr(fd) orelse return -1;
const want: u32 = @intCast(@min(buffer.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .read, .node = f.node, .offset = f.offset, .len = want, .flags = 0 };
const r = transact(request, &.{}, buffer) orelse return -1;
if (r.reply.status != 0) return -1;
f.offset += r.reply.len;
return @intCast(r.reply.len);
}
/// Write `data` at the current offset; returns the count or -1.
pub fn write(fd: i32, data: []const u8) isize {
const f = fdPtr(fd) orelse return -1;
const want: u32 = @intCast(@min(data.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .write, .node = f.node, .offset = f.offset, .len = want, .flags = 0 };
const r = transact(request, data[0..want], &.{}) orelse return -1;
if (r.reply.status != 0) return -1;
f.offset += r.reply.len;
return @intCast(r.reply.len);
}
/// Reposition the fd's offset. Returns the new offset or -1. (SEEK_END needs the
/// file size, which `stat` provides; handled by fetching it here.)
pub fn lseek(fd: i32, off: i64, whence: u32) i64 {
const f = fdPtr(fd) orelse return -1;
const base: i64 = switch (whence) {
SEEK_SET => 0,
SEEK_CURRENT => @intCast(f.offset),
SEEK_END => blk: {
const request = protocol.Request{ .operation = .status, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
var sbuf: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &sbuf) orelse return -1;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return -1;
const st = std.mem.bytesToValue(protocol.FileStatus, sbuf[0..@sizeOf(protocol.FileStatus)]);
break :blk @intCast(st.size);
},
else => return -1,
};
const pos = base + off;
if (pos < 0) return -1;
f.offset = @intCast(pos);
return pos;
}
/// Stat `path`. Returns 0 or -1.
pub fn stat(path: []const u8, out: *protocol.FileStatus) i32 {
// Open, stat by node, close — simple and enough for now.
const fd = open(path, 0);
if (fd < 0) return -1;
defer close(fd);
const f = fdPtr(fd).?;
const request = protocol.Request{ .operation = .status, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
var sbuf: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &sbuf) orelse return -1;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return -1;
out.* = std.mem.bytesToValue(protocol.FileStatus, sbuf[0..@sizeOf(protocol.FileStatus)]);
return 0;
}
/// Close an fd (best effort — tells the VFS to release the open file).
pub fn close(fd: i32) void {
const f = fdPtr(fd) orelse return;
const request = protocol.Request{ .operation = .close, .node = f.node, .offset = 0, .len = 0, .flags = 0 };
_ = transact(request, &.{}, &.{});
f.used = false;
}
/// Mount a filesystem backend (its server endpoint) at absolute path `target`;
/// the VFS then routes every path under `target` to that backend. Returns 0 or
/// -1. This is the one call that hands the VFS a capability (the backend).
pub fn mount(target: []const u8, backend: usize) i32 {
const h = vfs() orelse return -1;
const request = protocol.Request{ .operation = .mount, .node = 0, .offset = 0, .len = @intCast(target.len), .flags = 0 };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const tlen = @min(target.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..tlen], target[0..tlen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const result = ipc.callCap(h, message[0 .. protocol.request_size + tlen], &rbuf, backend) catch return -1;
if (result.len < protocol.reply_size) return -1;
return if (std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]).status == 0) 0 else -1;
}
/// A directory entry filled by `readdir`.
pub const DirEntry = struct {
kind: u32 = 0, // a protocol.NodeKind
size: u64 = 0,
name_buffer: [64]u8 = undefined,
name_len: usize = 0,
pub fn name(self: *const DirEntry) []const u8 {
return self.name_buffer[0..self.name_len];
}
};
/// Open a directory for reading with `readdir`. Returns an fd or -1.
pub fn opendir(path: []const u8) i32 {
return open(path, protocol.directory);
}
/// Read the next entry of a directory fd into `entry`; returns false at EOF or on
/// error. Advances the fd's cursor by one entry.
pub fn readdir(fd: i32, entry: *DirEntry) bool {
const f = fdPtr(fd) orelse return false;
const request = protocol.Request{ .operation = .readdir, .node = f.node, .offset = f.offset, .len = 0, .flags = 0 };
var buffer: [protocol.message_maximum]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return false;
if (r.reply.status != 0 or r.reply.len == 0) return false; // error or EOF
if (r.payload.len < protocol.directory_entry_size) return false;
const header = std.mem.bytesToValue(protocol.DirectoryEntry, r.payload[0..protocol.directory_entry_size]);
entry.kind = header.kind;
entry.size = header.size;
const source = r.payload[protocol.directory_entry_size..];
const nlen = @min(@min(@as(usize, header.name_len), source.len), entry.name_buffer.len);
@memcpy(entry.name_buffer[0..nlen], source[0..nlen]);
entry.name_len = nlen;
f.offset += 1;
return true;
}
/// Close a directory fd (same as `close`).
pub fn closedir(fd: i32) void {
close(fd);
}
+274
View File
@@ -0,0 +1,274 @@
//! runtime.fs — the danos-native file API. A program opens, reads, writes, and
//! lists files served by the user-space VFS (system/services/vfs), each call
//! marshalling a vfs-protocol request over IPC. This is the danos-native layer
//! danos programs use directly; it is also where the file operations that later
//! become `std.os.danos` are staged (see docs/zig-self-hosting.md). It replaces
//! the old POSIX `unistd` shim — a compatibility spelling danos does not need yet.
//!
//! Handles are *values*, not entries in a global descriptor table: a `File` /
//! `Directory` owns its VFS node id and (for files) a byte offset. So there is no
//! per-process fd limit and no shared table to synchronise — the danos-native
//! shape, unlike the POSIX fd model the old shim emulated.
const std = @import("std");
const ipc = @import("ipc.zig");
const protocol = @import("vfs-protocol");
/// The kind of a filesystem node — re-exported so a caller need not import the
/// wire protocol.
pub const Kind = protocol.NodeKind;
/// A node's metadata (the answer to a status request).
pub const Attributes = struct {
size: u64,
kind: Kind,
/// Modification time — Unix epoch seconds, UTC. 0 if the filesystem has none.
mtime: u64 = 0,
};
// Map a wire `NodeKind` value to the enum, defaulting anything unrecognised to
// `.regular` (the server is trusted, but a value outside the enum would be
// illegal to `@enumFromInt` directly).
fn kindFromWire(value: u32) Kind {
return switch (value) {
@intFromEnum(Kind.directory) => .directory,
@intFromEnum(Kind.character_device) => .character_device,
@intFromEnum(Kind.block_device) => .block_device,
@intFromEnum(Kind.symbolic_link) => .symbolic_link,
@intFromEnum(Kind.fifo) => .fifo,
@intFromEnum(Kind.socket) => .socket,
else => .regular,
};
}
/// How to open a path.
pub const OpenOptions = struct {
/// Create the file if it does not exist.
create: bool = false,
/// Open a directory node (for listing) rather than a file.
directory: bool = false,
/// Truncate an existing file to zero length on open (O_TRUNC) — replace its
/// contents rather than overwriting in place.
truncate: bool = false,
fn wireFlags(self: OpenOptions) u32 {
var f: u32 = 0;
if (self.create) f |= protocol.create;
if (self.directory) f |= protocol.directory;
if (self.truncate) f |= protocol.truncate;
return f;
}
};
// The VFS server endpoint, looked up once by well-known id and cached.
var vfs_handle: ipc.Handle = 0;
var vfs_resolved = false;
fn vfs() ?ipc.Handle {
if (!vfs_resolved) {
vfs_handle = ipc.lookup(.vfs) orelse return null;
vfs_resolved = true;
}
return vfs_handle;
}
const Result = struct { reply: protocol.Reply, payload: []u8 };
// One request/reply round trip: [Request header][send payload] -> VFS ->
// [Reply header][receive payload]. The receive payload lands in `out`.
fn transact(request: protocol.Request, send: []const u8, out: []u8) ?Result {
const h = vfs() orelse return null;
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const slen = @min(send.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..slen], send[0..slen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(h, message[0 .. protocol.request_size + slen], &rbuf) catch return null;
if (n < protocol.reply_size) return null;
const reply = std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]);
const rpl = @min(n - protocol.reply_size, out.len);
@memcpy(out[0..rpl], rbuf[protocol.reply_size..][0..rpl]);
return .{ .reply = reply, .payload = out[0..rpl] };
}
/// An open file: a VFS node plus a byte cursor. Read and write advance the cursor.
pub const File = struct {
node: u64,
offset: u64 = 0,
/// Read up to `buffer.len` bytes at the current offset; returns the count, or
/// null on error.
pub fn read(self: *File, buffer: []u8) ?usize {
const want: u32 = @intCast(@min(buffer.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .read, .node = self.node, .offset = self.offset, .len = want, .flags = 0 };
const r = transact(request, &.{}, buffer) orelse return null;
if (r.reply.status != 0) return null;
self.offset += r.reply.len;
return r.reply.len;
}
/// Write `data` at the current offset; returns the count written. A single
/// call is capped at the VFS payload size, so the return may be short — use
/// `writeAll` to write the whole slice. Null on error.
pub fn write(self: *File, data: []const u8) ?usize {
const want: u32 = @intCast(@min(data.len, protocol.maximum_payload));
const request = protocol.Request{ .operation = .write, .node = self.node, .offset = self.offset, .len = want, .flags = 0 };
const r = transact(request, data[0..want], &.{}) orelse return null;
if (r.reply.status != 0) return null;
self.offset += r.reply.len;
return r.reply.len;
}
/// Write all of `data`, looping past the per-call payload cap. Returns the
/// total written, or null if a write failed before any progress.
pub fn writeAll(self: *File, data: []const u8) ?usize {
var written: usize = 0;
while (written < data.len) {
const n = self.write(data[written..]) orelse return if (written == 0) null else written;
if (n == 0) return written; // no forward progress; stop rather than spin
written += n;
}
return written;
}
/// Move the read/write cursor to an absolute byte position.
pub fn seekTo(self: *File, position: u64) void {
self.offset = position;
}
/// This file's metadata.
pub fn attributes(self: *File) ?Attributes {
const request = protocol.Request{ .operation = .status, .node = self.node, .offset = 0, .len = 0, .flags = 0 };
var buffer: [@sizeOf(protocol.FileStatus)]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return null;
if (r.reply.status != 0 or r.payload.len < @sizeOf(protocol.FileStatus)) return null;
const status = std.mem.bytesToValue(protocol.FileStatus, buffer[0..@sizeOf(protocol.FileStatus)]);
return .{ .size = status.size, .kind = kindFromWire(status.kind), .mtime = status.mtime };
}
/// Release the VFS's open handle for this file.
pub fn close(self: *File) void {
const request = protocol.Request{ .operation = .close, .node = self.node, .offset = 0, .len = 0, .flags = 0 };
_ = transact(request, &.{}, &.{});
}
};
/// Open (or create, with `.create`) `path`. Returns the open file, or null.
pub fn open(path: []const u8, options: OpenOptions) ?File {
const request = protocol.Request{ .operation = .open, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = options.wireFlags() };
const r = transact(request, path, &.{}) orelse return null;
if (r.reply.status != 0) return null;
return .{ .node = r.reply.node };
}
/// A path's metadata without keeping it open (open -> status -> close).
pub fn attributes(path: []const u8) ?Attributes {
var file = open(path, .{}) orelse return null;
defer file.close();
return file.attributes();
}
/// Whether `path` resolves — handy as a readiness check (e.g. waiting for a mount
/// to come up before writing to it).
pub fn exists(path: []const u8) bool {
return attributes(path) != null;
}
/// One entry returned by `Directory.next`.
pub const Entry = struct {
kind: Kind = .regular,
size: u64 = 0,
name_buffer: [64]u8 = undefined,
name_len: usize = 0,
pub fn name(self: *const Entry) []const u8 {
return self.name_buffer[0..self.name_len];
}
};
/// An open directory being listed, cursor-advanced by `next`.
pub const Directory = struct {
node: u64,
cursor: u64 = 0,
/// Fill `entry` with the next directory entry; false at end of directory or
/// on error.
pub fn next(self: *Directory, entry: *Entry) bool {
const request = protocol.Request{ .operation = .readdir, .node = self.node, .offset = self.cursor, .len = 0, .flags = 0 };
var buffer: [protocol.message_maximum]u8 = undefined;
const r = transact(request, &.{}, &buffer) orelse return false;
if (r.reply.status != 0 or r.reply.len == 0) return false; // error or EOF
if (r.payload.len < protocol.directory_entry_size) return false;
const header = std.mem.bytesToValue(protocol.DirectoryEntry, r.payload[0..protocol.directory_entry_size]);
entry.kind = kindFromWire(header.kind);
entry.size = header.size;
const source = r.payload[protocol.directory_entry_size..];
const nlen = @min(@min(@as(usize, header.name_len), source.len), entry.name_buffer.len);
@memcpy(entry.name_buffer[0..nlen], source[0..nlen]);
entry.name_len = nlen;
self.cursor += 1;
return true;
}
/// Release the VFS's open handle for this directory.
pub fn close(self: *Directory) void {
var f = File{ .node = self.node };
f.close();
}
};
/// Open `path` as a directory for listing. Returns null if it isn't one / on error.
pub fn openDirectory(path: []const u8) ?Directory {
const file = open(path, .{ .directory = true }) orelse return null;
return .{ .node = file.node };
}
// A path-based request that returns only a status (mkdir, unlink).
fn pathOperation(operation: protocol.Operation, path: []const u8) bool {
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(path.len), .flags = 0 };
const r = transact(request, path, &.{}) orelse return false;
return r.reply.status == 0;
}
/// Create a directory at `path` (its parent must already exist). Returns true on
/// success. Only works under a mounted filesystem that supports directories.
pub fn makeDirectory(path: []const u8) bool {
return pathOperation(.mkdir, path);
}
/// Remove the file at `path`. Returns true on success. Directories are refused
/// (a separate directory-removal would have to check emptiness).
pub fn remove(path: []const u8) bool {
return pathOperation(.unlink, path);
}
/// Rename `old_path` to `new_path`. Both must be in the same directory (same-
/// directory, 8.3-name rename only for now). Returns true on success.
pub fn rename(old_path: []const u8, new_path: []const u8) bool {
const total = old_path.len + 1 + new_path.len;
if (total > protocol.maximum_payload) return false;
var payload: [protocol.maximum_payload]u8 = undefined;
@memcpy(payload[0..old_path.len], old_path);
payload[old_path.len] = 0;
@memcpy(payload[old_path.len + 1 ..][0..new_path.len], new_path);
const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 };
const r = transact(request, payload[0..total], &.{}) orelse return false;
return r.reply.status == 0;
}
/// Mount a filesystem backend (its server endpoint) at absolute path `target`;
/// the VFS then routes everything under `target` to that backend. This is the one
/// call that hands the VFS a capability (the backend endpoint). Returns true on
/// success.
pub fn mount(target: []const u8, backend: ipc.Handle) bool {
const h = vfs() orelse return false;
const request = protocol.Request{ .operation = .mount, .node = 0, .offset = 0, .len = @intCast(target.len), .flags = 0 };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const tlen = @min(target.len, protocol.maximum_payload);
@memcpy(message[protocol.request_size..][0..tlen], target[0..tlen]);
var rbuf: [protocol.message_maximum]u8 = undefined;
const result = ipc.callCap(h, message[0 .. protocol.request_size + tlen], &rbuf, backend) catch return false;
if (result.len < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, rbuf[0..protocol.reply_size]).status == 0;
}
+5
View File
@@ -46,6 +46,11 @@ pub const usb = @import("usb.zig");
/// usb-storage). See library/runtime/block.zig. /// usb-storage). See library/runtime/block.zig.
pub const block = @import("block.zig"); pub const block = @import("block.zig");
/// The danos-native file API (open/read/write/list over the user-space VFS) — the
/// layer danos programs use directly, and where the operations that later become
/// `std.os.danos` are staged. See docs/zig-self-hosting.md.
pub const fs = @import("fs.zig");
/// Re-exported so a user binary can `pub const panic = runtime.panic;`. /// Re-exported so a user binary can `pub const panic = runtime.panic;`.
pub const panic = start.panic; pub const panic = start.panic;
+18
View File
@@ -51,6 +51,24 @@ pub fn clock() u64 {
return @intCast(sc.systemCall0(.clock)); return @intCast(sc.systemCall0(.clock));
} }
/// Wall-clock time in Unix epoch seconds (UTC) — the real date/time, from the RTC.
/// Unlike `clock` (monotonic since boot), this tracks calendar time, so it is what a
/// filesystem stamps as a file's modification time. Formatting it into a calendar
/// date/timezone is user-space policy layered on top.
pub fn wallClock() u64 {
return @intCast(sc.systemCall0(.wall_clock));
}
/// Copy bytes out of the kernel's in-memory diagnostic log — the accumulated
/// stream of everything `write` (and the kernel itself) has emitted — starting at
/// `offset`, into `out`. Returns the number of bytes copied (0 at end of buffer).
/// A program reads the whole log by looping from offset 0, advancing by the return
/// value, until it gets 0. This is how the boot log is persisted to disk on a
/// headless/real machine where serial output is otherwise lost.
pub fn klogRead(offset: usize, out: []u8) usize {
return sc.systemCall3(.klog_read, offset, @intFromPtr(out.ptr), out.len);
}
/// End the process. Never returns. /// End the process. Never returns.
pub fn exit(code: usize) noreturn { pub fn exit(code: usize) noreturn {
_ = sc.systemCall1(.exit, code); _ = sc.systemCall1(.exit, code);
+2
View File
@@ -58,6 +58,8 @@ pub const SystemCall = enum(u64) {
signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on
process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself) process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself)
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
klog_read = 32, // klog_read(offset, ptr, len) -> bytes copied: copy the kernel RAM log buffer out to a user buffer (for persisting the boot log to disk)
wall_clock = 33, // wall_clock() -> Unix epoch seconds (UTC): the RTC wall-clock time, for filesystem timestamps (mtime). Monotonic time is `clock`.
_, _,
}; };
+89
View File
@@ -469,6 +469,95 @@ pub fn clockHz() u64 {
return apic.tscHz(); return apic.tscHz();
} }
// --- real-time clock (CMOS) --------------------------------------------------
//
// The battery-backed CMOS clock, read once at boot and thereafter anchored to the
// monotonic clock (see kernel/wall-clock.zig) — so this is never on a hot path and
// needs no lock. Wall-clock *seconds* are mechanism the kernel owns (the hardware's
// value), like the monotonic clock; calendars/timezones are policy layered on top.
fn cmosRead(register: u8) u8 {
io.outb(0x70, register);
return io.inb(0x71);
}
const RtcFields = struct { second: u8, minute: u8, hour: u8, day: u8, month: u8, year: u8 };
fn rtcRaw() RtcFields {
while (cmosRead(0x0A) & 0x80 != 0) {} // wait out any update in progress (status A bit 7)
return .{
.second = cmosRead(0x00),
.minute = cmosRead(0x02),
.hour = cmosRead(0x04),
.day = cmosRead(0x07),
.month = cmosRead(0x08),
.year = cmosRead(0x09),
};
}
fn bcdToBinary(v: u8) u8 {
return (v & 0x0F) + ((v >> 4) * 10);
}
fn isLeapYear(y: u32) bool {
return (y % 4 == 0 and y % 100 != 0) or (y % 400 == 0);
}
/// Read the CMOS real-time clock and convert it to Unix epoch seconds (UTC).
pub fn readRtcUnixSeconds() u64 {
// Read until two consecutive reads agree, so we never latch a half-updated time.
var a = rtcRaw();
while (true) {
const b = rtcRaw();
if (a.second == b.second and a.minute == b.minute and a.hour == b.hour and
a.day == b.day and a.month == b.month and a.year == b.year) break;
a = b;
}
const status_b = cmosRead(0x0B);
const binary_mode = status_b & 0x04 != 0; // else BCD
const hour_24 = status_b & 0x02 != 0; // else 12-hour with a PM bit
var second = a.second;
var minute = a.minute;
var hour_field = a.hour;
var day = a.day;
var month = a.month;
var year = a.year;
if (!binary_mode) {
second = bcdToBinary(second);
minute = bcdToBinary(minute);
hour_field = bcdToBinary(hour_field & 0x7F) | (hour_field & 0x80); // preserve the PM bit
day = bcdToBinary(day);
month = bcdToBinary(month);
year = bcdToBinary(year);
}
var hour: u32 = hour_field & 0x7F;
if (!hour_24) {
const pm = hour_field & 0x80 != 0;
hour %= 12; // 12 AM/PM -> 0
if (pm) hour += 12;
}
// The CMOS year is 0..99; QEMU and modern hardware mean 20xx (there is no
// reliable century register on QEMU). Treat < 70 as 20xx, else 19xx.
const full_year: u32 = if (year < 70) 2000 + @as(u32, year) else 1900 + @as(u32, year);
var days: u64 = 0;
var y: u32 = 1970;
while (y < full_year) : (y += 1) days += if (isLeapYear(y)) 366 else 365;
const month_lengths = [_]u8{ 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31 };
var m: u8 = 1;
while (m < month) : (m += 1) {
days += month_lengths[m - 1];
if (m == 2 and isLeapYear(full_year)) days += 1;
}
days += @as(u64, day) - 1;
return ((days * 24 + hour) * 60 + minute) * 60 + second;
}
/// Whether the CPU guarantees an **invariant** TSC (CPUID 0x80000007 EDX[8] on /// Whether the CPU guarantees an **invariant** TSC (CPUID 0x80000007 EDX[8] on
/// x86; the analogous architectural guarantee elsewhere). When false the TSC is not /// x86; the analogous architectural guarantee elsewhere). When false the TSC is not
/// used as the clocksource. /// used as the clocksource.
+9
View File
@@ -5,6 +5,7 @@ const parameters = @import("parameters");
const architecture = @import("architecture"); const architecture = @import("architecture");
const console = @import("console.zig"); const console = @import("console.zig");
const log = @import("log.zig"); const log = @import("log.zig");
const wall_clock = @import("wall-clock.zig");
const pmm = @import("pmm.zig"); const pmm = @import("pmm.zig");
const heap = @import("heap.zig"); const heap = @import("heap.zig");
const scheduler = @import("scheduler.zig"); const scheduler = @import("scheduler.zig");
@@ -62,6 +63,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
architecture.serialInit(); architecture.serialInit();
log.addSink(architecture.serialWrite); log.addSink(architecture.serialWrite);
if (architecture.debugconPresent()) log.addSink(architecture.debugconWrite); if (architecture.debugconPresent()) log.addSink(architecture.debugconWrite);
// Retain the whole stream in a RAM buffer too, so a user program can later
// read it back (klog_read) and persist the boot log to disk — the only way to
// see it on a headless/real machine with no host capturing serial.
log.addSink(log.ramSink);
// The **framebuffer** is deliberately *not* a log sink. It's a separate output // The **framebuffer** is deliberately *not* a log sink. It's a separate output
// surface — a bootstrap text console today, a graphics device driver later — so // surface — a bootstrap text console today, a graphics device driver later — so
@@ -281,6 +286,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
if (!architecture.clockSynchronized()) if (!architecture.clockSynchronized())
log.write("/system/kernel: WARNING: per-core TSCs are not synchronized; monotonic clock moved off the TSC\n"); log.write("/system/kernel: WARNING: per-core TSCs are not synchronized; monotonic clock moved off the TSC\n");
// Anchor wall-clock time: read the RTC once, now the monotonic clock is final.
wall_clock.init();
log.print("/system/kernel: wall clock {d} (Unix epoch seconds, UTC, from the RTC)\n", .{wall_clock.nowSeconds()});
// In a test build (`zig build -Dtest-case=<name>`), run that case and stop. // In a test build (`zig build -Dtest-case=<name>`), run that case and stop.
// Normal builds fall through to the idle halt. // Normal builds fall through to the idle halt.
if (build_options.test_case) |case| { if (build_options.test_case) |case| {
+33
View File
@@ -41,6 +41,39 @@ pub fn write(bytes: []const u8) void {
for (sinks[0..sink_count]) |sink| sink(bytes); for (sinks[0..sink_count]) |sink| sink(bytes);
} }
// --- the RAM sink: a retained copy of the whole diagnostic stream ------------
//
// A fixed in-image buffer that accumulates every logged byte, so a user program
// (`log-flush`, and init at shutdown) can read it back through `klog_read` and
// persist it to a file — the boot log survives on a headless/real machine that
// has no host capturing serial. It is a *sink like any other*: register it with
// `addSink(ramSink)` at boot. No allocation (works pre-heap and in a panic).
//
// It fills linearly and stops when full: the earliest output — the most valuable
// for diagnosing a boot — is kept, and the tail is still on the live serial sink.
// 256 KiB comfortably holds a full boot plus a long run (a boot is ~15 KiB).
const ram_capacity = 256 * 1024;
var ram_buffer: [ram_capacity]u8 = undefined;
var ram_len: usize = 0;
/// The RAM sink. Best-effort and self-guarding like every sink: appends what fits
/// and silently drops the rest once full. (Concurrency matches the other sinks —
/// the dominant writer, debug_write, already holds the kernel lock; a rare torn
/// append on a kernel-internal line is an accepted diagnostic imperfection.)
pub fn ramSink(bytes: []const u8) void {
const n = @min(ram_buffer.len - ram_len, bytes.len);
if (n != 0) {
@memcpy(ram_buffer[ram_len..][0..n], bytes[0..n]);
ram_len += n;
}
}
/// The accumulated log so far — what `klog_read` copies out.
pub fn ramSnapshot() []const u8 {
return ram_buffer[0..ram_len];
}
/// A formatted log line. Truncates past 256 bytes; the buffer is on the stack, so /// A formatted log line. Truncates past 256 bytes; the buffer is on the stack, so
/// this is safe to call from interrupt context and from a panic. /// this is safe to call from interrupt context and from a panic.
pub fn print(comptime fmt: []const u8, args: anytype) void { pub fn print(comptime fmt: []const u8, args: anytype) void {
+41
View File
@@ -34,6 +34,7 @@ const devices_broker = @import("devices-broker.zig");
const irq = @import("irq.zig"); const irq = @import("irq.zig");
const initial_ramdisk = @import("initial-ramdisk"); const initial_ramdisk = @import("initial-ramdisk");
const log = @import("log.zig"); const log = @import("log.zig");
const wall_clock = @import("wall-clock.zig");
const page_size = abi.page_size; const page_size = abi.page_size;
const SystemCall = abi.SystemCall; const SystemCall = abi.SystemCall;
@@ -206,6 +207,8 @@ fn system_call(state: *architecture.CpuState) void {
.signal_bind => systemSignalBind(state), .signal_bind => systemSignalBind(state),
.process_signal => systemProcessSignal(state), .process_signal => systemProcessSignal(state),
.timer_bind => systemTimerBind(state), .timer_bind => systemTimerBind(state),
.klog_read => systemKlogRead(state),
.wall_clock => systemWallClock(state),
_ => fail(state), _ => fail(state),
} }
} }
@@ -974,6 +977,37 @@ fn systemDebugWrite(state: *architecture.CpuState) void {
} }
} }
/// klog_read(offset, ptr, len) -> bytes copied: copy the kernel's in-memory
/// diagnostic log (the RAM sink in log.zig) out to the user buffer at `ptr`,
/// starting at `offset`. Returns the count copied — 0 once `offset` reaches the
/// end — so a program reads the whole log by looping from 0 until it gets 0.
///
/// The mirror of `debug_write`: the same overflow-safe user-half bounds check,
/// but the copy runs kernel -> user. Written under the kernel lock so the source
/// snapshot can't grow underneath the copy. A read-only diagnostic — it exposes
/// only the log the kernel already broadcasts to serial, nothing else.
fn systemKlogRead(state: *architecture.CpuState) void {
const offset = architecture.systemCallArg(state, 0);
const ptr = architecture.systemCallArg(state, 1);
const len = architecture.systemCallArg(state, 2);
// Confine the whole destination span to the user (low) half. `len <=
// user_half_end - ptr` bounds the length without an overflowing add.
if (ptr < user_half_end and len <= user_half_end - ptr) {
const flags = sync.enter();
defer sync.leave(flags);
const snapshot = log.ramSnapshot();
var n: usize = 0;
if (offset < snapshot.len) {
n = @min(len, snapshot.len - offset);
const dest: [*]u8 = @ptrFromInt(ptr);
@memcpy(dest[0..n], snapshot[offset..][0..n]);
}
architecture.setSystemCallResult(state, n);
} else {
fail(state);
}
}
/// mmap(len, prot) -> base: grant `len` bytes (rounded up to whole pages) of /// mmap(len, prot) -> base: grant `len` bytes (rounded up to whole pages) of
/// fresh, zeroed, writable+NX memory in the caller's mmap arena, and return the /// fresh, zeroed, writable+NX memory in the caller's mmap arena, and return the
/// base virtual address. `prot` is accepted but not yet honoured (grants are /// base virtual address. `prot` is accepted but not yet honoured (grants are
@@ -1304,3 +1338,10 @@ pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []c
fn systemClock(state: *architecture.CpuState) void { fn systemClock(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, architecture.nanos()); architecture.setSystemCallResult(state, architecture.nanos());
} }
/// wall_clock() -> Unix epoch seconds (UTC). The RTC value, read at boot and offset
/// by the monotonic clock (wall-clock.zig) — mechanism, not policy: calendars and
/// timezones layer on top in user space. Needed for filesystem timestamps (mtime).
fn systemWallClock(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, wall_clock.nowSeconds());
}
+14
View File
@@ -14,6 +14,7 @@ const boot_handoff = @import("boot-handoff");
const abi = @import("abi"); const abi = @import("abi");
const device_abi = @import("device-abi"); const device_abi = @import("device-abi");
const architecture = @import("architecture"); const architecture = @import("architecture");
const wall_clock = @import("wall-clock.zig");
const devices_broker = @import("devices-broker.zig"); const devices_broker = @import("devices-broker.zig");
const platform = @import("platform"); const platform = @import("platform");
const pmm = @import("pmm.zig"); const pmm = @import("pmm.zig");
@@ -66,6 +67,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
timer(); timer();
} else if (eql(case, "clock")) { } else if (eql(case, "clock")) {
clock(); clock();
} else if (eql(case, "wall-clock")) {
wallClock();
} else if (eql(case, "vmm")) { } else if (eql(case, "vmm")) {
vmm(); vmm();
} else if (eql(case, "heap")) { } else if (eql(case, "heap")) {
@@ -458,6 +461,17 @@ fn heapTest() void {
/// Verify the calibrated clocks: sane measured frequencies, monotonic uptime that /// Verify the calibrated clocks: sane measured frequencies, monotonic uptime that
/// advances with real ticks, and — the point of the TSC clock — nanosecond /// advances with real ticks, and — the point of the TSC clock — nanosecond
/// resolution far finer than the 1 ms tick, with the unit functions consistent. /// resolution far finer than the 1 ms tick, with the unit functions consistent.
fn wallClock() void {
log("DANOS-TEST-BEGIN: wall-clock\n", .{});
// The RTC was read and anchored at boot (kmain -> wall_clock.init()).
const seconds = wall_clock.nowSeconds();
log(" epoch: {d}\n", .{seconds});
// A plausible current wall-clock: after 2020-01-01 (1577836800) and before 2050
// (2524608000) — catches a broken CMOS read or a wrong epoch conversion.
check("wall clock reads a plausible current epoch", seconds > 1_577_836_800 and seconds < 2_524_608_000);
result();
}
fn clock() void { fn clock() void {
log("DANOS-TEST-BEGIN: clock\n", .{}); log("DANOS-TEST-BEGIN: clock\n", .{});
+26
View File
@@ -0,0 +1,26 @@
//! Wall-clock time: the CMOS real-time clock read once at boot and anchored to the
//! monotonic clock, so a query is a cheap arithmetic offset — no per-call CMOS poll,
//! no lock, no SMP hazard on the shared 0x70/0x71 ports.
//!
//! Wall-clock *seconds* are mechanism the kernel owns, exactly like the monotonic
//! clock ([[time-architecture]]): reading the hardware's value is not policy.
//! Calendars, timezones, and formatting layer on top in user space. It exists so the
//! filesystem can stamp real timestamps (mtime) — see docs/zig-self-hosting.md.
const architecture = @import("architecture");
var boot_unix_seconds: u64 = 0;
var boot_nanos: u64 = 0;
/// Read the RTC once and anchor it to the monotonic clock. Call at boot, after the
/// monotonic clock is calibrated.
pub fn init() void {
boot_unix_seconds = architecture.readRtcUnixSeconds();
boot_nanos = architecture.nanos();
}
/// The current wall-clock time in Unix epoch seconds (UTC): the boot RTC value plus
/// the monotonic time elapsed since. Zero until `init` runs.
pub fn nowSeconds() u64 {
return boot_unix_seconds + (architecture.nanos() -% boot_nanos) / 1_000_000_000;
}
+410 -12
View File
@@ -1,6 +1,8 @@
//! The FAT filesystem engine: mount a block device, walk the FAT and directory //! The FAT filesystem engine: mount a block device, walk the FAT and directory
//! structures, and read / write / create files. FAT12/16/32 (the type is //! structures, and read / write / create / truncate files, remove files, and make
//! detected from the cluster count). Pure logic over a `BlockDevice` interface — //! directories. FAT12/16/32 (the type is detected from the cluster count). The
//! crash-safe write order is data -> FAT -> directory; truncate and remove free the
//! cluster chain, then update the directory. Pure logic over a `BlockDevice` interface —
//! no IPC — so it is host-testable against a RAM-backed image (see the tests at //! no IPC — so it is host-testable against a RAM-backed image (see the tests at
//! the bottom). The fat.zig server wraps a real `.block` device in a BlockDevice //! the bottom). The fat.zig server wraps a real `.block` device in a BlockDevice
//! and serves this over the VFS protocol. //! and serves this over the VFS protocol.
@@ -39,6 +41,9 @@ pub const Node = struct {
first_cluster: u32, first_cluster: u32,
size: u32, size: u32,
is_directory: bool, is_directory: bool,
// Modification time (Unix epoch seconds, UTC), decoded from the directory
// entry's DOS write date/time. 0 if unset.
mtime: u64 = 0,
// The absolute sector and byte offset of this node's 8.3 directory entry, so // The absolute sector and byte offset of this node's 8.3 directory entry, so
// size/first-cluster changes can be written back. Absent for the root. // size/first-cluster changes can be written back. Absent for the root.
entry_sector: u64 = 0, entry_sector: u64 = 0,
@@ -61,6 +66,10 @@ pub const FileSystem = struct {
sector: [sector_size]u8 = undefined, sector: [sector_size]u8 = undefined,
fat_sector: [sector_size]u8 = undefined, fat_sector: [sector_size]u8 = undefined,
dir_sector: [sector_size]u8 = undefined, dir_sector: [sector_size]u8 = undefined,
// Wall-clock time (Unix epoch seconds) to stamp on create/write, set by the
// server before a mutating op. 0 leaves the on-disk timestamps untouched (host
// tests that don't care about time, and reads).
current_time_epoch: u64 = 0,
// Every filesystem-relative sector access adds the partition base. // Every filesystem-relative sector access adds the partition base.
fn blockRead(self: *FileSystem, lba: u64, buffer: []u8) bool { fn blockRead(self: *FileSystem, lba: u64, buffer: []u8) bool {
@@ -238,6 +247,21 @@ pub const FileSystem = struct {
return null; return null;
} }
// Free every cluster of the chain starting at `first`, returning them to the
// pool. A first < 2 (an empty file) frees nothing. Bounded against a corrupt
// cyclic chain by the cluster count so it can never loop forever.
fn freeChain(self: *FileSystem, first: u32) void {
var cluster = first;
var guard: u32 = 0;
const limit = self.geometry.cluster_count + 2;
while (cluster >= 2 and cluster < limit and guard < limit) : (guard += 1) {
const next = self.readFatEntry(cluster);
_ = self.writeFatEntry(cluster, on_disk.free_cluster);
if (self.isEndOfChain(next) or next < 2) break;
cluster = next;
}
}
// --- directory iteration ------------------------------------------------ // --- directory iteration ------------------------------------------------
// The absolute LBA of the `sector_index`th sector of directory `dir`, or null // The absolute LBA of the `sector_index`th sector of directory `dir`, or null
@@ -411,6 +435,7 @@ pub const FileSystem = struct {
.first_cluster = entry.firstCluster(), .first_cluster = entry.firstCluster(),
.size = entry.file_size, .size = entry.file_size,
.is_directory = entry.isDirectory(), .is_directory = entry.isDirectory(),
.mtime = on_disk.fatToEpoch(entry.write_date, entry.write_time),
.entry_sector = entry_sector, .entry_sector = entry_sector,
.entry_offset = entry_offset, .entry_offset = entry_offset,
.has_entry = true, .has_entry = true,
@@ -439,7 +464,7 @@ pub const FileSystem = struct {
/// The `cursor`th real entry of a directory (for readdir): its display name, /// The `cursor`th real entry of a directory (for readdir): its display name,
/// kind, and size. Returns null past the end. /// kind, and size. Returns null past the end.
pub const Listing = struct { name_buffer: [260]u8 = undefined, name_len: usize = 0, is_directory: bool = false, size: u32 = 0 }; pub const Listing = struct { name_buffer: [260]u8 = undefined, name_len: usize = 0, is_directory: bool = false, size: u32 = 0, mtime: u64 = 0 };
const ListContext = struct { target: u32, index: u32 = 0, out: *Listing, done: bool = false }; const ListContext = struct { target: u32, index: u32 = 0, out: *Listing, done: bool = false };
fn listVisit(context: *ListContext, entry: on_disk.DirectoryEntry, name: []const u8, entry_sector: u64, entry_offset: u32) bool { fn listVisit(context: *ListContext, entry: on_disk.DirectoryEntry, name: []const u8, entry_sector: u64, entry_offset: u32) bool {
@@ -451,6 +476,7 @@ pub const FileSystem = struct {
context.out.name_len = n; context.out.name_len = n;
context.out.is_directory = entry.isDirectory(); context.out.is_directory = entry.isDirectory();
context.out.size = entry.file_size; context.out.size = entry.file_size;
context.out.mtime = on_disk.fatToEpoch(entry.write_date, entry.write_time);
context.done = true; context.done = true;
return true; return true;
} }
@@ -546,6 +572,18 @@ pub const FileSystem = struct {
return consumed; return consumed;
} }
/// Truncate a file node to zero length: free its cluster chain and clear its
/// size and first cluster in the directory entry. This is O_TRUNC — the fix for
/// re-opening and overwriting an existing file, whose old (longer) contents
/// would otherwise linger past the new end (a silent-corruption bug for anything
/// that rewrites a file in place, like the boot-log flush).
pub fn truncate(self: *FileSystem, node: *Node) void {
self.freeChain(node.first_cluster);
node.first_cluster = 0;
node.size = 0;
self.updateEntry(node.*);
}
// Write a node's size and first cluster back into its 8.3 directory entry. // Write a node's size and first cluster back into its 8.3 directory entry.
fn updateEntry(self: *FileSystem, node: Node) void { fn updateEntry(self: *FileSystem, node: Node) void {
if (!node.has_entry) return; if (!node.has_entry) return;
@@ -553,15 +591,22 @@ pub const FileSystem = struct {
var entry = std.mem.bytesToValue(on_disk.DirectoryEntry, self.dir_sector[node.entry_offset .. node.entry_offset + @sizeOf(on_disk.DirectoryEntry)]); var entry = std.mem.bytesToValue(on_disk.DirectoryEntry, self.dir_sector[node.entry_offset .. node.entry_offset + @sizeOf(on_disk.DirectoryEntry)]);
entry.file_size = node.size; entry.file_size = node.size;
entry.setFirstCluster(node.first_cluster); entry.setFirstCluster(node.first_cluster);
// A write updates the modification time (leave it if no time is set, so host
// tests and reads don't zero it).
if (self.current_time_epoch != 0) {
const stamp = on_disk.epochToFatDateTime(self.current_time_epoch);
entry.write_date = stamp.date;
entry.write_time = stamp.time;
entry.last_access_date = stamp.date;
}
@memcpy(self.dir_sector[node.entry_offset .. node.entry_offset + @sizeOf(on_disk.DirectoryEntry)], std.mem.asBytes(&entry)); @memcpy(self.dir_sector[node.entry_offset .. node.entry_offset + @sizeOf(on_disk.DirectoryEntry)], std.mem.asBytes(&entry));
_ = self.blockWrite(node.entry_sector, &self.dir_sector); _ = self.blockWrite(node.entry_sector, &self.dir_sector);
} }
/// Create an 8.3-named file in directory `dir`. Returns the new (empty) node, // Add an 8.3 directory entry to `dir` with the given attributes, first cluster,
/// or null if the name is not 8.3-representable or no directory slot is free. // and size, reusing a free (0x00 or 0xE5) slot and growing the directory chain
pub fn createFile(self: *FileSystem, dir: Node, name: []const u8) ?Node { // if needed. Returns the new node (with its entry location) or null if full.
const raw = to83(name) orelse return null; fn addEntry(self: *FileSystem, dir: Node, raw: [11]u8, attributes: u8, first_cluster: u32, size: u32) ?Node {
// Find a free directory slot (a 0x00 or 0xE5 entry), growing the directory.
var sector_index: u32 = 0; var sector_index: u32 = 0;
while (self.dirSectorLba(dir, sector_index, true)) |lba| : (sector_index += 1) { while (self.dirSectorLba(dir, sector_index, true)) |lba| : (sector_index += 1) {
if (!self.blockRead(lba, &self.dir_sector)) return null; if (!self.blockRead(lba, &self.dir_sector)) return null;
@@ -572,13 +617,22 @@ pub const FileSystem = struct {
if (existing.isFree()) { if (existing.isFree()) {
var entry = std.mem.zeroes(on_disk.DirectoryEntry); var entry = std.mem.zeroes(on_disk.DirectoryEntry);
entry.name = raw; entry.name = raw;
entry.attributes = on_disk.attribute_archive; entry.attributes = attributes;
entry.file_size = size;
entry.setFirstCluster(first_cluster);
const stamp = on_disk.epochToFatDateTime(self.current_time_epoch);
entry.creation_date = stamp.date;
entry.creation_time = stamp.time;
entry.write_date = stamp.date;
entry.write_time = stamp.time;
entry.last_access_date = stamp.date;
@memcpy(self.dir_sector[offset .. offset + @sizeOf(on_disk.DirectoryEntry)], std.mem.asBytes(&entry)); @memcpy(self.dir_sector[offset .. offset + @sizeOf(on_disk.DirectoryEntry)], std.mem.asBytes(&entry));
if (!self.blockWrite(lba, &self.dir_sector)) return null; if (!self.blockWrite(lba, &self.dir_sector)) return null;
return .{ return .{
.first_cluster = 0, .first_cluster = first_cluster,
.size = 0, .size = size,
.is_directory = false, .is_directory = attributes & on_disk.attribute_directory != 0,
.mtime = self.current_time_epoch,
.entry_sector = lba, .entry_sector = lba,
.entry_offset = offset, .entry_offset = offset,
.has_entry = true, .has_entry = true,
@@ -590,6 +644,194 @@ pub const FileSystem = struct {
} }
return null; return null;
} }
/// Create an 8.3-named file in directory `dir`. Returns the new (empty) node,
/// or null if the name is not 8.3-representable or no directory slot is free.
pub fn createFile(self: *FileSystem, dir: Node, name: []const u8) ?Node {
const raw = to83(name) orelse return null;
return self.addEntry(dir, raw, on_disk.attribute_archive, 0, 0);
}
/// Create an 8.3-named subdirectory in `dir`: allocate and initialise its first
/// cluster with "." (itself) and ".." (the parent) entries, then add its
/// directory entry to `dir`. Returns the new directory node, or null (bad name,
/// no free cluster, or the directory is full). Long names are not created.
pub fn createDirectory(self: *FileSystem, dir: Node, name: []const u8) ?Node {
const raw = to83(name) orelse return null;
const cluster = self.allocateCluster() orelse return null;
self.zeroCluster(cluster);
// ".." points at the parent: 0 for the fixed root on FAT12/16, the root
// cluster on FAT32, else the parent's own first cluster.
const parent_cluster: u32 = if (dir.first_cluster != 0)
dir.first_cluster
else if (self.geometry.fat_type == .fat32)
self.geometry.root_cluster
else
0;
var dot = std.mem.zeroes(on_disk.DirectoryEntry);
dot.name = [_]u8{'.'} ++ ([_]u8{' '} ** 10);
dot.attributes = on_disk.attribute_directory;
dot.setFirstCluster(cluster);
var dotdot = std.mem.zeroes(on_disk.DirectoryEntry);
dotdot.name = [_]u8{ '.', '.' } ++ ([_]u8{' '} ** 9);
dotdot.attributes = on_disk.attribute_directory;
dotdot.setFirstCluster(parent_cluster);
var first_sector = [_]u8{0} ** sector_size;
const entry_size = @sizeOf(on_disk.DirectoryEntry);
@memcpy(first_sector[0..entry_size], std.mem.asBytes(&dot));
@memcpy(first_sector[entry_size .. 2 * entry_size], std.mem.asBytes(&dotdot));
if (!self.blockWrite(self.clusterSector(cluster, 0), &first_sector)) {
self.freeChain(cluster);
return null;
}
return self.addEntry(dir, raw, on_disk.attribute_directory, cluster, 0) orelse {
self.freeChain(cluster);
return null;
};
}
const EntryLoc = struct { sector: u64, offset: u32 };
// Mark a directory entry deleted in place (its name[0] set to 0xE5).
fn markDeleted(self: *FileSystem, sector: u64, offset: u32) void {
if (!self.blockRead(sector, &self.dir_sector)) return;
self.dir_sector[offset] = 0xE5;
_ = self.blockWrite(sector, &self.dir_sector);
}
/// Remove a file named `name` from directory `dir`: free its cluster chain and
/// mark its 8.3 entry — and any long-name entries immediately preceding it —
/// deleted, so the slots (and the long name) are reusable without a later entry
/// that reuses them inheriting the orphaned long name. Refuses a directory (a
/// separate rmdir would have to check emptiness). Returns true if removed.
pub fn removeFile(self: *FileSystem, dir: Node, name: []const u8) bool {
var run: [21]EntryLoc = undefined; // the long-name entries before the 8.3 one
var run_len: usize = 0;
var long_name: [260]u8 = undefined;
var long_len: usize = 0;
var sector_index: u32 = 0;
while (self.dirSectorLba(dir, sector_index, false)) |lba| : (sector_index += 1) {
if (!self.blockRead(lba, &self.dir_sector)) return false;
var i: u32 = 0;
while (i < entries_per_sector) : (i += 1) {
const offset = i * @sizeOf(on_disk.DirectoryEntry);
const entry = std.mem.bytesToValue(on_disk.DirectoryEntry, self.dir_sector[offset .. offset + @sizeOf(on_disk.DirectoryEntry)]);
if (entry.isEnd()) return false;
if (entry.name[0] == 0xE5) {
run_len = 0;
long_len = 0;
continue;
}
if (entry.isLongName()) {
if (run_len < run.len) {
run[run_len] = .{ .sector = lba, .offset = offset };
run_len += 1;
}
const lfn = std.mem.bytesToValue(on_disk.LongNameEntry, self.dir_sector[offset .. offset + @sizeOf(on_disk.LongNameEntry)]);
const order = lfn.order & 0x1F;
if (order >= 1 and order <= 20) {
var chunk: [13]u8 = undefined;
const n = longNameChars(lfn, &chunk);
const start = (order - 1) * 13;
if (start + n <= long_name.len) {
@memcpy(long_name[start .. start + n], chunk[0..n]);
if (lfn.order & 0x40 != 0) long_len = start + n;
}
}
continue;
}
if (entry.isVolumeLabel()) {
run_len = 0;
long_len = 0;
continue;
}
// A real 8.3 entry.
var short: [12]u8 = undefined;
const display = if (long_len > 0) long_name[0..long_len] else format83(entry.name, &short);
if (nameMatches(display, name)) {
if (entry.isDirectory()) return false; // not for directories
self.freeChain(entry.firstCluster());
self.markDeleted(lba, offset);
var r: usize = 0;
while (r < run_len) : (r += 1) self.markDeleted(run[r].sector, run[r].offset);
return true;
}
run_len = 0;
long_len = 0;
}
}
return false;
}
/// Rename `old_name` to `new_name` within the SAME directory `dir`, rewriting
/// the 8.3 entry's name in place. Refuses if `old_name` is missing, `new_name`
/// is not 8.3-representable, or `new_name` already exists. Any long-name entries
/// on the old file are dropped (the file takes its new 8.3 name); cross-directory
/// and long-name-preserving rename are not implemented. Returns true on success.
pub fn rename(self: *FileSystem, dir: Node, old_name: []const u8, new_name: []const u8) bool {
const raw = to83(new_name) orelse return false;
if (self.findChild(dir, new_name) != null) return false; // target already exists
var run: [21]EntryLoc = undefined; // the long-name entries before the 8.3 one
var run_len: usize = 0;
var long_name: [260]u8 = undefined;
var long_len: usize = 0;
var sector_index: u32 = 0;
while (self.dirSectorLba(dir, sector_index, false)) |lba| : (sector_index += 1) {
if (!self.blockRead(lba, &self.dir_sector)) return false;
var i: u32 = 0;
while (i < entries_per_sector) : (i += 1) {
const offset = i * @sizeOf(on_disk.DirectoryEntry);
const entry = std.mem.bytesToValue(on_disk.DirectoryEntry, self.dir_sector[offset .. offset + @sizeOf(on_disk.DirectoryEntry)]);
if (entry.isEnd()) return false;
if (entry.name[0] == 0xE5) {
run_len = 0;
long_len = 0;
continue;
}
if (entry.isLongName()) {
if (run_len < run.len) {
run[run_len] = .{ .sector = lba, .offset = offset };
run_len += 1;
}
const lfn = std.mem.bytesToValue(on_disk.LongNameEntry, self.dir_sector[offset .. offset + @sizeOf(on_disk.LongNameEntry)]);
const order = lfn.order & 0x1F;
if (order >= 1 and order <= 20) {
var chunk: [13]u8 = undefined;
const n = longNameChars(lfn, &chunk);
const start = (order - 1) * 13;
if (start + n <= long_name.len) {
@memcpy(long_name[start .. start + n], chunk[0..n]);
if (lfn.order & 0x40 != 0) long_len = start + n;
}
}
continue;
}
if (entry.isVolumeLabel()) {
run_len = 0;
long_len = 0;
continue;
}
// A real 8.3 entry.
var short: [12]u8 = undefined;
const display = if (long_len > 0) long_name[0..long_len] else format83(entry.name, &short);
if (nameMatches(display, old_name)) {
var updated = entry;
updated.name = raw;
@memcpy(self.dir_sector[offset .. offset + @sizeOf(on_disk.DirectoryEntry)], std.mem.asBytes(&updated));
if (!self.blockWrite(lba, &self.dir_sector)) return false;
// Drop the old long name, if any, so the new 8.3 name is what shows.
var r: usize = 0;
while (r < run_len) : (r += 1) self.markDeleted(run[r].sector, run[r].offset);
return true;
}
run_len = 0;
long_len = 0;
}
}
return false;
}
}; };
// --- tests: a RAM-backed FAT16 image ---------------------------------------- // --- tests: a RAM-backed FAT16 image ----------------------------------------
@@ -705,3 +947,159 @@ test "create, write, read back a file through the engine" {
try std.testing.expectEqualStrings("HELLO.TXT", listing.name_buffer[0..listing.name_len]); try std.testing.expectEqualStrings("HELLO.TXT", listing.name_buffer[0..listing.name_len]);
try std.testing.expect(fs.listEntry(fs.rootNode(), 1) == null); try std.testing.expect(fs.listEntry(fs.rootNode(), 1) == null);
} }
test "truncate frees the chain and zeroes the file" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
var node = fs.createFile(fs.rootNode(), "BIG.BIN").?;
var payload: [2000]u8 = undefined;
for (&payload, 0..) |*b, i| b.* = @truncate(i);
_ = fs.writeFile(&node, 0, &payload);
const cluster = node.first_cluster;
try std.testing.expect(cluster >= 2);
fs.truncate(&node);
try std.testing.expectEqual(@as(u32, 0), node.size);
try std.testing.expectEqual(@as(u32, 0), node.first_cluster);
// The old first cluster is free again.
try std.testing.expectEqual(on_disk.free_cluster, fs.readFatEntry(cluster));
// Re-resolve: the persisted entry is empty, and reads produce nothing.
const resolved = fs.resolve("/BIG.BIN").?;
try std.testing.expectEqual(@as(u32, 0), resolved.size);
var buf: [16]u8 = undefined;
try std.testing.expectEqual(@as(usize, 0), fs.readFile(resolved, 0, &buf));
}
test "overwrite after truncate leaves no stale tail (the O_TRUNC corruption fix)" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
// Write a long file, then truncate-and-rewrite a short one — the O_TRUNC flow.
var node = fs.createFile(fs.rootNode(), "LOG.TXT").?;
_ = fs.writeFile(&node, 0, "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"); // 32 bytes
fs.truncate(&node);
_ = fs.writeFile(&node, 0, "bb");
// Size is the short length — no lingering old bytes past the new end.
const resolved = fs.resolve("/LOG.TXT").?;
try std.testing.expectEqual(@as(u32, 2), resolved.size);
var buf: [8]u8 = undefined;
const n = fs.readFile(resolved, 0, &buf);
try std.testing.expectEqualStrings("bb", buf[0..n]);
}
test "remove a file frees its slot and its cluster chain" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
var node = fs.createFile(fs.rootNode(), "GONE.TXT").?;
var payload: [1000]u8 = undefined;
for (&payload, 0..) |*b, i| b.* = @truncate(i);
_ = fs.writeFile(&node, 0, &payload);
const cluster = fs.resolve("/GONE.TXT").?.first_cluster;
try std.testing.expect(cluster >= 2);
try std.testing.expect(fs.removeFile(fs.rootNode(), "GONE.TXT"));
// Gone from the directory, its cluster free, root empty again.
try std.testing.expect(fs.resolve("/GONE.TXT") == null);
try std.testing.expectEqual(on_disk.free_cluster, fs.readFatEntry(cluster));
try std.testing.expect(fs.listEntry(fs.rootNode(), 0) == null);
// Removing a missing file reports false.
try std.testing.expect(!fs.removeFile(fs.rootNode(), "GONE.TXT"));
// A directory is refused (it is not a file).
_ = fs.createDirectory(fs.rootNode(), "ADIR").?;
try std.testing.expect(!fs.removeFile(fs.rootNode(), "ADIR"));
}
test "create a subdirectory with . and .. and a file inside" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
const made = fs.createDirectory(fs.rootNode(), "SUB").?;
try std.testing.expect(made.is_directory);
try std.testing.expect(made.first_cluster >= 2);
// It resolves as a directory, with "." and ".." as its first two entries.
const dir = fs.resolve("/SUB").?;
try std.testing.expect(dir.is_directory);
const dot = fs.listEntry(dir, 0).?;
try std.testing.expectEqualStrings(".", dot.name_buffer[0..dot.name_len]);
const dotdot = fs.listEntry(dir, 1).?;
try std.testing.expectEqualStrings("..", dotdot.name_buffer[0..dotdot.name_len]);
// A file created inside is reachable by its full path.
var child = fs.createFile(dir, "INNER.TXT").?;
_ = fs.writeFile(&child, 0, "hi");
const inner = fs.resolve("/SUB/INNER.TXT").?;
try std.testing.expectEqual(@as(u32, 2), inner.size);
// The root lists SUB as a directory.
const listing = fs.listEntry(fs.rootNode(), 0).?;
try std.testing.expectEqualStrings("SUB", listing.name_buffer[0..listing.name_len]);
try std.testing.expect(listing.is_directory);
}
test "rename a file in place, keeping its contents" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
var node = fs.createFile(fs.rootNode(), "OLD.TXT").?;
_ = fs.writeFile(&node, 0, "content");
try std.testing.expect(fs.rename(fs.rootNode(), "OLD.TXT", "NEW.TXT"));
try std.testing.expect(fs.resolve("/OLD.TXT") == null);
const renamed = fs.resolve("/NEW.TXT").?;
var buf: [16]u8 = undefined;
const n = fs.readFile(renamed, 0, &buf);
try std.testing.expectEqualStrings("content", buf[0..n]);
// Refuse a collision with an existing name.
_ = fs.createFile(fs.rootNode(), "OTHER.TXT").?;
try std.testing.expect(!fs.rename(fs.rootNode(), "NEW.TXT", "OTHER.TXT"));
// Refuse a non-8.3 target name.
try std.testing.expect(!fs.rename(fs.rootNode(), "NEW.TXT", "toolongbasename.txt"));
// Refuse a missing source.
try std.testing.expect(!fs.rename(fs.rootNode(), "NOPE.TXT", "X.TXT"));
// After the refused renames, NEW.TXT is untouched.
try std.testing.expect(fs.resolve("/NEW.TXT") != null);
}
test "a create stamps the modification time" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
formatFat16(bytes);
var disk = RamDisk{ .bytes = bytes };
var fs = FileSystem.mount(disk.device()).?;
fs.current_time_epoch = 1_700_000_000; // an even-second UTC time
var node = fs.createFile(fs.rootNode(), "STAMP.TXT").?;
_ = fs.writeFile(&node, 0, "hi");
// The persisted entry carries the stamped mtime (even seconds round-trip exactly),
// as does a fresh listing.
try std.testing.expectEqual(@as(u64, 1_700_000_000), fs.resolve("/STAMP.TXT").?.mtime);
try std.testing.expectEqual(@as(u64, 1_700_000_000), fs.listEntry(fs.rootNode(), 0).?.mtime);
}
+55 -15
View File
@@ -6,6 +6,7 @@
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const fs = runtime.fs;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void { fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined; var line: [128]u8 = undefined;
@@ -14,38 +15,37 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
_ = init; _ = init;
const unistd = @import("posix").unistd;
// Wait for /mnt/usb to be mounted — the fat server races us at boot (it must // Wait for /mnt/usb to be mounted — the fat server races us at boot (it must
// bring up the whole USB storage chain first). // bring up the whole USB storage chain first).
var dir: i32 = -1; var opened: ?fs.Directory = null;
var tries: u32 = 0; var tries: u32 = 0;
while (dir < 0 and tries < 1400) : (tries += 1) { while (opened == null and tries < 1400) : (tries += 1) {
dir = unistd.opendir("/mnt/usb"); opened = fs.openDirectory("/mnt/usb");
if (dir < 0) runtime.system.sleep(50); if (opened == null) runtime.system.sleep(50);
} }
if (dir < 0) { var dir = opened orelse {
_ = runtime.system.write("fat-test: /mnt/usb never became available\n"); _ = runtime.system.write("fat-test: /mnt/usb never became available\n");
return; return;
} };
var count: u32 = 0; var count: u32 = 0;
var entry: unistd.DirEntry = .{}; var entry: fs.Entry = .{};
while (unistd.readdir(dir, &entry)) { while (dir.next(&entry)) {
writeLine("fat-test: entry '{s}' kind={d} size={d}\n", .{ entry.name(), entry.kind, entry.size }); writeLine("fat-test: entry '{s}' kind={d} size={d}\n", .{ entry.name(), @intFromEnum(entry.kind), entry.size });
count += 1; count += 1;
if (count > 32) break; if (count > 32) break;
} }
unistd.closedir(dir); dir.close();
writeLine("fat-test: listed {d} entries\n", .{count}); writeLine("fat-test: listed {d} entries\n", .{count});
// Read a known file off the boot volume through the mount (best effort): the // Read a known file off the boot volume through the mount (best effort): the
// kernel image is an ELF, so its first bytes are the ELF magic. // kernel image is an ELF, so its first bytes are the ELF magic.
const fd = unistd.open("/mnt/usb/system/kernel", 0); if (fs.open("/mnt/usb/system/kernel", .{})) |opened_file| {
if (fd >= 0) { var file = opened_file;
var magic: [4]u8 = undefined; var magic: [4]u8 = undefined;
const n = unistd.read(fd, &magic); const n = file.read(&magic) orelse 0;
unistd.close(fd); file.close();
if (n == 4 and magic[0] == 0x7F and magic[1] == 'E' and magic[2] == 'L' and magic[3] == 'F') { if (n == 4 and magic[0] == 0x7F and magic[1] == 'E' and magic[2] == 'L' and magic[3] == 'F') {
_ = runtime.system.write("fat-test: read /mnt/usb/system/kernel ELF magic ok\n"); _ = runtime.system.write("fat-test: read /mnt/usb/system/kernel ELF magic ok\n");
} else { } else {
@@ -53,6 +53,46 @@ pub fn main(init: runtime.process.Init) void {
} }
} }
// Exercise directory + file mutation through the mount: mkdir, create a file
// inside it, read it back, then remove it — proof mkdir/unlink reach the engine.
if (fs.makeDirectory("/mnt/usb/TESTDIR")) {
var wrote = false;
if (fs.open("/mnt/usb/TESTDIR/HELLO.TXT", .{ .create = true, .truncate = true })) |created| {
var f = created;
wrote = (f.writeAll("mutation-ok") orelse 0) == "mutation-ok".len;
f.close();
}
// The created file carries a real modification time (stamped from the RTC).
var mtime_ok = false;
if (fs.attributes("/mnt/usb/TESTDIR/HELLO.TXT")) |attrs| {
writeLine("fat-test: mtime {d}\n", .{attrs.mtime});
mtime_ok = attrs.mtime > 1_577_836_800; // after 2020-01-01
}
if (mtime_ok) _ = runtime.system.write("fat-test: mtime ok\n");
// Rename it, then read from the new name and confirm the old name is gone.
const renamed = fs.rename("/mnt/usb/TESTDIR/HELLO.TXT", "/mnt/usb/TESTDIR/RENAMED.TXT");
const old_gone = !fs.exists("/mnt/usb/TESTDIR/HELLO.TXT");
if (renamed and old_gone) _ = runtime.system.write("fat-test: rename ok\n");
var readback = false;
if (fs.open("/mnt/usb/TESTDIR/RENAMED.TXT", .{})) |reopened| {
var f = reopened;
var buf: [16]u8 = undefined;
const got = f.read(&buf) orelse 0;
f.close();
readback = std.mem.eql(u8, buf[0..got], "mutation-ok");
}
const removed = fs.remove("/mnt/usb/TESTDIR/RENAMED.TXT");
const gone = !fs.exists("/mnt/usb/TESTDIR/RENAMED.TXT");
if (wrote and mtime_ok and renamed and old_gone and readback and removed and gone) {
_ = runtime.system.write("fat-test: mutations ok\n");
} else {
writeLine("fat-test: mutations FAILED (wrote={} mtime={} renamed={} oldgone={} read={} removed={} gone={})\n", .{ wrote, mtime_ok, renamed, old_gone, readback, removed, gone });
}
} else {
_ = runtime.system.write("fat-test: mkdir /mnt/usb/TESTDIR failed\n");
}
if (count > 0) { if (count > 0) {
while (true) { while (true) {
_ = runtime.system.write("fat-test: ok\n"); _ = runtime.system.write("fat-test: ok\n");
+50 -9
View File
@@ -14,7 +14,6 @@ const runtime = @import("runtime");
const engine = @import("engine.zig"); const engine = @import("engine.zig");
const on_disk = @import("on-disk.zig"); const on_disk = @import("on-disk.zig");
const protocol = runtime.vfs_protocol; const protocol = runtime.vfs_protocol;
const unistd = @import("posix").unistd;
const dma = runtime.dma; const dma = runtime.dma;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void { fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
@@ -106,7 +105,7 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
// comes up). From here the VFS routes /mnt/usb/... to this server. // comes up). From here the VFS routes /mnt/usb/... to this server.
var tries: u32 = 0; var tries: u32 = 0;
while (tries < 100) : (tries += 1) { while (tries < 100) : (tries += 1) {
if (unistd.mount(mount_point, endpoint) == 0) { if (runtime.fs.mount(mount_point, endpoint)) {
writeLine("/system/services/fat: mounted {s}\n", .{mount_point}); writeLine("/system/services/fat: mounted {s}\n", .{mount_point});
return true; return true;
} }
@@ -116,16 +115,31 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
return true; // still serve directly, even if the namespace mount didn't take return true; // still serve directly, even if the namespace mount didn't take
} }
const ParentLeaf = struct { parent: []const u8, leaf: []const u8 };
// Split a path into its parent directory and final component: "/a/b" -> ("/a",
// "b"); "/b" -> ("/", "b"); "b" -> ("/", "b").
fn splitParent(path: []const u8) ParentLeaf {
const slash = std.mem.lastIndexOfScalar(u8, path, '/');
return .{
.parent = if (slash) |s| (if (s == 0) "/" else path[0..s]) else "/",
.leaf = if (slash) |s| path[s + 1 ..] else path,
};
}
fn handleOpen(out: []u8, path: []const u8, flags: u32) usize { fn handleOpen(out: []u8, path: []const u8, flags: u32) usize {
var node = filesystem.resolve(path); var node = filesystem.resolve(path);
if (node == null and flags & protocol.create != 0) { if (node == null and flags & protocol.create != 0) {
const slash = std.mem.lastIndexOfScalar(u8, path, '/'); const split = splitParent(path);
const parent_path = if (slash) |s| (if (s == 0) "/" else path[0..s]) else "/"; const parent = filesystem.resolve(split.parent) orelse return fail(out);
const leaf = if (slash) |s| path[s + 1 ..] else path; node = filesystem.createFile(parent, split.leaf);
const parent = filesystem.resolve(parent_path) orelse return fail(out); }
node = filesystem.createFile(parent, leaf); var resolved = node orelse return fail(out);
// O_TRUNC: replace an existing file's contents rather than overwriting in place
// (frees the old chain, so a shorter rewrite leaves no stale tail).
if (flags & protocol.truncate != 0 and !resolved.is_directory) {
filesystem.truncate(&resolved);
} }
const resolved = node orelse return fail(out);
const index = allocOpen() orelse return fail(out); const index = allocOpen() orelse return fail(out);
open_nodes[index] = .{ .used = true, .node = resolved }; open_nodes[index] = .{ .used = true, .node = resolved };
return writeReply(out, .{ .status = 0, .node = index }, &.{}); return writeReply(out, .{ .status = 0, .node = index }, &.{});
@@ -138,6 +152,10 @@ fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.i
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]); const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..]; const payload = message[protocol.request_size..];
// Stamp create/write with the current wall-clock time (mtime). Cheap, and it
// keeps the engine pure (it takes the time as data, not a syscall).
filesystem.current_time_epoch = runtime.system.wallClock();
switch (request.operation) { switch (request.operation) {
.open => return handleOpen(out, payload[0..@min(payload.len, request.len)], request.flags), .open => return handleOpen(out, payload[0..@min(payload.len, request.len)], request.flags),
.read => { .read => {
@@ -156,7 +174,7 @@ fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.i
.status => { .status => {
const o = openAt(request.node) orelse return fail(out); const o = openAt(request.node) orelse return fail(out);
const kind: protocol.NodeKind = if (o.node.is_directory) .directory else .regular; const kind: protocol.NodeKind = if (o.node.is_directory) .directory else .regular;
const status = protocol.FileStatus{ .size = o.node.size, .kind = @intFromEnum(kind) }; const status = protocol.FileStatus{ .size = o.node.size, .kind = @intFromEnum(kind), .mtime = o.node.mtime };
return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&status)); return writeReply(out, .{ .status = 0, .len = @sizeOf(protocol.FileStatus) }, std.mem.asBytes(&status));
}, },
.readdir => { .readdir => {
@@ -176,6 +194,29 @@ fn onMessage(message: []const u8, out: []u8, sender: u32, capability: ?runtime.i
if (openAt(request.node)) |o| o.used = false; if (openAt(request.node)) |o| o.used = false;
return writeReply(out, .{ .status = 0 }, &.{}); return writeReply(out, .{ .status = 0 }, &.{});
}, },
.mkdir => {
const split = splitParent(payload[0..@min(payload.len, request.len)]);
const parent = filesystem.resolve(split.parent) orelse return fail(out);
if (filesystem.createDirectory(parent, split.leaf) == null) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
.unlink => {
const split = splitParent(payload[0..@min(payload.len, request.len)]);
const parent = filesystem.resolve(split.parent) orelse return fail(out);
if (!filesystem.removeFile(parent, split.leaf)) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
.rename => {
const both = payload[0..@min(payload.len, request.len)];
const sep = std.mem.indexOfScalar(u8, both, 0) orelse return fail(out);
const old_split = splitParent(both[0..sep]);
const new_split = splitParent(both[sep + 1 ..]);
// Same-directory rename only.
if (!std.mem.eql(u8, old_split.parent, new_split.parent)) return fail(out);
const parent = filesystem.resolve(old_split.parent) orelse return fail(out);
if (!filesystem.rename(parent, old_split.leaf, new_split.leaf)) return fail(out);
return writeReply(out, .{ .status = 0 }, &.{});
},
// A backend is never itself a mount target. // A backend is never itself a mount target.
.mount, .unmount => return fail(out), .mount, .unmount => return fail(out),
} }
+90
View File
@@ -202,6 +202,96 @@ pub fn geometryOf(sector: []const u8) ?Geometry {
}; };
} }
// --- DOS date/time <-> Unix epoch --------------------------------------------
//
// FAT stamps a file's modification time as two 16-bit DOS fields. There is no
// timezone, so danos treats them as UTC. `date`: year-1980(7)|month(4)|day(5);
// `time`: hour(5)|minute(6)|(second/2)(5).
fn isLeapYear(year: u32) bool {
return (year % 4 == 0 and year % 100 != 0) or (year % 400 == 0);
}
const days_in_month = [_]u8{ 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31 };
/// Convert a FAT date+time to Unix epoch seconds (UTC). Returns 0 for an unset
/// (zero) date.
pub fn fatToEpoch(date: u16, time: u16) u64 {
if (date == 0) return 0;
const day: u32 = date & 0x1F;
const month: u32 = (date >> 5) & 0x0F;
const year: u32 = 1980 + (date >> 9);
if (month < 1 or month > 12 or day < 1) return 0;
const second: u32 = @as(u32, time & 0x1F) * 2;
const minute: u32 = (time >> 5) & 0x3F;
const hour: u32 = (time >> 11) & 0x1F;
var days: u64 = 0;
var y: u32 = 1970;
while (y < year) : (y += 1) days += if (isLeapYear(y)) 366 else 365;
var m: u32 = 1;
while (m < month) : (m += 1) {
days += days_in_month[m - 1];
if (m == 2 and isLeapYear(year)) days += 1;
}
days += day - 1;
return ((days * 24 + hour) * 60 + minute) * 60 + second;
}
pub const FatDateTime = struct { date: u16, time: u16 };
/// Convert Unix epoch seconds (UTC) to a FAT date+time. Returns {0,0} for epoch 0 or
/// any time before 1980 (which DOS cannot represent).
pub fn epochToFatDateTime(epoch: u64) FatDateTime {
if (epoch == 0) return .{ .date = 0, .time = 0 };
var remaining = epoch;
const second: u32 = @intCast(remaining % 60);
remaining /= 60;
const minute: u32 = @intCast(remaining % 60);
remaining /= 60;
const hour: u32 = @intCast(remaining % 24);
remaining /= 24;
var days: u32 = @intCast(remaining); // whole days since 1970-01-01
var year: u32 = 1970;
while (true) {
const y_days: u32 = if (isLeapYear(year)) 366 else 365;
if (days < y_days) break;
days -= y_days;
year += 1;
}
if (year < 1980) return .{ .date = 0, .time = 0 };
var month: u32 = 1;
while (true) {
var m_days: u32 = days_in_month[month - 1];
if (month == 2 and isLeapYear(year)) m_days += 1;
if (days < m_days) break;
days -= m_days;
month += 1;
}
const day = days + 1;
return .{
.date = @intCast(((year - 1980) << 9) | (month << 5) | day),
.time = @intCast((hour << 11) | (minute << 5) | (second / 2)),
};
}
test "FAT date/time <-> Unix epoch round trip" {
// Even-second UTC times (FAT stores seconds/2, so even seconds round-trip exactly).
for ([_]u64{ 1_577_836_800, 1_700_000_000, 1_262_304_000, 1_783_971_244 }) |epoch| {
const fat = epochToFatDateTime(epoch);
try std.testing.expectEqual(epoch, fatToEpoch(fat.date, fat.time));
}
// Absolute check: 1577836800 is 2020-01-01 00:00:00 UTC.
const y2020 = epochToFatDateTime(1_577_836_800);
try std.testing.expectEqual(@as(u16, 2020), 1980 + (y2020.date >> 9));
try std.testing.expectEqual(@as(u16, 1), (y2020.date >> 5) & 0x0F); // month
try std.testing.expectEqual(@as(u16, 1), y2020.date & 0x1F); // day
// 0 is "unset" both ways.
try std.testing.expectEqual(@as(u64, 0), fatToEpoch(0, 0));
try std.testing.expectEqual(@as(u16, 0), epochToFatDateTime(0).date);
}
test "on-disk struct sizes match the specification" { test "on-disk struct sizes match the specification" {
try std.testing.expectEqual(@as(usize, 36), @sizeOf(BiosParameterBlock)); try std.testing.expectEqual(@as(usize, 36), @sizeOf(BiosParameterBlock));
try std.testing.expectEqual(@as(usize, 26), @sizeOf(ExtendedBootRecord16)); try std.testing.expectEqual(@as(usize, 26), @sizeOf(ExtendedBootRecord16));
+41 -3
View File
@@ -21,6 +21,11 @@ const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const power = runtime.power_protocol; const power = runtime.power_protocol;
/// Where the kernel boot log is persisted on the USB FAT volume — an 8.3 name at
/// the mount root (see system/services/log-flush). init writes it at shutdown;
/// the log-flush one-shot writes it once at boot.
const log_path = "/mnt/usb/DANOS.LOG";
/// The system services init brings up at boot, in order. This is init's policy — the /// The system services init brings up at boot, in order. This is init's policy — the
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent /// microkernel keeps such choices in user space, not the kernel. Drivers are absent
/// on purpose: the device manager owns those. (A future init reads this from a /// on purpose: the device manager owns those. (A future init reads this from a
@@ -65,6 +70,14 @@ pub fn main() void {
} }
} }
// Once the storage stack is up, a one-shot copies the boot log to the USB
// volume (/mnt/usb/DANOS.LOG) so it can be read on another machine — the only
// way to see it on a headless/real board with no host capturing serial. Fire
// and forget: it polls for the mount itself, and is deliberately NOT one of
// init's supervised children (a transient one-shot must not be stopped-and-
// waited-for during shutdown).
_ = runtime.system.spawn("log-flush");
// Subscribe to power events (retry: the power service registers well after // Subscribe to power events (retry: the power service registers well after
// init starts). Best-effort — without it, a `terminate` signal still // init starts). Best-effort — without it, a `terminate` signal still
// triggers the same shutdown path. // triggers the same shutdown path.
@@ -115,11 +128,36 @@ fn subscribePower() void {
_ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {}; _ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {};
} }
/// The stop sequence: terminate each child in reverse spawn order (vfs last — /// Copy the whole kernel log to /mnt/usb/DANOS.LOG (the same file log-flush
/// other services may flush through it), waiting up to a deadline for each to /// writes at boot), so a poweroff captures the fullest log. Best-effort: if the
/// exit before killing it, then ask the power service to enter S5. /// USB volume is not mounted, the open fails and it does nothing. Must run while
/// the storage services are still alive (see shutDown).
fn flushKernelLog() void {
// Truncate on open so this fuller flush replaces the boot-time one cleanly.
var file = runtime.fs.open(log_path, .{ .create = true, .truncate = true }) orelse return; // no USB volume
defer file.close();
var chunk: [4096]u8 = undefined;
var offset: usize = 0;
while (true) {
const got = runtime.system.klogRead(offset, &chunk);
if (got == 0) break; // reached the end of the accumulated log
if (file.writeAll(chunk[0..got]) == null) break; // storage went away
offset += got;
}
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, "/system/services/init: flushed log to {s} ({d} bytes)\n", .{ log_path, offset }) catch "");
}
/// The stop sequence: persist the log while storage is still up, then terminate
/// each child in reverse spawn order (vfs last — other services may flush through
/// it), waiting up to a deadline for each to exit before killing it, then ask the
/// power service to enter S5.
fn shutDown() void { fn shutDown() void {
_ = runtime.system.write("/system/services/init: shutting down\n"); _ = runtime.system.write("/system/services/init: shutting down\n");
// Persist the fullest log to the USB volume BEFORE tearing anything down: the
// reverse-order stop loop below kills the fat server (children[3]) first, so
// /mnt/usb must be written while it is still mounted.
flushKernelLog();
var i = child_count; var i = child_count;
while (i > 0) { while (i > 0) {
i -= 1; i -= 1;
+67
View File
@@ -0,0 +1,67 @@
//! system/services/log-flush — a one-shot that copies the kernel's in-memory
//! diagnostic log to a file on the mounted USB FAT volume, so the boot log
//! survives to be read on another machine. On a headless or real board there is
//! no host capturing serial, so without this the log is lost at power-off; this
//! is the on-disk equivalent of QEMU's `-serial file:`.
//!
//! It reads the whole kernel log back through `klog_read` (the RAM sink in
//! system/kernel/log.zig) and writes it to /mnt/usb/DANOS.LOG. The name is 8.3
//! (FAT short-name rule: base <= 8, extension <= 3) and lives at the mount root
//! (there is no mkdir on the FAT path yet). init spawns this once the boot
//! services are up; init itself repeats the flush at shutdown for a fuller log.
//!
//! If no USB volume is mounted — no stick, or the initial-ramdisk sweep that
//! spawns every bundled binary bare with no VFS — it waits briefly, then exits
//! silently, deranging no other test's output.
const std = @import("std");
const runtime = @import("runtime");
const fs = runtime.fs;
const log_path = "/mnt/usb/DANOS.LOG";
/// Copy the whole kernel log to the open file, looping klog_read -> write until
/// the log is exhausted. Returns the number of bytes written.
fn drainKernelLog(file: *fs.File) usize {
var chunk: [4096]u8 = undefined;
var offset: usize = 0;
while (true) {
const got = runtime.system.klogRead(offset, &chunk);
if (got == 0) break; // reached the end of the accumulated log
if (file.writeAll(chunk[0..got]) == null) break; // storage went away
offset += got;
}
return offset;
}
pub fn main() void {
// Wait for the fat server to mount /mnt/usb (it must bring up the whole USB
// storage chain first, so it races us at boot). Bounded: if the mount never
// appears — no volume, or the no-VFS ramdisk sweep — give up silently.
var ready = false;
var tries: u32 = 0;
while (tries < 1400) : (tries += 1) {
if (fs.openDirectory("/mnt/usb")) |directory| {
var dir = directory;
dir.close();
ready = true;
break;
}
runtime.system.sleep(50);
}
if (!ready) return; // /mnt/usb never became available — nothing to persist to
// Truncate on open: each flush replaces the file, so a shorter log on a later
// boot of the same stick leaves no stale tail from a previous, longer one.
var file = fs.open(log_path, .{ .create = true, .truncate = true }) orelse return;
const written = drainKernelLog(&file);
file.close();
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, "log-flush: wrote {d} bytes to {s}\n", .{ written, log_path }) catch return);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+18 -5
View File
@@ -4,12 +4,11 @@
//! header followed by an inline payload (read bytes, or a FileStatus). Everything fits //! header followed by an inline payload (read bytes, or a FileStatus). Everything fits
//! in one IPC message (<= ipc MESSAGE_MAXIMUM = 256 bytes). //! in one IPC message (<= ipc MESSAGE_MAXIMUM = 256 bytes).
//! //!
//! This is a danos-native contract, so it uses danos names throughout — the POSIX //! This is a danos-native contract, so it uses danos names throughout. The client
//! spellings (`stat`, `O_CREAT`, ...) live only in the POSIX layer //! side is `runtime.fs` (library/runtime/fs.zig), which programs use directly.
//! (library/posix/unistd.zig), which translates to these.
//! //!
//! This is user-space only — the kernel knows nothing of files or paths; it only moves the bytes. //! This is user-space only — the kernel knows nothing of files or paths; it only moves the bytes.
//! Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server). //! Shared by library/runtime/fs.zig (client) and system/services/vfs/vfs.zig (server).
pub const Operation = enum(u32) { pub const Operation = enum(u32) {
open, // open(path) -> node id open, // open(path) -> node id
@@ -22,6 +21,13 @@ pub const Operation = enum(u32) {
readdir, // readdir(dir_node, cursor=offset) -> one DirectoryEntry (len==0 => EOF) readdir, // readdir(dir_node, cursor=offset) -> one DirectoryEntry (len==0 => EOF)
mount, // mount(prefix payload, capability = backend endpoint) mount, // mount(prefix payload, capability = backend endpoint)
unmount, // unmount(prefix payload) unmount, // unmount(prefix payload)
// Appended for filesystem mutation (Phase 2). Path-based (the path is the
// payload); a mounted backend handles them, the flat ramfs refuses them.
mkdir, // mkdir(path payload) -> status
unlink, // unlink(path payload) -> status
// rename: the payload is the old path, a single 0x00 separator, then the new
// path. Same-directory rename only (the router requires both under one mount).
rename, // rename(old\0new payload) -> status
}; };
/// The type of a filesystem node, aligned to the FSH file-type table /// The type of a filesystem node, aligned to the FSH file-type table
@@ -75,6 +81,9 @@ pub const FileStatus = extern struct {
size: u64, size: u64,
kind: u32, kind: u32,
_padding: u32 = 0, _padding: u32 = 0,
/// Modification time — Unix epoch seconds, UTC. 0 if the backend has none (the
/// flat ramfs). Filled from the FAT directory entry's write date/time.
mtime: u64 = 0,
}; };
pub const message_maximum: usize = 256; pub const message_maximum: usize = 256;
@@ -83,11 +92,15 @@ pub const reply_size: usize = @sizeOf(Reply);
/// Largest inline payload that still fits one IPC message alongside a header. /// Largest inline payload that still fits one IPC message alongside a header.
pub const maximum_payload: usize = message_maximum - request_size; pub const maximum_payload: usize = message_maximum - request_size;
/// Open flags (danos-native; the POSIX layer maps `O_CREAT` onto `create`). /// Open flags (danos-native; `runtime.fs.OpenOptions` maps its booleans onto these).
pub const create: u32 = 1; pub const create: u32 = 1;
/// Open a directory (for readdir) rather than a file. A mounted backend uses /// Open a directory (for readdir) rather than a file. A mounted backend uses
/// this to open a directory node; the flat ramfs ignores it. /// this to open a directory node; the flat ramfs ignores it.
pub const directory: u32 = 2; pub const directory: u32 = 2;
/// Truncate the file to zero length on open (O_TRUNC): replace its contents rather
/// than overwriting in place, so a shorter new file leaves no stale tail. A mounted
/// backend frees the old cluster chain; the flat ramfs ignores it.
pub const truncate: u32 = 4;
test "protocol struct sizes and node kinds" { test "protocol struct sizes and node kinds" {
const std = @import("std"); const std = @import("std");
+18 -18
View File
@@ -1,26 +1,26 @@
//! /system/services/vfs/vfs-test — a client that proves the VFS round trip end to end: open a //! /system/services/vfs/vfs-test — a client that proves the VFS round trip end to end: open a
//! file through the `runtime` file API, write to it, seek back, read it, and compare. //! file through the `runtime.fs` file API, write to it, seek back, read it, and compare.
//! On success it heartbeats "vfstest: ok" so the kernel test can observe it; //! On success it heartbeats "vfstest: ok" so the kernel test can observe it;
//! on failure it reports what went wrong. Shipped in the initial_ramdisk alongside vfs. //! on failure it reports what went wrong. Shipped in the initial_ramdisk alongside vfs.
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const fs = runtime.fs;
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
const u = @import("posix").unistd;
const payload = "hello-vfs"; const payload = "hello-vfs";
// The "park" role (the vfs-client-death test): open a file, then hold the // The "park" role (the vfs-client-death test): open a file, then hold the
// handle forever without closing — the kill and the VFS's release-on-death // handle forever without closing — the kill and the VFS's release-on-death
// are the point. // are the point.
if (init.arguments.count > 1) { if (init.arguments.count > 1) {
var fd: i32 = -1; var parked: ?fs.File = null;
var tries: u32 = 0; var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) { while (parked == null and tries < 200) : (tries += 1) {
fd = u.open("parked", u.O_CREAT); parked = fs.open("parked", .{ .create = true });
if (fd < 0) runtime.system.sleep(20); if (parked == null) runtime.system.sleep(20);
} }
if (fd < 0) { if (parked == null) {
_ = runtime.system.write("vfstest: park open failed\n"); _ = runtime.system.write("vfstest: park open failed\n");
return; return;
} }
@@ -31,28 +31,28 @@ pub fn main(init: runtime.process.Init) void {
} }
// The VFS server may not have registered yet — retry open until it's up. // The VFS server may not have registered yet — retry open until it's up.
var fd: i32 = -1; var opened: ?fs.File = null;
var tries: u32 = 0; var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) { while (opened == null and tries < 200) : (tries += 1) {
fd = u.open("greeting", u.O_CREAT); opened = fs.open("greeting", .{ .create = true });
if (fd < 0) runtime.system.sleep(20); if (opened == null) runtime.system.sleep(20);
} }
if (fd < 0) { var greeting = opened orelse {
_ = runtime.system.write("vfstest: open failed\n"); _ = runtime.system.write("vfstest: open failed\n");
return; return;
} };
if (u.write(fd, payload) != @as(isize, payload.len)) { if ((greeting.write(payload) orelse 0) != payload.len) {
_ = runtime.system.write("vfstest: write failed\n"); _ = runtime.system.write("vfstest: write failed\n");
return; return;
} }
_ = u.lseek(fd, 0, u.SEEK_SET); greeting.seekTo(0);
var buffer: [32]u8 = undefined; var buffer: [32]u8 = undefined;
const n = u.read(fd, &buffer); const n = greeting.read(&buffer) orelse 0;
u.close(fd); greeting.close();
if (n == @as(isize, payload.len) and std.mem.eql(u8, buffer[0..@intCast(n)], payload)) { if (n == payload.len and std.mem.eql(u8, buffer[0..n], payload)) {
while (true) { while (true) {
_ = runtime.system.write("vfstest: ok\n"); _ = runtime.system.write("vfstest: ok\n");
runtime.system.sleep(1000); runtime.system.sleep(1000);
+57
View File
@@ -161,6 +161,44 @@ fn forwardRequest(out: []u8, backend: ipc.Handle, request: protocol.Request, pay
return copy; return copy;
} }
/// Forward a path-based operation (mkdir, unlink) under a mount to its backend and
/// relay the reply. No handle is created — these operate by path and return only a
/// status.
fn forwardPath(out: []u8, backend: ipc.Handle, operation: protocol.Operation, relative: []const u8) usize {
const request = protocol.Request{ .operation = operation, .node = 0, .offset = 0, .len = @intCast(relative.len), .flags = 0 };
var message: [protocol.message_maximum]u8 = undefined;
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
const rel = relative[0..@min(relative.len, protocol.maximum_payload)];
@memcpy(message[protocol.request_size..][0..rel.len], rel);
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0 .. protocol.request_size + rel.len], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Forward a rename to its backend: the payload is the mount-relative old path, a
/// 0x00 separator, then the mount-relative new path. Relays the backend's reply.
fn forwardRename(out: []u8, backend: ipc.Handle, old_relative: []const u8, new_relative: []const u8) usize {
const total = old_relative.len + 1 + new_relative.len;
var message: [protocol.message_maximum]u8 = undefined;
if (protocol.request_size + total > message.len) return fail(out);
const request = protocol.Request{ .operation = .rename, .node = 0, .offset = 0, .len = @intCast(total), .flags = 0 };
@memcpy(message[0..protocol.request_size], std.mem.asBytes(&request));
var p = protocol.request_size;
@memcpy(message[p..][0..old_relative.len], old_relative);
p += old_relative.len;
message[p] = 0;
p += 1;
@memcpy(message[p..][0..new_relative.len], new_relative);
p += new_relative.len;
var reply: [protocol.message_maximum]u8 = undefined;
const n = ipc.call(backend, message[0..p], &reply) catch return fail(out);
const copy = @min(n, out.len);
@memcpy(out[0..copy], reply[0..copy]);
return copy;
}
/// Best-effort close of a backend node (used when a dead client's forwarding /// Best-effort close of a backend node (used when a dead client's forwarding
/// handles are swept — the backend must not leak the vfs's opens). /// handles are swept — the backend must not leak the vfs's opens).
fn forwardClose(backend: ipc.Handle, backend_node: u64) void { fn forwardClose(backend: ipc.Handle, backend_node: u64) void {
@@ -305,6 +343,25 @@ fn handle(message: []const u8, out: []u8, sender: u32, capability: ?ipc.Handle)
} }
return writeReply(out, .{ .status = 0 }, &.{}); return writeReply(out, .{ .status = 0 }, &.{});
}, },
.mkdir, .unlink => {
const name = payload[0..@min(payload.len, request.len)];
if (longestMount(name)) |m| return forwardPath(out, m.backend, request.operation, m.relative);
// Only a mounted backend has real directories; the flat ramfs cannot
// create or remove them (and a bare-name path is not a mount target).
return fail(out);
},
.rename => {
const both = payload[0..@min(payload.len, request.len)];
const sep = std.mem.indexOfScalar(u8, both, 0) orelse return fail(out);
const old_path = both[0..sep];
const new_path = both[sep + 1 ..];
const mo = longestMount(old_path) orelse return fail(out);
const mn = longestMount(new_path) orelse return fail(out);
// Both paths must live under the same mount — cross-filesystem rename is
// not supported.
if (mo.backend != mn.backend) return fail(out);
return forwardRename(out, mo.backend, mo.relative, mn.relative);
},
} }
} }
+44
View File
@@ -108,6 +108,11 @@ CASES = [
{"name": "clock", {"name": "clock",
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# Wall-clock: the CMOS RTC read at boot gives a plausible current epoch (the
# foundation for filesystem mtime).
{"name": "wall-clock",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "vmm", {"name": "vmm",
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -320,6 +325,31 @@ CASES = [
"timeout": 150, "timeout": 150,
"expect": r"fat: mounted /mnt/usb[\s\S]*fat-test: ok", "expect": r"fat: mounted /mnt/usb[\s\S]*fat-test: ok",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the
# fat-test client, after listing, makes a directory, writes+reads a file inside
# it, then removes the file, exercising the whole VFS -> fat mutation path.
{"name": "fat-mutations",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: mutations ok",
"fail": r"fat-test: mutations FAILED|fat-test: mkdir .* failed|DANOS-TEST-RESULT: FAIL"},
# Phase 2c: rename through the mount — fat-test renames the file it created
# before removing it, and confirms the old name is gone.
{"name": "fat-rename",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: rename ok",
"fail": r"fat-test: mutations FAILED|DANOS-TEST-RESULT: FAIL"},
# Phase 2d: filesystem timestamps — a freshly-created file's mtime is a real
# current wall-clock time (stamped from the RTC), read back through stat.
{"name": "fat-mtime",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"expect": r"fat-test: mtime ok",
"fail": r"fat-test: mutations FAILED|DANOS-TEST-RESULT: FAIL"},
# Boot-from-USB smoke: the whole system now boots off the FAT32 image on a # Boot-from-USB smoke: the whole system now boots off the FAT32 image on a
# usb-storage device (OVMF -> \EFI\BOOT\BOOTX64.efi -> kernel), so the kernel # usb-storage device (OVMF -> \EFI\BOOT\BOOTX64.efi -> kernel), so the kernel
# reaching its PASS marker at all proves the USB boot path end to end. Reuses # reaching its PASS marker at all proves the USB boot path end to end. Reuses
@@ -368,6 +398,20 @@ CASES = [
r"init: shutting down[\s\S]*" r"init: shutting down[\s\S]*"
r"power: entering S5", r"power: entering S5",
"fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"}, "fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"},
# M8: the boot log is persisted to the USB FAT volume. Reuses the orderly-
# shutdown build (full tree + power button): init spawns log-flush at boot,
# which copies the kernel log to /mnt/usb/DANOS.LOG once /mnt/usb is mounted
# (first marker); then the power button drives init's own pre-teardown flush
# (second marker), proving both triggers write the file while storage is up.
{"name": "log-flush",
"build_case": "orderly-shutdown",
"smp": 4,
"timeout": 150,
"qmp_after": {"delay": 8, "command": "system_powerdown"},
"expect": r"log-flush: wrote \d+ bytes to /mnt/usb/DANOS\.LOG[\s\S]*"
r"init: flushed log to /mnt/usb/DANOS\.LOG[\s\S]*"
r"power: entering S5",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers + # M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers +
# reports its _HID devices — the two PS/2 nodes must appear with resources # reports its _HID devices — the two PS/2 nodes must appear with resources
# (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/discovery.md). # (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/discovery.md).