9 Commits
Author SHA1 Message Date
Daniel Samson 7082699f5f gitignore: /var/log — real-hardware log pulls stay out of the tree
(An earlier append landed on a line missing its newline, mangling
'.github/' and silently un-ignoring var/ — which let a stray add sweep
pulled logs into the tree. Repaired, and the logs untracked again.)
2026-07-21 20:09:29 +01:00
Daniel Samson f0611ef8ac logger: logger.log becomes the completeness receipt
Its only content was the redundant announce line — the useful shutdown
marker was written AFTER the final drain, so it reached serial and the
ring but never a file, and a truncated boot could only be inferred from
what was missing. The marker now enters the ring BEFORE the drain, so
the drain carries it into logger.log: a directory whose logger.log ends
with 'shutting down; final flush' is complete through shutdown; one
without it was cut early. The sequence-count epilogue stays serial-only,
after the drain, by design.
2026-07-21 20:05:13 +01:00
Daniel Samson eb6e8edafe apic: time-bound the TSC warp check — 47 s of AP bring-up becomes ~0.3 s
First real per-process logs off the stick (the logging track paying for
itself): kernel.log showed 16-core bring-up costing 47 s — per-core gaps
of 2-21 s — on a machine whose clocksource is the TSC, so every AP runs
the pairwise warp check. QEMU always picks HPET, so the harness never
executed this path at all.

The check was bounded by ITERATIONS: 1<<20 warp ticks, each a locked
read-modify-write on a cacheline two cores fight over — microseconds
under real contention, not the nanosecond the '~1 ms' comment assumed —
and the 1<<32-PAUSE rendezvous 'bound' is ~2 minutes on modern Intel
(PAUSE ~140 cycles). Both are now bounded by TIME measured on the TSC
itself: ~5 ms of pairwise hammering per core (Linux's check_tsc_warp
budget — ample to catch a lagging TSC) and a ~100 ms rendezvous window.
2026-07-21 19:47:53 +01:00
Daniel Samson 59ba95a315 usb: don't reset an enabled SuperSpeed port; name every setup failure
Real-PC diagnose boot (photo + OCR): the boot stick connects at
SuperSpeed on port 21 and 'device setup failed' lands in the SAME
millisecond — an instant failure, not a timeout. setupDevice reset
every port unconditionally; that is required to enable USB2 ports, but
a SuperSpeed port that trained its link is ALREADY enabled (xHCI
advances USB3 ports to Enabled, no reset — spec 4.3), and driving a hot
reset into the live link drops PED mid-reset on real silicon. QEMU
tolerates the spurious reset, which is why the harness never saw it.

An enabled speed>=4 port now skips the reset (a not-yet-enabled SS link
still gets one). Every setup step names its failure — port reset with
the PORTSC value, Enable Slot, device-slot exhaustion, Address Device
with its completion code — so the on-screen transcript of the next
failure identifies the exact xHCI command instead of one blanket line.
block.open's give-up window drops 60 s -> 30 s (the slowest observed
healthy chain completed at ~24 s); a machine whose stick failed setup
should not sit a further minute pretending otherwise.
2026-07-21 19:37:24 +01:00
Daniel Samson f587e7e05e console: page-wrap instead of scrolling — never read the framebuffer
The scroll path copied every pixel row up by one glyph height, READING
video memory — and VRAM reads are uncached-slow on real hardware.
Measured on the 16-core PC: ~90 seconds to bring the cores online,
almost entirely boot-transcript lines each paying a whole-screen scroll
copy. (QEMU never shows this: its 'VRAM' is host RAM.)

When the screen fills, the console now clears and restarts at the top —
writes only, once per screenful. The transcript reads the same as it
streams; only the scrollback illusion is gone, which a boot console
never needed.
2026-07-21 19:22:59 +01:00
Daniel Samson b541921218 build: -Ddiagnose — boot without the display so the transcript stays on screen
The display service claiming the framebuffer suppresses the on-screen
boot transcript — correctly in normal operation, but on a serial-less
machine being debugged, the timeline vanishes just when it matters. A
diagnose image (zig build -Ddiagnose=true) has init skip the display
service and demo: the timestamped transcript stays on screen
indefinitely, and the power button still runs the orderly shutdown (so
the logger's files land when storage works).

First use immediately caught a real bug IN QEMU: an intermittent (~1
in 3) usb-storage READ CAPACITY failure after a ~21 s stall — after
which usb-storage and fat both exit cleanly and nothing retries: one
transient early-boot USB failure leaves the system permanently without
storage (and therefore without logs). That no-retry policy is a prime
suspect for real hardware never mounting /var, and is Track B's first
work item.
2026-07-21 19:15:14 +01:00
Daniel Samson a91365b3d9 kernel: the boot transcript on screen, every line timestamped
Without serial and without working USB storage, a slow real-hardware
boot is undiagnosable — 'stabbing in the dark'. Two changes end that:

The log renderer stamps every line with boot-relative seconds
([  12.045] ...), so every surface — serial, debugcon, and now the
screen — is a readable timeline. And the framebuffer console registers
as an ordinary log sink at boot: kernel AND userspace lines (device
bring-up, fat mounts, logger announcements) show live on screen until
the display service claims the framebuffer, which flips the console's
suppression and silences the sink automatically — the display-owns-the-
screen design is unchanged in normal operation; the console now simply
narrates the part of boot that happens before there IS a display.

On the machine that motivated this, the next boot will show by eye
where the minutes go — including whether the USB chain ever brings
storage up, and whether screen drawing itself crawls (the latent
non-write-combining framebuffer suspect: if these very lines paint
slowly, that's the answer).
2026-07-21 19:09:34 +01:00
Daniel Samson ffa45edc8b boot: the capsule — one-file system image first, manifest and walk as fallbacks
Real-firmware finding: the per-file /system tree walk boots in seconds
under OVMF but stalls for MINUTES on real firmware — the cost is not
bytes (USB 3 moves the ~4 MB instantly) but firmware filesystem
OPERATIONS: ~30 opens, each an uncached directory-chain walk in an
unoptimized firmware FAT driver. This is why every real OS loader
(winload, GRUB) reads many files through its own filesystem code over
Block I/O rather than the firmware's file protocol.

The loader now reads boot/system.img — the bundled binaries packed into
ONE v2 initial_ramdisk (tools/pack-system-image.py, derived from the
same bundled list in the same build graph, so tree and capsule cannot
drift) — with a single open + sequential read, the one firmware file
I/O shape that is fast everywhere. The manifest (open each listed path
by name) and the tree walk remain as fallbacks, so a hand-assembled
stick without the capsule still boots. The running system is identical
in all three cases: the kernel receives the same in-RAM table.

The load phase now brackets itself with unconditional on-screen
breadcrumbs ('EFI: loading the system...' / '...starting the kernel'),
because this phase stalling behind a silent black screen — kernel
status is serial-only by design — already cost a real-hardware
debugging session.

Direction (settled with the user): this capsule becomes the BOOTSTRAP
capsule — kernel + init + the storage-bring-up set — once a
spawn-from-memory syscall lets init and the device manager load
everything else from the stick's real file tree at runtime through
danos's own storage stack: file-granular updates (rebuild one binary,
copy one file), the initramfs shape.
2026-07-21 18:57:33 +01:00
Daniel Samson d446ddd2ed boot: survive hand-written sticks — case-fold the /system walk, skip host litter
The loader ENUMERATES /system now, so it sees whatever the stick's
directory entries literally store — and a hand-copied stick differs
from our generated image: firmware returns bare 8.3 short entries
UPPERCASE (INIT, SYSTEM), and host OSes leave litter next to every file
(macOS '._' AppleDouble forks, .fseventsd). Verified in QEMU/OVMF: an
uppercase-stored volume booted to a dead kernel-only system before this
change and boots fully (display up, logger writing /var/log) after.

The walk now lowers ASCII names (the danos tree is canonically
lowercase; FAT lookups are case-insensitive by definition), skips any
dot-prefixed entry, and treats malformed or unreadable entries as
skip-this-file instead of abort-the-whole-walk. Reader.find compares
case-insensitively as belt and braces.
2026-07-21 18:01:12 +01:00
13 changed files with 332 additions and 56 deletions
+2 -1
View File
@@ -6,4 +6,5 @@ zig-out/
.idea/
.claude/
.github/
.github/
/var/log/
+125 -17
View File
@@ -380,11 +380,19 @@ const Bundled = struct {
data: []align(8) u8,
};
/// Walk the boot volume's /system tree and pack every regular file (except the
/// kernel image itself — the only top-level file) into an in-RAM v2
/// initial_ramdisk image, entries named by full FHS path. This is what makes the
/// volume's file structure the single source of truth: there is no packed
/// ramdisk artifact on disk, and init travels in the table like everything else.
/// Gather the boot volume's user binaries into an in-RAM v2 initial_ramdisk
/// image, entries named by full FHS path — the volume's file structure is the
/// single source of truth (no packed ramdisk artifact; init travels in the
/// table like everything else).
///
/// Two strategies, most portable first:
/// 1. /system/manifest (written by the build): each listed path is opened BY
/// NAME — the case-insensitive lookup every firmware FAT driver gets
/// right, and the only file access the pre-tree loader ever used.
/// 2. No manifest: ENUMERATE the /system tree. Portable in principle, but
/// firmware differs in what names enumeration returns (bare 8.3 entries
/// come back uppercase on some drivers), so this is the fallback for
/// hand-assembled sticks, not the primary path.
fn loadSystemTree(bs: *uefi.tables.BootServices, boot_information: *BootInformation) !void {
const loaded = (try bs.handleProtocol(uefi.protocol.LoadedImage, uefi.handle)) orelse
return error.NoLoadedImage;
@@ -395,12 +403,30 @@ fn loadSystemTree(bs: *uefi.tables.BootServices, boot_information: *BootInformat
const root = try fs.openVolume();
defer _ = root.close() catch {};
const system_directory = try root.open(system_directory_name, .read, .{});
defer _ = system_directory.close() catch {};
// Unconditional breadcrumb (con_out, independent of -Dserial): this phase
// is where a slow firmware stalls, and a silent black screen here already
// cost a real-hardware debugging session.
log("EFI: loading the system...\r\n");
// The capsule (boot\system.img) first: one open + one sequential read is
// the only firmware file I/O shape that is fast everywhere. It is already
// the kernel's wire format — hand it over as-is.
if (loadCapsule(bs, root, boot_information)) {
log("EFI: system image loaded, starting the kernel\r\n");
return;
}
var list: [maximum_bundled]Bundled = undefined;
var count: usize = 0;
try walkDirectory(bs, system_directory, "/system", 0, &list, &count);
loadByManifest(bs, root, &list, &count) catch {
count = 0; // a torn manifest read leaves partial entries; start over
};
if (count == 0) {
const system_directory = try root.open(system_directory_name, .read, .{});
defer _ = system_directory.close() catch {};
try walkDirectory(bs, system_directory, "/system", 0, &list, &count);
}
if (count == 0) return error.NoBinaries;
// Assemble the v2 image: header, entry table, then the blobs.
@@ -426,7 +452,72 @@ fn loadSystemTree(bs: *uefi.tables.BootServices, boot_information: *BootInformat
boot_information.initial_ramdisk_base = @intFromPtr(image.ptr);
boot_information.initial_ramdisk_len = total;
progress("EFI: /system tree loaded\r\n");
log("EFI: /system tree loaded, starting the kernel\r\n");
}
/// The boot capsule: the bundled binaries as one v2 initial_ramdisk image.
const capsule_file_name = std.unicode.utf8ToUtf16LeStringLiteral("boot\\system.img");
/// Load boot\system.img whole and hand it to the kernel unmodified — it is
/// already the initial_ramdisk wire format. Returns false (capsule absent or
/// unreadable or wrong magic) to let the caller fall back to per-file loading.
fn loadCapsule(bs: *uefi.tables.BootServices, root: *uefi.protocol.File, boot_information: *BootInformation) bool {
const file = root.open(capsule_file_name, .read, .{}) catch return false;
defer _ = file.close() catch {};
const image = readWholeFile(bs, file) catch return false;
if (image.len < @sizeOf(initial_ramdisk.Header) or
std.mem.bytesToValue(initial_ramdisk.Header, image[0..@sizeOf(initial_ramdisk.Header)]).magic != initial_ramdisk.magic)
{
_ = bs.freePool(image.ptr) catch {};
return false;
}
boot_information.initial_ramdisk_base = @intFromPtr(image.ptr);
boot_information.initial_ramdisk_len = image.len;
return true;
}
/// The manifest path, and a scratch limit for its UTF-16 conversion.
const manifest_file_name = std.unicode.utf8ToUtf16LeStringLiteral("system\\manifest");
/// Load every binary the manifest lists, opening each path by name from the
/// volume root. A listed-but-unopenable file is skipped (the kernel reports the
/// absence); a missing manifest errors so the caller falls back to the walk.
fn loadByManifest(bs: *uefi.tables.BootServices, root: *uefi.protocol.File, list: *[maximum_bundled]Bundled, count: *usize) !void {
const manifest_handle = try root.open(manifest_file_name, .read, .{});
var manifest_open = true;
defer if (manifest_open) {
_ = manifest_handle.close() catch {};
};
const manifest = try readWholeFile(bs, manifest_handle);
_ = manifest_handle.close() catch {};
manifest_open = false;
defer _ = bs.freePool(manifest.ptr) catch {};
var lines = std.mem.tokenizeAny(u8, manifest, "\r\n");
while (lines.next()) |line| {
if (line.len < 2 or line[0] != '/') continue;
if (line.len >= initial_ramdisk.maximum_name) continue;
if (count.* == maximum_bundled) return;
// "/system/services/init" -> UTF-16 "system\services\init".
var name16: [initial_ramdisk.maximum_name]u16 = undefined;
var i: usize = 0;
for (line[1..]) |c| {
name16[i] = if (c == '/') '\\' else c;
i += 1;
}
name16[i] = 0;
const file = root.open(@ptrCast(name16[0..i :0]), .read, .{}) catch continue;
defer _ = file.close() catch {};
const data = readWholeFile(bs, file) catch continue;
var entry: *Bundled = &list[count.*];
@memcpy(entry.path[0..line.len], line);
entry.path_len = line.len;
entry.data = data;
count.* += 1;
}
}
/// Recursively collect the regular files below `directory` into `list`. Top-level
@@ -448,17 +539,34 @@ fn walkDirectory(
const info: *const uefi.protocol.File.Info.File = @ptrCast(@alignCast(&info_buffer));
const name16 = info.getFileName();
// Convert the (ASCII in practice) UTF-16 name; skip "." and "..".
// Convert the (ASCII in practice) UTF-16 name. A hostile-shaped entry
// (too long, non-ASCII) is SKIPPED, never fatal — one odd file on a
// hand-written stick must not cost the whole boot. Names are lowered:
// the danos tree is canonically lowercase and FAT lookups are
// case-insensitive, but firmware ENUMERATION returns whatever the
// directory stores — an 8.3 short entry comes back uppercase ("INIT"),
// which would otherwise poison every path comparison downstream.
var name_buffer: [initial_ramdisk.maximum_name]u8 = undefined;
var name_length: usize = 0;
var name_ok = true;
while (name16[name_length] != 0) : (name_length += 1) {
if (name_length == name_buffer.len) return error.NameTooLong;
if (name_length == name_buffer.len) {
name_ok = false;
break;
}
const c = name16[name_length];
if (c > 0x7F) return error.UnsupportedName;
name_buffer[name_length] = @intCast(c);
if (c > 0x7F) {
name_ok = false;
break;
}
name_buffer[name_length] = std.ascii.toLower(@intCast(c));
}
if (!name_ok) continue;
const name = name_buffer[0..name_length];
if (std.mem.eql(u8, name, ".") or std.mem.eql(u8, name, "..")) continue;
// Skip dot entries: "." / ".." and host-OS litter (macOS "._*" AppleDouble
// resource forks, ".fseventsd", ".Spotlight-V100") a copied-onto stick
// accumulates — none of it is a danos binary.
if (name.len == 0 or name[0] == '.') continue;
if (info.attribute.directory) {
if (depth == maximum_tree_depth) continue;
@@ -474,12 +582,12 @@ fn walkDirectory(
if (count.* == maximum_bundled) return error.TooManyBinaries;
var entry: *Bundled = &list[count.*];
const path = try std.fmt.bufPrint(&entry.path, "{s}/{s}", .{ prefix, name });
const path = std.fmt.bufPrint(&entry.path, "{s}/{s}", .{ prefix, name }) catch continue; // path too long: skip the file, keep the boot
entry.path_len = path.len;
const file = try directory.open(name16, .read, .{});
const file = directory.open(name16, .read, .{}) catch continue;
defer _ = file.close() catch {};
entry.data = try readWholeFile(bs, file);
entry.data = readWholeFile(bs, file) catch continue; // unreadable/empty: skip
count.* += 1;
}
}
+47 -2
View File
@@ -235,6 +235,8 @@ fn addBootImage(
b: *std.Build,
kernel_bin: std.Build.LazyPath,
efi_bin: std.Build.LazyPath,
manifest: std.Build.LazyPath,
capsule: std.Build.LazyPath,
bundled: []const BundledBinary,
) std.Build.LazyPath {
const mk_fat = b.addSystemCommand(&.{"python3"});
@@ -245,6 +247,10 @@ fn addBootImage(
mk_fat.addFileArg(efi_bin);
mk_fat.addArg("system/kernel");
mk_fat.addFileArg(kernel_bin);
mk_fat.addArg("system/manifest");
mk_fat.addFileArg(manifest);
mk_fat.addArg("boot/system.img");
mk_fat.addFileArg(capsule);
for (bundled) |item| {
mk_fat.addArg(item.path);
mk_fat.addFileArg(item.binary);
@@ -462,6 +468,10 @@ pub fn build(b: *std.Build) void {
// QEMU test harness (test/qemu_test.py, which asserts on serial markers) turn
// it on; a flashable `zig build` image leaves it out. See serial.zig.
const serial = b.option(bool, "serial", "Compile the serial-console log sink into the kernel (default: off; run-x86-64 and the test harness enable it)") orelse false;
// The diagnose boot: init skips the display service (and demo), so the
// on-screen boot transcript is never suppressed — the full timestamped
// timeline stays on the screen for real-hardware debugging by eye.
const diagnose = b.option(bool, "diagnose", "Boot without the display service so the timestamped boot transcript stays on screen (real-hardware debugging)") orelse false;
// --- Kernel: freestanding x86_64 ELF, jumped to by the bootloader ---
// SSE2 is part of the x86_64 baseline and UEFI leaves it enabled at handoff,
@@ -505,6 +515,7 @@ pub fn build(b: *std.Build) void {
// the heartbeat stays present under test.
const init_options = b.addOptions();
init_options.addOption(bool, "serial", serial);
init_options.addOption(bool, "diagnose", diagnose);
programModule(init_exe).addImport("build_options", init_options.createModule());
// --- the rest of the /system tree: services, drivers, test fixtures ---
@@ -622,6 +633,40 @@ pub fn build(b: *std.Build) void {
.{ .path = "system/tests/thread-test", .binary = thread_test_exe.getEmittedBin() },
};
// The boot manifest: the FHS path of every bundled binary, one per line. The
// EFI loader reads THIS by name and opens each listed path by name — FAT
// name lookup is case-insensitive and firmware-portable, unlike directory
// ENUMERATION, whose returned names vary by firmware (bare 8.3 entries come
// back uppercase on some FAT drivers). The tree walk remains only as the
// loader's fallback for hand-assembled sticks without a manifest.
var manifest_text: std.ArrayListUnmanaged(u8) = .empty;
for (bundled) |item| {
manifest_text.append(b.allocator, '/') catch @panic("OOM");
manifest_text.appendSlice(b.allocator, item.path) catch @panic("OOM");
manifest_text.append(b.allocator, '\n') catch @panic("OOM");
}
const manifest_files = b.addWriteFiles();
const manifest_file = manifest_files.add("manifest", manifest_text.items);
const manifest_install = b.addInstallFileWithDir(manifest_file, .prefix, "system/manifest");
b.getInstallStep().dependOn(&manifest_install.step);
// The boot capsule: the same bundled list packed into ONE file (v2
// initial_ramdisk format), because a single open + sequential read is the
// only firmware file I/O shape that is fast everywhere — a per-file tree
// walk measured MINUTES on real firmware. The loader tries this first,
// then the manifest, then the walk; the running system cannot tell the
// difference (it always receives the same in-RAM table). Derived from the
// tree in the same build graph, so the two cannot drift.
const mk_capsule = b.addSystemCommand(&.{"python3"});
mk_capsule.addFileArg(b.path("tools/pack-system-image.py"));
const capsule_img = mk_capsule.addOutputFileArg("system.img");
for (bundled) |item| {
mk_capsule.addArg(item.path);
mk_capsule.addFileArg(item.binary);
}
const capsule_install = b.addInstallFile(capsule_img, "boot/system.img");
b.getInstallStep().dependOn(&capsule_install.step);
// Install every bundled binary to its FHS home, so zig-out is a true image of
// the filesystem — the same tree make-fat-image.py lays out on the boot volume.
for (bundled) |item| {
@@ -669,7 +714,7 @@ pub fn build(b: *std.Build) void {
// binaries at their FHS paths. QEMU presents this image as a USB mass-storage
// device the guest boots from (see run-x86-64 and the test harness), and the
// danos fat driver mounts the same image at /mnt/usb.
const fat_image = addBootImage(b, exe.getEmittedBin(), efiexe.getEmittedBin(), &bundled);
const fat_image = addBootImage(b, exe.getEmittedBin(), efiexe.getEmittedBin(), manifest_file, capsule_img, &bundled);
const fat_image_install = b.addInstallFile(fat_image, "danos-usb.img");
b.getInstallStep().dependOn(&fat_image_install.step);
@@ -678,7 +723,7 @@ pub fn build(b: *std.Build) void {
// log captured to serial0 — without baking serial into the image users flash.
// Built lazily (only when `run-x86-64` is requested), and never installed.
const exe_serial = addKernel(b, kernel_target, optimize, kernel_modules, test_case, true);
const fat_image_serial = addBootImage(b, exe_serial.getEmittedBin(), efiexe.getEmittedBin(), &bundled);
const fat_image_serial = addBootImage(b, exe_serial.getEmittedBin(), efiexe.getEmittedBin(), manifest_file, capsule_img, &bundled);
// `zig build check-fat-image` — validate the produced image is a real FAT32
// with the EFI stub present (the builder's own --verify, no external tools).
+4 -1
View File
@@ -61,7 +61,10 @@ pub fn open() ?Device {
// enumeration, mass-storage bring-up) must complete first, which can take
// tens of seconds under emulation.
var attempts: usize = 0;
while (attempts < 1200) : (attempts += 1) {
// 30 s covers the slowest observed healthy chain (a flaky QEMU enumeration
// completed at ~24 s); a machine whose stick genuinely failed setup should
// not sit a further minute pretending otherwise.
while (attempts < 600) : (attempts += 1) {
if (ipc.lookup(.block)) |handle| return .{ .endpoint = handle };
system.sleep(50);
}
@@ -620,8 +620,15 @@ pub const Controller = struct {
.parameter = device.input_context.physical,
.control = trbControl(.address_device, @as(u32, device.slot_id) << 24),
});
const code = self.awaitCommand(physical) orelse return false;
return code == @intFromEnum(CompletionCode.success);
const code = self.awaitCommand(physical) orelse {
std.log.info("port {d} setup: Address Device timed out", .{device.port});
return false;
};
if (code != @intFromEnum(CompletionCode.success)) {
std.log.info("port {d} setup: Address Device completion code {d}", .{ device.port, code });
return false;
}
return true;
}
/// Reset the port, enable a slot, and address the device on it: after this the
@@ -629,9 +636,25 @@ pub const Controller = struct {
/// or null on any failure. The EP0 MPS is taken from the speed default and
/// corrected from the device descriptor by `refreshMaxPacketSize0` if needed.
pub fn setupDevice(self: *Controller, port: u32, speed: u32) ?*Device {
if (!self.resetPort(port)) return null;
const slot_id = self.enableSlot() orelse return null;
const device = self.allocateDevice() orelse return null;
// A SuperSpeed port that has trained its link is ALREADY enabled — the
// xHCI advances USB3 ports to Enabled with no reset (spec 4.3). Driving
// a hot reset into a live SS link drops PED mid-reset on real silicon
// (observed: "setup failed" in the same millisecond as "connected").
// Only a not-yet-enabled port — every USB2 device, or a stuck SS link —
// needs the reset to enable.
const already_enabled = speed >= 4 and self.portStatus(port) & portsc_enabled != 0;
if (!already_enabled and !self.resetPort(port)) {
std.log.info("port {d} setup: port reset failed (PORTSC 0x{x:0>8})", .{ port, self.portStatus(port) });
return null;
}
const slot_id = self.enableSlot() orelse {
std.log.info("port {d} setup: Enable Slot failed", .{port});
return null;
};
const device = self.allocateDevice() orelse {
std.log.info("port {d} setup: no free device slot", .{port});
return null;
};
device.* = .{
.used = true,
.slot_id = slot_id,
@@ -652,6 +675,7 @@ pub const Controller = struct {
return device;
}
fn abandon(self: *Controller, device: *Device) ?*Device {
_ = self;
device.used = false;
+6 -3
View File
@@ -76,17 +76,20 @@ pub const Reader = struct {
/// Look a binary up by name: an exact path match wins; otherwise a unique
/// basename match ("fat" finds "/system/services/fat") keeps pre-path callers
/// working. The returned Item's name is always the stored full path.
/// working. Comparisons are ASCII case-insensitive — the entries come from a
/// FAT volume, whose name lookups are case-insensitive by definition (and
/// whose short entries store uppercase). The returned Item's name is always
/// the stored full path.
pub fn find(self: Reader, name: []const u8) ?Item {
var i: u32 = 0;
while (i < self.count) : (i += 1) {
const item = self.entry(i) orelse continue;
if (std.mem.eql(u8, item.name, name)) return item;
if (std.ascii.eqlIgnoreCase(item.name, name)) return item;
}
i = 0;
while (i < self.count) : (i += 1) {
const item = self.entry(i) orelse continue;
if (std.mem.eql(u8, basename(item.name), name)) return item;
if (std.ascii.eqlIgnoreCase(basename(item.name), name)) return item;
}
return null;
}
+26 -11
View File
@@ -529,8 +529,22 @@ var warp_ap_ready: u32 = 0;
var warp_stop: u32 = 0;
var warp_checks: u32 = 0; // completed per-AP rendezvous count (for the tsc-sync test)
const warp_rounds: u32 = 1 << 20; // locked reads on the BSP: ~1 ms at GHz rates
const warp_spin_limit: u64 = 1 << 32; // bound every rendezvous wait so a lost core can't hang boot
// The warp check is bounded by TIME, not iterations: a warp tick is a locked
// read-modify-write on a cacheline two cores are fighting over — microseconds
// under real contention, not the nanosecond an uncontended count assumes (a
// 1<<20-round budget measured 2-21 SECONDS per core on a 16-core machine), and
// a PAUSE costs ~140 cycles on modern Intel, so an iteration-counted await
// mis-measures by two orders of magnitude too. ~5 ms of pairwise hammering per
// core is plenty to catch a lagging TSC (Linux's check_tsc_warp budget), and
// ~100 ms is a generous rendezvous window for a healthy core.
const warp_check_ns: u64 = 5_000_000; // per-AP pairwise check duration
const warp_await_ns: u64 = 100_000_000; // rendezvous wait before giving up
/// TSC ticks for `ns` nanoseconds (valid whenever the warp check runs: the TSC
/// is the clocksource, so tsc_hz is calibrated).
fn warpTicksFor(ns: u64) u64 {
return @intCast(@as(u128, ns) * tsc_hz / 1_000_000_000);
}
fn warpTick() void {
while (@cmpxchgWeak(u32, &warp_lock, 0, 1, .acquire, .monotonic) != null) asm volatile ("pause");
@@ -545,11 +559,11 @@ fn warpTick() void {
@atomicStore(u32, &warp_lock, 0, .release);
}
/// Spin (bounded) until `flag` is nonzero; false on timeout.
/// Spin (time-bounded) until `flag` is nonzero; false on timeout.
fn warpAwait(flag: *u32) bool {
var spins: u64 = 0;
while (@atomicLoad(u32, flag, .acquire) == 0) : (spins += 1) {
if (spins >= warp_spin_limit) return false;
const deadline = rdtsc() +% warpTicksFor(warp_await_ns);
while (@atomicLoad(u32, flag, .acquire) == 0) {
if (rdtsc() -% deadline < (1 << 62)) return false; // past the deadline
asm volatile ("pause");
}
return true;
@@ -568,8 +582,8 @@ pub fn checkWarpSource() void {
@atomicStore(u32, &warp_bsp_ready, 0, .release);
return;
}
var i: u32 = 0;
while (i < warp_rounds) : (i += 1) warpTick();
const deadline = rdtsc() +% warpTicksFor(warp_check_ns);
while (rdtsc() -% deadline >= (1 << 62)) warpTick(); // until the time budget is spent
@atomicStore(u32, &warp_stop, 1, .release);
@atomicStore(u32, &warp_bsp_ready, 0, .release);
warp_checks += 1;
@@ -583,9 +597,10 @@ pub fn checkWarpTarget() void {
if (clock_source != .tsc) return;
if (!warpAwait(&warp_bsp_ready)) return;
@atomicStore(u32, &warp_ap_ready, 1, .release);
var spins: u64 = 0;
while (@atomicLoad(u32, &warp_stop, .acquire) == 0) : (spins += 1) {
if (spins >= warp_spin_limit) return;
// The BSP owns the budget; this bound only protects against a lost BSP.
const deadline = rdtsc() +% warpTicksFor(2 * warp_check_ns + warp_await_ns);
while (@atomicLoad(u32, &warp_stop, .acquire) == 0) {
if (rdtsc() -% deadline < (1 << 62)) return; // past the deadline
warpTick();
}
}
+6 -10
View File
@@ -138,14 +138,16 @@ pub const Console = struct {
}
}
/// Shift the visible text up one glyph row and clear the freed bottom row,
/// leaving the cursor on that now-blank last line.
/// The screen is full: start a fresh page at the top. NEVER scroll by
/// copying pixel rows — that READS the framebuffer, and VRAM reads are
/// uncached-slow on real hardware (measured: 16-core bring-up took ~90 s
/// purely from boot lines each paying a whole-screen scroll copy). A page
/// clear is writes only, and only once per screenful.
fn scroll(self: *Console) void {
const visible = self.rows * glyph_h;
var y: u32 = 0;
while (y + glyph_h < visible) : (y += 1) self.copyRow(y, y + glyph_h);
while (y < visible) : (y += 1) self.fillRow(y, self.bg);
self.row = self.rows - 1;
self.row = 0;
}
inline fn rowPtr(self: *Console, y: u32) [*]volatile u32 {
@@ -163,10 +165,4 @@ pub const Console = struct {
while (x < self.fb.width) : (x += 1) row[x] = color;
}
fn copyRow(self: *Console, destination_y: u32, source_y: u32) void {
const destination = self.rowPtr(destination_y);
const source = self.rowPtr(source_y);
var x: u32 = 0;
while (x < self.fb.width) : (x += 1) destination[x] = source[x];
}
};
+7
View File
@@ -162,6 +162,13 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// uncached crawl). Routine boot output goes only to the log; this console now exists for
// early-boot and fatal (`fatal`/panic) output, until the display service takes over.
console.init(fb);
// The console joins the log sinks: the boot transcript — kernel AND
// userspace lines, each timestamped by the renderer — shows on screen
// until the display service claims the framebuffer (which flips the
// console's `suppressed` and silences this sink). On a machine with no
// serial this is the only live view of the boot, and a slow boot becomes
// diagnosable by eye: the timeline is right there.
if (console.present()) log.addSink(console.write);
log.write(if (console.present())
"/system/kernel: framebuffer ready (early-boot + fatal fallback; the display service drives it in normal operation)\n"
else
+12 -4
View File
@@ -126,7 +126,7 @@ fn appendLocked(pid: u32, name: []const u8, level: abi.KlogLevel, now: u64, byte
const line_complete = newline != null or level != .raw;
if (line.len != 0 or line_complete)
_ = ring.append(pid, name, level, now, line, line.len > abi.klog_maximum_message);
render(pid, name, level, line, line_complete);
render(pid, name, level, now, line, line_complete);
rest = if (newline) |i| rest[i + 1 ..] else rest[rest.len..];
}
}
@@ -136,7 +136,7 @@ fn appendLocked(pid: u32, name: []const u8, level: abi.KlogLevel, now: u64, byte
/// own "name: " prefixes until the std.log migration). Leveled (std.log)
/// records get a kernel-rendered "<name>: " prefix at line start — err/warn/
/// debug also get their level spelled out.
fn render(pid: u32, name: []const u8, level: abi.KlogLevel, line: []const u8, line_complete: bool) void {
fn render(pid: u32, name: []const u8, level: abi.KlogLevel, now: u64, line: []const u8, line_complete: bool) void {
if (sink_count == 0) return;
if (line.len == 0 and !line_complete) return;
// Compose the whole rendered piece first and emit it in ONE sink call per
@@ -149,6 +149,14 @@ fn render(pid: u32, name: []const u8, level: abi.KlogLevel, line: []const u8, li
used += 1;
at_line_start = true;
}
if (at_line_start) {
// Every line starts with its boot-relative time: the live transcript
// (serial AND the on-screen boot console) is a readable timeline —
// which is how a slow real-hardware boot gets diagnosed by eye.
const seconds = now / 1_000_000_000;
const millis = (now / 1_000_000) % 1000;
used += (std.fmt.bufPrint(buffer[used..], "[{d:>4}.{d:0>3}] ", .{ seconds, millis }) catch buffer[used..used]).len;
}
if (at_line_start and level != .raw) {
used += place(buffer[used..], name);
used += place(buffer[used..], ": ");
@@ -169,8 +177,8 @@ fn render(pid: u32, name: []const u8, level: abi.KlogLevel, line: []const u8, li
open_line_pid = pid;
}
/// newline + name + ": warning: " + a full payload line + newline.
const render_buffer_size = 1 + abi.maximum_process_name + 11 + abi.klog_maximum_message + 1;
/// newline + timestamp + name + ": warning: " + a full payload line + newline.
const render_buffer_size = 1 + 16 + abi.maximum_process_name + 11 + abi.klog_maximum_message + 1;
fn place(destination: []u8, bytes: []const u8) usize {
const n = @min(destination.len, bytes.len);
+9 -1
View File
@@ -27,7 +27,15 @@ const build_options = @import("build_options");
/// kernel. Drivers are absent on purpose: the device manager owns those. (A
/// future init reads this from a manifest under /system/services instead of a
/// hardcoded list.)
const boot_services = [_][]const u8{
const boot_services = if (build_options.diagnose) [_][]const u8{
// The diagnose boot: no display service, so the kernel's on-screen boot
// transcript is never suppressed — the timestamped timeline (USB bring-up,
// storage, logger) stays readable on real hardware with no serial.
"/system/services/input",
"/system/services/device-manager",
"/system/services/fat",
"/system/services/logger",
} else [_][]const u8{
"/system/services/input",
"/system/services/device-manager",
"/system/services/fat",
+6 -1
View File
@@ -112,9 +112,14 @@ fn onNotification(badge: u64) void {
}
fn onTerminate() void {
// Final drain: everything still in the ring, then close (= flush) all files.
// The completeness receipt FIRST: this record enters the ring before the
// final drain, so the drain carries it into logger.log — a directory whose
// logger.log ends with this marker is complete through shutdown; one that
// doesn't was cut early and may be missing tails.
_ = system.write("logger: shutting down; final flush\n");
drain();
closeAll();
// Serial-only epilogue (after the drain, so it reaches no file — by design).
var line: [96]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "logger: flushed through sequence {d}\n", .{next_expected_sequence}) catch return);
}
+53
View File
@@ -0,0 +1,53 @@
#!/usr/bin/env python3
"""Pack the bundled binaries into the boot capsule (boot/system.img) — the
single file the EFI loader reads in one sequential pass, which is the only
shape firmware file I/O is fast at (a per-file tree walk measured minutes on
real firmware). The format is the v2 initial_ramdisk (system/initial-ramdisk.zig):
entries named by full FHS path, so the running system is identical whether the
loader read the capsule or walked the tree.
Usage: pack-system-image.py <out.img> [<path> <file>]...
Layout (little-endian): Header{magic "DNR2", count}, Entry{name[64], offset, len}*N, blobs.
"""
import struct
import sys
MAGIC = 0x32524E44 # "DNR2"
HEADER = struct.Struct("<II")
ENTRY = struct.Struct("<64sQQ")
def main() -> int:
out_path = sys.argv[1]
rest = sys.argv[2:]
if len(rest) % 2 != 0:
sys.stderr.write("usage: pack-system-image.py <out.img> [<path> <file>]...\n")
return 2
items = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)]
table_end = HEADER.size + len(items) * ENTRY.size
entries = b""
blobs = []
offset = table_end
for path, source in items:
name = path if path.startswith("/") else "/" + path
encoded = name.encode("ascii")
if len(encoded) > 63:
sys.stderr.write(f"pack-system-image: path too long (>63): {name}\n")
return 2
with open(source, "rb") as f:
data = f.read()
entries += ENTRY.pack(encoded, offset, len(data))
blobs.append(data)
offset += len(data)
with open(out_path, "wb") as f:
f.write(HEADER.pack(MAGIC, len(items)))
f.write(entries)
for blob in blobs:
f.write(blob)
return 0
if __name__ == "__main__":
sys.exit(main())