8 Commits
Author SHA1 Message Date
Daniel Samson 0b25cd2c94 docs: multi-volume is built — storage architecture + rationale (S3)
Flip the storage docs from "one FAT volume today" to the built state:
the manager adopts every device, probes each device's whole partition
table, and spawns one range-confined FAT per volume — several volumes
across several devices, or several partitions on one device's channel.
The boot volume is identified by content; a filesystem installs the
/system rewrites only when it resolves /system/configuration on its own
media. No filesystem binds a shared service name any more — clients route
through the kernel mount table.

Correct two now-stale claims in the rationale, in the honest direction:
"multi-volume providers are reserved, not implemented" becomes built
(multi-namespace-per-provider is the untried NVMe case); and the cloned-
duplicate "second mounts suffixed" was never built — S3 mounts both,
they collide on the shared content id-path (last wins), each boot claim
logged, and distinguishing them is S4 arbitration.
2026-08-10 02:35:43 +01:00
Daniel Samson d59279422e test: partitioned-image tool + shared-channel multi-volume proof (S3)
The two-volumes case proves multiple DEVICES; this proves multiple
volumes on ONE device. make-partitioned-image.py lays several FAT32
partitions (each from make-fat-image) behind a classic MBR; the new
partitioned-volume case attaches one such disk (two partitions,
da7a0001 at lba 2048, da7a0002 at lba 83968) as a single usb-storage
device.

partition.allVolumes walks the table and the manager spawns a confined
fat per partition on the SAME block channel, each clamped to its own LBA
range by usb-storage's per-badge range table — so partition B's fat
cannot read partition A's blocks. The case asserts two mount lines at
two distinct non-zero base_lbas on one device; against the pre-uncap
allVolumes (S3 step 2, capped to one partition) only da7a0001 mounts.

No image is committed — the disk is generated per run. Full suite
130/130 (128 + two-volumes + partitioned-volume).
2026-08-10 02:35:32 +01:00
Daniel Samson bf9f8560c6 fat/harness: filesystems coexist without the shared vfs name; two-volume proof (S3)
A second usb-storage device (a generated data volume, serial da7a0001,
an empty FAT with no /system) plugged in beside the boot volume: the
volume manager adopts both devices and spawns a confined fat per volume,
each mounted at its own content id-path.

The test surfaced a real coexistence bug. Every filesystem bound the
single "vfs" contract name under /protocol; the second volume's fat lost
the race, service.run refused-and-exited on the held name, and that
volume never mounted. Clients don't reach filesystems by that name —
fs_resolve routes a path to its backing endpoint through the kernel
mount table by prefix — and nothing consumes "vfs", so the fix is to
bind no shared name: the harness's service_name now defaults to null.
This is the "this fades" the harness comment anticipated for the
volume-manager era; a filesystem's endpoint still serves as its mount
backend without a name.

fat logs "is a data volume" for the non-system branch so the test can
positively assert content-based detection. make-fat-image gains
--serial/--label (default unchanged) so a second image gets a distinct
id-path; the data image is generated per run, never committed. The case
fails against the pre-fix harness (the data volume's fat exits on the
refused bind) — toggle-demonstrated.

Full suite 129/129 (128 + two-volumes); the single-volume path is
unaffected by dropping the vestigial name bind.
2026-08-10 02:10:06 +01:00
Daniel Samson da7dcce64e kernel/vfs: raise the mount ceiling for N volumes; refuse (not drop) a full table (S3)
Multi-volume makes the mount table the bottleneck: each volume installs
one id-path mount and the system volume two FHS rewrites, so at the
volume manager's maximum_volumes (16) the old cap of 8 is far too low.
Raise maximum_mounts to 32 (headroom over the ~20-mount worst case) and
declare its bounds block; drop it from the bounds allowlist.

Fix a latent bug the higher pressure would expose: installMount silently
dropped a mount when the table was full, and mountBackend returned true
anyway — a full table was reported as a successful mount. installMount
now returns whether it placed the mount, and mountBackend propagates a
false so the mounting filesystem's harness logs "could not mount
<prefix>". At-limit is now a refusal that is observed, not a silent
success. (The full-table path has no host unit test: vfs.zig's tests are
not wired into the host aggregate — its import graph reaches the
freestanding kernel — so the correction rests on the propagated return
and the truthful bounds block.)
2026-08-10 01:52:47 +01:00
Daniel Samson 9750db14da fat: install the FHS boot rewrites only on the system volume (S3)
A volume backs /system/configuration and /system/logs only when it
actually carries the /system tree — decided by content (resolve
/system/configuration on its own media at mount), not by spawn order.
The boot volume takes the branch and installs the two rewrites; a data
volume resolves null, mounts only at its id-path, and never shadows the
running system's config or logs with a dead mount.

This retires the "resolve-at-bring-up quirk" the single-volume step
deferred around: that diagnosis was wrong. A boot probe confirmed
resolve() works the instant mount() returns — /system, /system/
configuration, /system/kernel, /system/services all resolve at bring-up
(mount reads LBA 0 through the same block path, so a directory read
cannot fail where the boot-sector read succeeded). No deferral needed.

Behavior-preserving on the single boot volume (it carries the system
tree, so it still installs all three mounts): suite stays 128/128.
2026-08-10 01:37:32 +01:00
Daniel Samson b2a5a0a3c6 volume-manager: adopt every device, a filesystem per partition (S3)
Lift the one-device/one-volume cap. bringUpVolume now probes the whole
partition table (allVolumes uncapped) and spawns a confined filesystem
per volume; pollTick loops it to adopt every present, not-yet-adopted
device each tick.

The subtlety is adopt-once-and-keep: a device is recorded in the table
the first time it is seen and kept until it leaves the tree, even when it
carries no servable volume or its geometry cannot be read. Dropping an
unservable device would make openAnyStorage hand back the same one every
tick and starve the devices behind it; keeping it lets the scan advance
past it. A genuine removal frees the slot; a re-insert (fresh device id)
is probed anew.

The boot image is a single bare-FAT volume, so the full suite is
unchanged at 128/128.
2026-08-10 01:26:47 +01:00
Daniel Samson 7efe7b72d8 volume-manager: N-volume device+volume tables, behavior-preserving (S3)
Replace the single `var volume: ?Volume` and file-global supervision
state with two fixed tables: devices[maximum_devices] owning each adopted
block channel once, and volumes[maximum_volumes] each carrying its own
identity, id, mount prefix, and supervision fields (restarts, spawn_ns,
failed, restart_pending, restart_due_ns). A monotonic next_volume_id
never reuses ids, so a stale hello can't address the wrong child.

Lookups (deviceById, volumeById, volumeByPid, firstUsedVolume) and
claims (claimDevice, claimVolumeIndex) replace the ad-hoc singletons.
pollTick reconciles devices first (removeDevice drops their volumes),
then per-volume restarts, then idle bring-up.

This step stays one-device/one-volume on purpose: bringUpVolume adopts
the first device and caps allVolumes to a single partition, so behavior
is identical and the full suite stays 128/128. Uncapping and adopt-all
land next.
2026-08-10 01:13:52 +01:00
Daniel Samson d4b544d66b volume-manager: partition.allVolumes — every partition, not just the first (S3)
The multi-volume enabler. allVolumes(reader, device_blocks, out) appends every
volume on the device to the caller's buffer and returns the count: GPT
enumerates all valid entries (gptFirstVolume becomes gptAllVolumes), the MBR walk
collects all fitting partitions, and a bare FAT is the single whole-device volume
— each with the same per-entry overflow-safe range validation (the confinement
invariant the driver's clamp rests on) and fatIdentity-over-disk-signature
preference. firstVolume is now the one-element case of allVolumes, so the S1
behavior and its ten tests are unchanged. New host test: a two-partition MBR
yields two volumes with distinct identities (index 0 vs 1); it FAILS when
allVolumes is capped to one (the old firstVolume semantics), passes at 11/11.
2026-08-10 00:57:56 +01:00
11 changed files with 604 additions and 242 deletions
@@ -11,15 +11,21 @@
> identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:
> `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional
> override), a volume's mount path IS its content id (`/volumes/<id>`), with the
> label as display metadata a `volumes` query returns. **Still pending**: the
> `filesystem UUID` rung (needs a non-FAT engine), multi-volume (one FAT volume
> today; fat's boot rewrites are unconditional until S3 makes them
> content-conditional), the volume manager *consuming* `medium_changed` (removal
> is detected by device-presence polling; the event is published but only a
> card-reader medium change needs the subscription), and the remount-on-replug
> label as display metadata a `volumes` query returns. Multi-volume is **built**:
> the manager adopts every storage device, probes each device's whole partition
> table, and spawns one range-confined FAT per volume — several volumes across
> several devices, or several partitions sharing one device's channel — each at
> its own `/volumes/<id>` path with its own supervision. The boot volume is
> identified by **content** (a volume backs `/system/configuration` + `/system/logs`
> only when it resolves `/system/configuration` on its own media), so it works as
> any partition of any device. **Still pending**: the `filesystem UUID` rung and
> a second engine (exFAT, S4); the volume manager *consuming* `medium_changed`
> (removal is detected by device-presence polling; the event is published but only
> a card-reader medium change needs the subscription); the remount-on-replug
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller
> device, so it is bench-verified). A few markers below are left where a duty is
> still pending.
> device, so it is bench-verified); and arbitration when two volumes both resolve
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
> markers below are left where a duty is still pending.
## The model
@@ -121,12 +127,16 @@ its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there, plus
the two `/system` hierarchy rewrites it installs in place (unconditional this
increment; S3 makes them content-conditional across volumes). It no longer
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there. It
installs the two `/system` hierarchy rewrites (`/system/configuration`,
`/system/logs`) only when it is the boot volume — decided by **content**: it
resolves `/system/configuration` on its own media at mount, so a data volume
mounts at its id-path alone and never shadows the running system. It no longer
self-acquires a volume — the V3b flip made it receive its volume id and block
channel from the volume manager, consistent with "it never discovers devices"
above.
above. Because several volumes now serve at once, no filesystem binds a shared
service name; clients reach each through the kernel mount table (`fs_resolve`
routes by prefix to the backing endpoint).
**Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart
@@ -145,9 +145,13 @@ matrix-proven shape; genuinely open.
**The pressure points, honestly:**
1. **Multi-volume providers are reserved, not implemented.** The volume
manager flow assumes one provider, one volume; NVMe namespaces make
endpoint-per-volume real work with hardware demanding it.
1. **Multi-volume is built; multi-namespace-per-provider is untried.** The
volume manager adopts every device and spawns one range-confined FAT per
partition — several volumes across several devices, or several partitions
sharing one device's channel, both proven on USB. What is untried is a single
provider exposing several volumes as *namespaces* (NVMe): the endpoint and
per-badge range machinery generalizes, but no such driver exists yet to
exercise it.
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
one operation in flight, one bounce buffer — fine for a USB2 stick,
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
@@ -233,13 +237,15 @@ matrix-proven shape; genuinely open.
Consequences, each mechanical once identity keys the map: **moving a drive
to a different port changes nothing** — same identity, same mount point,
whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (a map lookup
once the map exists; today's single volume re-probes and remounts at the
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the
recorded identity of the volume carrying `/system/configuration`, findable
on any port; and **duplicate identity is a policy case, not a surprise** —
two cloned sticks at once: first keeps the mapped name, second mounts
suffixed and is logged loudly, never silently shadowed. Unknown identities
returned in a SATA dock; **replug remounts at the same path** (the id-path is
content-derived, so a volume returns to `/volumes/<id>` wherever it reappears;
remount-on-replug end-to-end is bench-verified, not QEMU-tested, because QEMU
can't re-present the boot-controller device); **the boot volume** is the volume
that resolves `/system/configuration` on its own media, findable on any port or
partition; and **duplicate identity is a known S4 gap** — two cloned sticks
share one content id, so today they collide on `/volumes/<id>` (the kernel
remount-replaces; the last wins) and each boot-volume claim is logged loudly.
Distinguishing them with a suffix is arbitration, deferred to S4. Unknown identities
mount under a derived name (sanitized label, else generated) at
`/volumes/<name>` — the hierarchy's documented home for attached media,
which stands: `/system` is what danos IS; attached media is what it isn't.
+8 -4
View File
@@ -61,10 +61,14 @@ pub fn Server(comptime Engine: type) type {
/// caller does the filesystem-specific bring-up (find the block
/// device, set up DMA, mount the engine) and returns a `Volume`.
bringUp: *const fn (endpoint: ipc.Handle) ?Volume,
/// The vfs contract name to bind. A filesystem serving one volume
/// binds "vfs" today; the volume-manager era hands each per-volume
/// process its own establishment and this fades.
service_name: ?[]const u8 = "vfs",
/// A contract name to bind under /protocol, or null to bind none. In
/// the volume-manager era every filesystem is a per-volume process and
/// clients reach it through the kernel mount table — fs_resolve routes
/// a path to its backing endpoint by prefix — so no filesystem binds a
/// shared name. Two volumes would collide on one: the second's bind is
/// refused and service.run would exit, so its volume never mounts. The
/// endpoint still serves as the mount backend without a name.
service_name: ?[]const u8 = null,
};
// --- the harness's own state, one set per instantiation ---------------
+25 -6
View File
@@ -55,7 +55,17 @@ fn tokenIndex(t: u64) u64 {
// --- the mount table ---------------------------------------------------------
pub const maximum_mounts = 8;
/// bound: prefixes mounted in the kernel VFS table at once
/// decided-by: ours
/// protects: the `mounts` table below
/// at-limit: refuse - installMount returns false and mountBackend propagates it;
/// the mounting filesystem's harness logs "could not mount <prefix>" and the
/// mount simply does not exist (no silent success). Budget: the initrd's
/// top-level dirs (/system, /test) plus one id-path mount per volume and the
/// system volume's two FHS rewrites — a few over the volume manager's
/// maximum_volumes (16); 32 leaves headroom.
/// observed-by: the harness "file-system: could not mount <prefix>" ring line
pub const maximum_mounts = 32;
const maximum_prefix = 64;
const maximum_rewrite = 32;
@@ -156,7 +166,10 @@ pub fn setInitialRamdisk(image: []const u8) void {
for (directories[0..directory_count], 0..) |*d, index| {
const parent = parentOf(d.slice());
d.parent = directoryIndex(parent) orelse index;
if (parent.len == 1) installMount(d.slice(), .kernel_initrd, null, "");
// Boot-time install of one mount per top-level initrd dir (/system, /test):
// provably few, far under maximum_mounts, so a full table here is
// impossible — but discard the result explicitly rather than assume it.
if (parent.len == 1) _ = installMount(d.slice(), .kernel_initrd, null, "");
}
}
@@ -167,7 +180,12 @@ fn directoryIndex(path: []const u8) ?usize {
return null;
}
fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, rewrite: []const u8) void {
/// Install (or remount-replace) a prefix. Returns false when the table is full
/// and no slot could be claimed — the caller must surface that, never report a
/// dropped mount as success. A remount of an already-mounted prefix reuses its
/// slot and always succeeds; a /protocol remount is refused-as-noop (returns
/// true: the first mount stands, nothing is dropped).
fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, rewrite: []const u8) bool {
// Remount replaces: a restarted backend re-mounts its prefix.
var slot: ?*Mount = null;
for (&mounts) |*m| {
@@ -176,19 +194,20 @@ fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, re
// restarted FAT retakes /volumes/usb; letting it retake /protocol
// would hand the whole naming layer to whoever asked second.
// First mount wins, and init (PID 1) is always first.
if (std.mem.eql(u8, prefix, protocol_root)) return;
if (std.mem.eql(u8, prefix, protocol_root)) return true;
if (m.backend) |old| ipc.dropRef(old);
slot = m;
break;
}
if (slot == null and !m.used) slot = m;
}
const m = slot orelse return;
const m = slot orelse return false;
m.* = .{ .used = true, .kind = kind, .backend = backend };
@memcpy(m.prefix[0..prefix.len], prefix);
m.prefix_len = prefix.len;
@memcpy(m.rewrite[0..rewrite.len], rewrite);
m.rewrite_len = rewrite.len;
return true;
}
// --- resolve -----------------------------------------------------------------
@@ -406,7 +425,7 @@ pub fn mountBackend(prefix: []const u8, backend: *ipc.Endpoint, rewrite: []const
if (!isInitrdCarveOut(prefix)) return false;
}
}
installMount(prefix, .backend, backend, rewrite);
if (!installMount(prefix, .backend, backend, rewrite)) return false; // table full
for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefixSlice(), prefix)) m.owner = owner;
}
+16 -7
View File
@@ -173,16 +173,25 @@ fn fatBringUp(endpoint: ipc.Handle) ?Harness.Volume {
};
std.log.info("mounted FAT ({s}, {d} clusters, partition lba {d})", .{ @tagName(filesystem.geometry.fat_type), filesystem.geometry.cluster_count, filesystem.base_lba });
// The volume mounts at its id-path (argv[2]), plus the two FHS rewrites so
// hierarchy paths (the logger's /system/logs) stay decoupled from which
// volume backs them. This single-volume increment's one volume IS the boot
// volume, so it installs both unconditionally; S3 (multi-volume) makes the
// rewrites content-conditional — installed only by whichever volume carries
// the system, decided by content, not order.
// Every volume mounts at its own id-path (argv[2]). The boot/system volume —
// the one carrying the /system tree — ADDITIONALLY installs the two FHS
// rewrites, so hierarchy paths (config reads, the logger's persistent
// /system/logs) stay decoupled from which volume backs them. Detection is by
// CONTENT, not spawn order: a volume is the system volume iff /system/
// configuration resolves on its own media. A data volume has no /system, so it
// mounts only at its id-path and never shadows the running system's config or
// logs with a dead mount.
mount_specs[0] = .{ .prefix = volume_mount_prefix };
var mount_count: usize = 1;
if (filesystem.resolve("/system/configuration") != null) {
std.log.info("volume {d} carries the system tree; backing /system/configuration and /system/logs", .{my_volume_id});
mount_specs[1] = .{ .prefix = "/system/configuration", .rewrite = "/system/configuration" };
mount_specs[2] = .{ .prefix = "/system/logs", .rewrite = "/system/logs" };
return .{ .engine = &filesystem, .mounts = mount_specs[0..3], .flush = flushIfDirty };
mount_count = 3;
} else {
std.log.info("volume {d} is a data volume; mounted at {s}", .{ my_volume_id, volume_mount_prefix });
}
return .{ .engine = &filesystem, .mounts = mount_specs[0..mount_count], .flush = flushIfDirty };
}
pub fn main(init: process.Init) void {
+80 -28
View File
@@ -172,39 +172,41 @@ fn setLabelFromUtf16(id: *Identity, name_bytes: []const u8) void {
id.label_len = @intCast(out);
}
/// The first GPT volume, or null if LBA 1 is not a valid GPT header or no entry
/// validates. The header CRC-32 and the per-entry overflow-safe range check are
/// the confinement-safety guards the driver's clamp rests on — the invariant
/// firstVolume documents for MBR, extended to untrusted GPT metadata. The
/// entry-array CRC is deferred (correctness-only; the range check carries safety).
fn gptFirstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
/// Append every valid GPT volume to `out` (up to `out.len`), returning the count
/// (0 if LBA 1 is not a valid GPT header). The header CRC-32 and the per-entry
/// overflow-safe range check are the confinement-safety guards the driver's clamp
/// rests on — the invariant documented for MBR, extended to untrusted GPT
/// metadata. The entry-array CRC is deferred (correctness-only; the range check
/// carries safety).
fn gptAllVolumes(reader: SectorReader, device_blocks: u64, out: []Volume) usize {
var header: [sector_bytes]u8 = undefined;
if (!reader.read(1, &header)) return null;
if (!std.mem.eql(u8, header[0..8], gpt_signature)) return null;
if (!reader.read(1, &header)) return 0;
if (!std.mem.eql(u8, header[0..8], gpt_signature)) return 0;
const header_size = std.mem.readInt(u32, header[12..16], .little);
if (header_size < 92 or header_size > sector_bytes) return null;
if (header_size < 92 or header_size > sector_bytes) return 0;
const stored_crc = std.mem.readInt(u32, header[16..20], .little);
var check: [sector_bytes]u8 = undefined;
@memcpy(check[0..header_size], header[0..header_size]);
@memset(check[16..20], 0);
if (crc32(check[0..header_size]) != stored_crc) return null;
if (crc32(check[0..header_size]) != stored_crc) return 0;
const entry_lba = std.mem.readInt(u64, header[72..80], .little);
const num_entries = std.mem.readInt(u32, header[80..84], .little);
const entry_size = std.mem.readInt(u32, header[84..88], .little);
if (entry_size != 128 and entry_size != 256 and entry_size != 512) return null;
if (entry_lba == 0 or entry_lba >= device_blocks) return null;
if (entry_size != 128 and entry_size != 256 and entry_size != 512) return 0;
if (entry_lba == 0 or entry_lba >= device_blocks) return 0;
const scan = @min(num_entries, gpt_entry_scan_maximum);
var sector_buf: [sector_bytes]u8 = undefined;
var loaded: u64 = std.math.maxInt(u64);
var count: usize = 0;
var i: u32 = 0;
while (i < scan) : (i += 1) {
while (i < scan and count < out.len) : (i += 1) {
const abs = @as(u64, i) * entry_size;
const lba = entry_lba + abs / sector_bytes;
const off = @as(usize, @intCast(abs % sector_bytes));
if (lba != loaded) {
if (!reader.read(lba, &sector_buf)) return null;
if (!reader.read(lba, &sector_buf)) break; // return what we have
loaded = lba;
}
const entry = sector_buf[off..][0..128]; // the fields we read live in the first 128 bytes
@@ -224,9 +226,10 @@ fn gptFirstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
if (start == 0 or end < start or end >= device_blocks) continue;
var id = Identity{ .rung = .gpt_guid, .key = std.mem.readInt(u128, entry[16..32], .little) };
setLabelFromUtf16(&id, entry[56..128]);
return .{ .base_lba = start, .block_count = end - start + 1, .identity = id };
out[count] = .{ .base_lba = start, .block_count = end - start + 1, .identity = id };
count += 1;
}
return null;
return count;
}
/// Trim trailing spaces (FAT labels are space-padded) and copy into the display
@@ -259,18 +262,22 @@ fn fatIdentity(reader: SectorReader, start_lba: u64) ?Identity {
return id;
}
/// The first volume on the device `reader` addresses, whose whole-device size is
/// `device_blocks`, or null if none is found. A GPT disk (protective MBR) is
/// handled by GPT, authoritatively — its null is final. Otherwise an MBR with a
/// non-empty entry yields that partition's [start, size); otherwise a boot
/// signature with no partitions is treated as a bare FAT spanning the device.
pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
/// Append every volume on the device `reader` addresses, whose whole-device size
/// is `device_blocks`, to `out` (up to `out.len`), returning the count. A GPT
/// disk (protective MBR) is enumerated by GPT, authoritatively — a zero count is
/// final. Otherwise every fitting MBR entry is a volume; a boot signature with no
/// partition entries is a bare FAT spanning the whole device. Each volume's
/// [start, count) is validated overflow-safe (the confinement invariant the
/// driver's clamp rests on), and each prefers its FAT serial identity over the
/// disk signature.
pub fn allVolumes(reader: SectorReader, device_blocks: u64, out: []Volume) usize {
var block0: [sector_bytes]u8 = undefined;
if (!reader.read(0, &block0)) return null;
if (!hasBootSignature(&block0)) return null;
if (isProtectiveMbr(&block0)) return gptFirstVolume(reader, device_blocks);
if (!reader.read(0, &block0)) return 0;
if (!hasBootSignature(&block0)) return 0;
if (isProtectiveMbr(&block0)) return gptAllVolumes(reader, device_blocks, out);
var count: usize = 0;
var index: u8 = 0;
while (index < 4) : (index += 1) {
while (index < 4 and count < out.len) : (index += 1) {
const entry = block0[446 + @as(usize, index) * 16 ..][0..16];
const kind = entry[4];
const start = std.mem.readInt(u32, entry[8..12], .little);
@@ -283,10 +290,25 @@ pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
// device (usb-storage.zig resolveTransfer), which only holds because the
// range handed down is validated here. The subtraction cannot overflow.
if (start > device_blocks or device_blocks - start < size) continue;
return .{ .base_lba = start, .block_count = size, .identity = fatIdentity(reader, start) orelse mbrIdentity(&block0, index) };
out[count] = .{ .base_lba = start, .block_count = size, .identity = fatIdentity(reader, start) orelse mbrIdentity(&block0, index) };
count += 1;
}
if (count == 0 and out.len > 0) {
// No partition entries: a bare FAT spanning the device.
return .{ .base_lba = 0, .block_count = device_blocks, .identity = fatIdentity(reader, 0) orelse mbrIdentity(&block0, 0) };
out[0] = .{ .base_lba = 0, .block_count = device_blocks, .identity = fatIdentity(reader, 0) orelse mbrIdentity(&block0, 0) };
return 1;
}
return count;
}
/// firstVolume is allVolumes into a one-element buffer.
const one_volume_slot = 1;
/// The first volume on the device, or null — the single-volume case of
/// `allVolumes`, kept for callers that want just one.
pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
var one: [one_volume_slot]Volume = undefined;
return if (allVolumes(reader, device_blocks, &one) > 0) one[0] else null;
}
/// A read-only RAM disk over a byte slice of sectors, for the host tests.
@@ -507,3 +529,33 @@ test "GPT with 256-byte entries reads the non-128 offset arithmetic correctly" {
try std.testing.expectEqual(Rung.gpt_guid, v.identity.rung);
try std.testing.expectEqual(@as(u128, 0xF00D), v.identity.key);
}
// A fixture-sized volume buffer for the multi-volume tests, named so the bounds
// gate (which flags literal array lengths) stays quiet: a test input.
const test_volume_slots = 4;
test "allVolumes returns every fitting MBR partition with distinct identities" {
var block0 = [_]u8{0} ** 512;
block0[510] = 0x55;
block0[511] = 0xAA;
std.mem.writeInt(u32, block0[440..444], 0xDEADBEEF, .little);
// partition 0: start 2048, size 1000
block0[446 + 4] = 0x0c;
std.mem.writeInt(u32, block0[446 + 8 ..][0..4], 2048, .little);
std.mem.writeInt(u32, block0[446 + 12 ..][0..4], 1000, .little);
// partition 1: start 4096, size 2000
block0[462 + 4] = 0x0c;
std.mem.writeInt(u32, block0[462 + 8 ..][0..4], 4096, .little);
std.mem.writeInt(u32, block0[462 + 12 ..][0..4], 2000, .little);
const disk = RamDisk{ .sectors = &block0 };
var vols: [test_volume_slots]Volume = undefined;
const n = allVolumes(disk.reader(), 200000, &vols);
try std.testing.expectEqual(@as(usize, 2), n); // both partitions, not just the first
try std.testing.expectEqual(@as(u64, 2048), vols[0].base_lba);
try std.testing.expectEqual(@as(u64, 4096), vols[1].base_lba);
// distinct rung-4 identities (no FAT VBR at those LBAs): index 0 vs 1.
try std.testing.expectEqual((@as(u128, 0xDEADBEEF) << 8) | 0, vols[0].identity.key);
try std.testing.expectEqual((@as(u128, 0xDEADBEEF) << 8) | 1, vols[1].identity.key);
// firstVolume (the 1-buffer case) still returns just the first.
try std.testing.expectEqual(@as(u64, 2048), firstVolume(disk.reader(), 200000).?.base_lba);
}
+242 -156
View File
@@ -8,10 +8,12 @@
//! supervises the filesystems it spawns, exactly as the device manager
//! supervises drivers.
//!
//! This increment (V3b) is the flip: the FAT service stops acquiring its own
//! volume and is spawned here instead, confined to its partition, and handed
//! its channel over the volume-manager protocol. Single volume for now; the
//! mount map (volumes.csv) and multi-volume land next.
//! The manager holds a table of adopted storage DEVICES and a table of the
//! VOLUMES on them: it adopts every storage device the device-manager tree
//! carries, probes each one's whole partition table, and spawns one filesystem
//! process per volume — each confined to its partition's badge-scoped block
//! range, each supervised with its own budget. A device leaving the tree takes
//! its volumes with it.
const std = @import("std");
const channel = @import("channel");
@@ -35,22 +37,37 @@ const Serve = volume_manager_protocol.Protocol.Provider(void);
const Invocation = envelope.Invocation;
const Answer = envelope.Answer;
/// The single volume this increment handles: its provider channel, its block
/// sub-range, its identity, the id it is addressed by, and the filesystem
/// process serving it (0 until spawned; reset on death for respawn).
const Volume = struct {
storage: block.Device,
storage_device_id: u64, // the device-manager id this volume's provider serves
base_lba: u64,
block_count: u64,
identity: partition.Identity,
id: u64,
binary: []const u8, // the service binary, from filesystems.csv by signature
mount_prefix: []const u8, // the volume-root mount path (its id-path, or a volumes.csv override)
filesystem_pid: u32 = 0,
/// One adopted storage device: the block channel to its provider (opened once and
/// shared — refcounted per confined filesystem via the hello reply) and the
/// device-manager id it serves. A device leaving the tree takes its volumes.
const StorageDevice = struct {
used: bool = false,
device_id: u64 = 0,
channel: block.Device = undefined,
};
const volume_id: u64 = 1;
/// One volume: which device serves it, its block sub-range, its content
/// identity, the id it is addressed by, the service binary + mount path it was
/// spawned with, the filesystem process serving it, and its own supervision
/// budget (so one volume's crash loop never touches another's).
const Volume = struct {
used: bool = false,
device_id: u64 = 0,
base_lba: u64 = 0,
block_count: u64 = 0,
identity: partition.Identity = .{ .rung = .anonymous },
id: u64 = 0,
binary: []const u8 = "",
mount_prefix: []const u8 = "",
filesystem_pid: u32 = 0,
// Per-volume supervision, mirroring the device manager's: a clean exit is not
// restarted, a fault restarts with backoff, a fast crash loop gives up.
restarts: u32 = 0,
spawn_ns: u64 = 0,
failed: bool = false,
restart_pending: bool = false,
restart_due_ns: u64 = 0,
};
// The mount map, read from configuration at boot (the policy home, storage-
// architecture.md): filesystems.csv (content signature -> service binary) and
@@ -80,47 +97,77 @@ var filesystem_rules: [maximum_filesystem_rules]filesystem_map.Rule = undefined;
var filesystem_rule_count: usize = 0;
var volume_rules: [maximum_volume_rules]volume_map.Override = undefined;
var volume_rule_count: usize = 0;
/// The composed default mount path (/volumes/<id>) for the current volume; a
/// volumes.csv override is used in place and needs no buffer (it is already a
/// slice into volumes_source). One buffer suffices while the manager serves one
/// volume (multi-volume gives each its own in S3).
/// bound: bytes of a composed /volumes/<id> mount path
/// decided-by: ours
/// protects: the mount_prefix_buf below
/// protects: the per-volume mount_prefix buffers below
/// at-limit: truncate - bufPrint fails; the volume mounts at a fallback path (logged)
/// observed-by: the fallback path in the log
const mount_path_maximum = 64;
var mount_prefix_buf: [mount_path_maximum]u8 = undefined;
/// bound: volumes the manager serves at once
/// decided-by: ours
/// protects: the volumes table and its per-volume mount-path buffers
/// at-limit: truncate - a further partition is left unserved and logged (real
/// machines carry a handful of volumes, far under this)
/// observed-by: the "volume table full" log line
const maximum_volumes = 16;
/// bound: storage devices the manager adopts at once
/// decided-by: ours
/// protects: the devices table
/// at-limit: truncate - a further device is left unadopted and logged
/// observed-by: the "device table full" log line
const maximum_devices = 8;
var devices = [_]StorageDevice{.{}} ** maximum_devices;
var volumes = [_]Volume{.{}} ** maximum_volumes;
/// Each volume's composed default mount path lives in its slot's buffer; a
/// volumes.csv override is used in place (a slice into volumes_source, no buffer).
var mount_prefix_bufs: [maximum_volumes][mount_path_maximum]u8 = undefined;
var next_volume_id: u64 = 1; // monotonic — never reused, so a stale id can't address the wrong child
var service_endpoint: ipc.Handle = 0;
var manager_handle: ?ipc.Handle = null;
var bounce: memory.DmaRegion = undefined;
var bounce_ready = false;
/// The currently-mounted volume, or null while no storage is present. The whole
/// removal lifecycle is this field going null and back: the poll sees the
/// storage provider leave the device tree (a pulled stick), kills the filesystem
/// and clears this; when it returns, the poll re-acquires and re-mounts.
var volume: ?Volume = null;
var logged_no_volume = false;
/// How often the poll checks whether the storage provider is present. Fast
/// enough that an unplug unmounts promptly; the poll is a bare device-manager
/// enumerate, no channel work, so it is cheap to run continuously.
/// How often the poll checks device presence and fires due restarts. Fast enough
/// that an unplug unmounts promptly; the poll is a bare device-manager enumerate,
/// no channel work, so it is cheap to run continuously.
const poll_interval_ms = 500;
// Filesystem supervision, mirroring the device manager's (device-manager.zig):
// a clean exit is not restarted, a fault restarts with backoff, and a fast
// crash loop gives up rather than spinning. Without this a faulting filesystem
// respawns in a zero-delay loop.
// Filesystem supervision, mirroring the device manager's (device-manager.zig).
const fast_death_ns: u64 = 2_000_000_000;
const crash_loop_cap: u32 = 3;
const backoff_base_ms: u64 = 300;
var fs_restarts: u32 = 0;
var fs_spawn_ns: u64 = 0;
var fs_failed = false;
/// A fat restart is due at `restart_due_ns`; the poll loop performs it once the
/// backoff has elapsed (one timer, folded into the poll — no second timer).
var restart_pending = false;
var restart_due_ns: u64 = 0;
/// bytes to format a u64 volume id as decimal (20 digits fit)
const id_decimal_bytes = 24;
// --- table lookups -----------------------------------------------------------
fn deviceById(id: u64) ?*StorageDevice {
for (&devices) |*d| if (d.used and d.device_id == id) return d;
return null;
}
fn claimDevice() ?*StorageDevice {
for (&devices) |*d| if (!d.used) return d;
return null;
}
fn volumeById(id: u64) ?*Volume {
for (&volumes) |*v| if (v.used and v.id == id) return v;
return null;
}
fn volumeByPid(pid: u32) ?*Volume {
for (&volumes) |*v| if (v.used and v.filesystem_pid == pid) return v;
return null;
}
fn firstUsedVolume() ?*Volume {
for (&volumes) |*v| if (v.used) return v;
return null;
}
fn claimVolumeIndex() ?usize {
for (&volumes, 0..) |*v, i| if (!v.used) return i;
return null;
}
// --- device-manager plumbing -------------------------------------------------
fn deviceManager() ?ipc.Handle {
if (manager_handle) |h| return h;
@@ -131,13 +178,12 @@ fn deviceManager() ?ipc.Handle {
const OpenedStorage = struct { device_id: u64, device: block.Device };
/// The first mass-storage provider whose block channel actually opens, with its
/// device id. A device-manager tree can carry more than one entry of the
/// mass-storage identity — a phantom that no driver is bound to answers a
/// consumer hello with NO channel — so this tries each and takes the first that
/// yields a channel, exactly as a filesystem's own acquisition loop does.
/// Called only when there is no volume (an insertion), so the hellos it makes
/// are not per-poll churn.
/// The first mass-storage provider whose block channel opens and is NOT already
/// adopted, with its device id. A device-manager tree can carry more than one
/// entry of the mass-storage identity — a phantom that no driver is bound to
/// answers a consumer hello with NO channel — so this tries each and takes the
/// first that yields a channel. Skips already-adopted devices so a re-poll does
/// not re-open a device it already serves.
fn openAnyStorage() ?OpenedStorage {
const manager = deviceManager() orelse return null;
const Entry = device_manager_protocol.ChildEntry;
@@ -157,6 +203,7 @@ fn openAnyStorage() ?OpenedStorage {
const entry = std.mem.bytesToValue(Entry, tail[index * @sizeOf(Entry) ..][0..@sizeOf(Entry)]);
if (entry.device_id == device_manager_protocol.no_device) continue;
if ((entry.identity >> 16) & 0xff != 0x08 or (entry.identity >> 8) & 0xff != 0x06) continue;
if (deviceById(entry.device_id) != null) continue; // already adopted
const exchanged = driver.helloOn(manager, .consumer, entry.device_id, null, true) orelse continue;
const provider = exchanged.channel orelse continue; // a phantom / not-yet-bound entry
return .{ .device_id = entry.device_id, .device = .{ .endpoint = provider } };
@@ -167,7 +214,7 @@ fn openAnyStorage() ?OpenedStorage {
/// Whether `device_id` is still in the device-manager tree — a bare enumerate,
/// no consumer-hello, so it is cheap to call every poll. This is how removal is
/// detected: the specific device the mounted volume sits on disappears.
/// detected: the specific device a mounted volume sits on disappears.
fn isDevicePresent(device_id: u64) bool {
const manager = deviceManager() orelse return false;
const Entry = device_manager_protocol.ChildEntry;
@@ -191,59 +238,88 @@ fn isDevicePresent(device_id: u64) bool {
}
}
/// Spawn the filesystem for `v`, confine it to the volume's range, and record
/// its pid. The confinement is defined for the fresh pid BEFORE the filesystem
/// runs, so its first read is already bounded; the volume manager is the
/// confinement controller (it defines the first range on the device).
// --- lifecycle ---------------------------------------------------------------
/// Spawn the filesystem for `v`, confine it to the volume's range on its device's
/// channel, and record its pid. The confinement is defined for the fresh pid
/// BEFORE the filesystem runs, so its first read is already bounded; the volume
/// manager is the confinement controller (it defines the first range on the
/// device).
fn spawnFilesystem(v: *Volume) void {
if (fs_failed) return;
const pid = process.spawnSupervised(v.binary, &.{ "1", v.mount_prefix }, service_endpoint) orelse {
if (v.failed) return;
const dev = deviceById(v.device_id) orelse return; // its device left — poll will clean up
var id_str_buf: [id_decimal_bytes]u8 = undefined;
const id_str = std.fmt.bufPrint(&id_str_buf, "{d}", .{v.id}) catch "1";
const pid = process.spawnSupervised(v.binary, &.{ id_str, v.mount_prefix }, service_endpoint) orelse {
_ = logging.write("volume-manager: could not spawn the filesystem; retrying\n");
armRestart();
armRestart(v);
return;
};
if (!v.storage.defineRange(pid, v.base_lba, v.block_count)) {
if (!dev.channel.defineRange(pid, v.base_lba, v.block_count)) {
_ = logging.write("volume-manager: could not confine the filesystem to its volume; retrying\n");
_ = process.kill(pid);
armRestart();
armRestart(v);
return;
}
v.filesystem_pid = pid;
fs_spawn_ns = time.clock();
v.spawn_ns = time.clock();
std.log.info("volume 0x{x} -> {s} (pid {d}), lba {d}, {d} blocks", .{ v.identity.key, v.binary, pid, v.base_lba, v.block_count });
}
/// Schedule a fat restart after backoff; the poll loop performs it once due.
fn armRestart() void {
const delay = if (fs_restarts == 0) backoff_base_ms else backoff_base_ms << @intCast(@min(fs_restarts - 1, 5));
restart_due_ns = time.clock() + delay * 1_000_000;
restart_pending = true;
/// Schedule a restart for `v` after backoff; the poll loop performs it once due.
fn armRestart(v: *Volume) void {
const delay = if (v.restarts == 0) backoff_base_ms else backoff_base_ms << @intCast(@min(v.restarts - 1, 5));
v.restart_due_ns = time.clock() + delay * 1_000_000;
v.restart_pending = true;
}
/// A storage provider just appeared: open its channel, read block 0, parse the
/// volume, and spawn its filesystem. On any failure the channel is closed (so a
/// present-but-unreadable device does not leak a handle every poll) and `volume`
/// stays null — the next poll retries. A fresh medium gets a fresh supervision
/// budget.
fn bringUpVolume() void {
/// Compose a volume's mount path (its id-path `/volumes/<id>`, or a volumes.csv
/// override) into its slot's buffer, and return the slice.
fn composeMountPrefix(slot: usize, identity: partition.Identity) []const u8 {
var id_buf: [volume_map.id_maximum]u8 = undefined;
const id = volume_map.idString(identity, &id_buf);
return volume_map.overrideFor(volume_rules[0..volume_rule_count], id) orelse
(std.fmt.bufPrint(&mount_prefix_bufs[slot], "/volumes/{s}", .{id}) catch "/volumes/unknown");
}
/// Adopt the next present, not-yet-adopted storage device: take its channel,
/// probe its whole partition table, and spawn a filesystem per volume it carries.
/// Returns true when it consumed a device (so the caller can loop to adopt every
/// present device in one tick), false when none remain or the device table is full.
///
/// A device is adopted exactly once and kept until it leaves the tree — even when
/// it carries no volume we can serve, or its geometry cannot be read. Keeping the
/// empty/unreadable device adopted (rather than dropping and re-probing) is what
/// lets openAnyStorage advance PAST it to the devices behind it; dropping it would
/// make openAnyStorage hand back the same unservable device every tick and starve
/// the rest. A genuine removal frees the slot (removeDevice); a re-insert gets a
/// fresh device id and is probed anew.
fn bringUpVolume() bool {
if (!bounce_ready) {
bounce = memory.dmaAlloc(512, memory.dma_coherent | memory.dma_shareable) orelse return;
bounce = memory.dmaAlloc(512, memory.dma_coherent | memory.dma_shareable) orelse return false;
bounce_ready = true;
}
const opened = openAnyStorage() orelse return;
const opened = openAnyStorage() orelse return false;
const dev = claimDevice() orelse {
_ = logging.write("volume-manager: device table full; a storage device is left unadopted\n");
_ = ipc.close(opened.device.endpoint);
return false;
};
dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device };
const device = opened.device;
// Attach the read buffer to THIS device (a no-op without an enforcing IOMMU).
// The handle is kept, not closed, so it can be re-attached to the next
// device after a replug.
// The handle is kept, not closed, so it can be re-attached after a replug. A
// failed attach or geometry read leaves the device adopted but empty — we just
// cannot read it, and the slot still watches it for removal.
if (bounce.handle) |handle| {
if (!device.attach(handle)) {
_ = ipc.close(device.endpoint);
return;
_ = logging.write("volume-manager: could not attach the read buffer to a storage device; no volume served\n");
return true;
}
}
const geometry = device.geometry() orelse {
_ = ipc.close(device.endpoint);
return;
_ = logging.write("volume-manager: could not read a storage device's geometry; no volume served\n");
return true;
};
const ProbeReader = struct {
device: block.Device,
@@ -257,77 +333,88 @@ fn bringUpVolume() void {
};
var probe = ProbeReader{ .device = device };
const reader = partition.SectorReader{ .context = &probe, .readFn = ProbeReader.readSector };
const found = partition.firstVolume(reader, geometry.block_count) orelse {
if (!logged_no_volume) {
_ = logging.write("volume-manager: storage present but no recognizable volume\n");
logged_no_volume = true;
var found: [maximum_volumes]partition.Volume = undefined;
const n = partition.allVolumes(reader, geometry.block_count, found[0..]);
if (n == 0) {
std.log.info("device {d} present but carries no recognizable volume", .{dev.device_id});
return true;
}
_ = ipc.close(device.endpoint);
return;
};
for (found[0..n]) |fv| {
// Pick the service binary from the volume's content signature. A signature
// no filesystems.csv row serves goes unserved (logged), like an unbound
// device — the manager does not guess.
const binary = filesystem_map.match(filesystem_rules[0..filesystem_rule_count], found.signature) orelse {
if (!logged_no_volume) {
const binary = filesystem_map.match(filesystem_rules[0..filesystem_rule_count], fv.signature) orelse {
_ = logging.write("volume-manager: no filesystem serves this volume's content; unserved\n");
logged_no_volume = true;
}
_ = ipc.close(device.endpoint);
return;
continue;
};
// The mount path is the volume's identity id (/volumes/<id>), or a
// volumes.csv override pinning it to a chosen path. The id is content-derived,
// so the path is stable and never a port or a label.
var id_buf: [volume_map.id_maximum]u8 = undefined;
const id = volume_map.idString(found.identity, &id_buf);
const mount_prefix = volume_map.overrideFor(volume_rules[0..volume_rule_count], id) orelse
(std.fmt.bufPrint(&mount_prefix_buf, "/volumes/{s}", .{id}) catch "/volumes/unknown");
logged_no_volume = false;
fs_restarts = 0;
fs_failed = false;
restart_pending = false;
volume = .{ .storage = device, .storage_device_id = opened.device_id, .base_lba = found.base_lba, .block_count = found.block_count, .identity = found.identity, .id = volume_id, .binary = binary, .mount_prefix = mount_prefix };
spawnFilesystem(&volume.?);
const slot = claimVolumeIndex() orelse {
_ = logging.write("volume-manager: volume table full; a volume is left unserved\n");
break;
};
volumes[slot] = .{
.used = true,
.device_id = dev.device_id,
.base_lba = fv.base_lba,
.block_count = fv.block_count,
.identity = fv.identity,
.id = next_volume_id,
.binary = binary,
.mount_prefix = composeMountPrefix(slot, fv.identity),
};
next_volume_id += 1;
spawnFilesystem(&volumes[slot]);
}
return true;
}
/// The storage provider left the device tree (a pulled stick): kill the
/// filesystem so its mounts are retired. Retirement is lazy, not an eager
/// death-time sweep — killing the process marks the filesystem's backend
/// endpoint dead, and the VFS router drops each mount that endpoint backed on
/// the next path resolution under it (that resolve frees the slot and returns
/// not_found). Then drop the now-dead channel and clear the volume; the next
/// poll that sees storage return re-mounts.
fn removeVolume() void {
const v = volume orelse return;
/// Close a device's channel and free its slot. No volumes are touched (the caller
/// ensures none remain, or there never were any).
fn dropDevice(dev: *StorageDevice) void {
_ = ipc.close(dev.channel.endpoint);
dev.* = .{};
}
/// Retire one volume: kill its filesystem so its mounts are retired. Retirement
/// is lazy, not an eager death-time sweep — killing the process marks the
/// filesystem's backend endpoint dead, and the VFS router drops each mount that
/// endpoint backed on the next path resolution under it (that resolve frees the
/// slot and returns not_found). Then free the volume slot.
fn removeVolumeState(v: *Volume) void {
std.log.info("storage for volume {d} removed; unmounting", .{v.id});
if (v.filesystem_pid != 0) _ = process.kill(v.filesystem_pid);
_ = ipc.close(v.storage.endpoint);
volume = null;
restart_pending = false;
fs_restarts = 0;
fs_failed = false;
v.* = .{};
}
/// One poll tick. Removal is checked FIRST and supersedes a pending restart: if
/// the device is gone there is nothing to restart fat onto, and respawning it
/// against the dead channel would just churn until the crash cap. Only once the
/// device is confirmed present does a due restart fire.
/// A storage device left the tree (a pulled stick): retire every volume it served
/// and drop its channel. One removal path, whether the device is pulled cleanly
/// or vanishes.
fn removeDevice(dev: *StorageDevice) void {
for (&volumes) |*v| {
if (v.used and v.device_id == dev.device_id) removeVolumeState(v);
}
dropDevice(dev);
}
/// One poll tick. Device removal is reconciled FIRST and supersedes a pending
/// restart: a volume whose device left is retired before its restart could fire,
/// so nothing respawns against a dead channel. Then due restarts fire for present
/// volumes; then, if no device is adopted, a present device is brought up.
fn pollTick() void {
if (volume) |v| {
// Serving: watch for the specific device leaving (a pulled stick).
if (!isDevicePresent(v.storage_device_id)) {
removeVolume();
return;
for (&devices) |*dev| {
if (dev.used and !isDevicePresent(dev.device_id)) removeDevice(dev);
}
if (restart_pending and time.clock() >= restart_due_ns) {
restart_pending = false;
spawnFilesystem(&volume.?);
for (&volumes) |*v| {
if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) {
v.restart_pending = false;
spawnFilesystem(v);
}
} else {
// Idle: try to bring a present storage device up.
bringUpVolume();
}
// Adopt every present, not-yet-adopted storage device. Each call consumes at
// most one device (openAnyStorage skips the adopted), so the loop terminates
// once none remain; the maximum_devices guard is insurance against a logic
// slip, never the normal exit.
var adopted: usize = 0;
while (adopted < maximum_devices and bringUpVolume()) : (adopted += 1) {}
}
/// A filesystem announces itself for the volume it was spawned to serve. Reply
@@ -335,26 +422,26 @@ fn pollTick() void {
/// badge) as the call's returned capability. No channel means the volume is not
/// ready — the filesystem retries.
fn onHello(_: void, invocation: Invocation(volume_manager_protocol.Hello), _: Answer(void)) isize {
const v = volume orelse return 0; // not probed yet — retryable, no cap
if (invocation.target != v.id) return 0; // unknown volume — retryable
const v = volumeById(invocation.target) orelse return 0; // not probed yet — retryable, no cap
if (invocation.sender != v.filesystem_pid) {
// Not the filesystem we spawned for this volume. Refuse: only the
// confined filesystem gets the channel.
// Not the filesystem we spawned for this volume. Refuse: only the confined
// filesystem gets the channel.
std.log.info("refused hello for volume {d} from process {d}", .{ invocation.target, invocation.sender });
return -envelope.EPERM;
}
service.replyWithCapability(v.storage.endpoint);
const dev = deviceById(v.device_id) orelse return 0; // its device left — retryable
service.replyWithCapability(dev.channel.endpoint);
std.log.info("handed volume {d} to pid {d}", .{ v.id, invocation.sender });
return 0;
}
/// Answer a `volumes` query with the mounted volume's descriptor — its id (its
/// mount path is /volumes/<id> unless overridden), its actual mount path, and
/// its display label. This is how a shell or file manager reads a volume's
/// friendly name: software keys on the id, a UI shows the label. An empty reply
/// Answer a `volumes` query with a mounted volume's descriptor — its id (its
/// mount path is /volumes/<id> unless overridden), its actual mount path, and its
/// display label. Software keys on the id; a UI shows the label. Returns the first
/// mounted volume for now; a full enumerate is a later refinement. Empty reply
/// means no volume is mounted.
fn onVolumes(_: void, _: Invocation(volume_manager_protocol.Volumes), answer: Answer(void)) isize {
const v = volume orelse return 0;
const v = firstUsedVolume() orelse return 0;
var id_buf: [volume_map.id_maximum]u8 = undefined;
const info = volume_manager_protocol.VolumeInfo{
.id = volume_map.idString(v.identity, &id_buf),
@@ -381,9 +468,9 @@ fn readConfig(path: []const u8, buf: []u8) usize {
defer file.close();
var used: usize = 0;
while (used < buf.len) {
const n = file.read(buf[used..]) orelse break;
if (n == 0) break;
used += n;
const nn = file.read(buf[used..]) orelse break;
if (nn == 0) break;
used += nn;
}
return used;
}
@@ -427,23 +514,22 @@ fn onNotification(badge: u64) void {
// reclaimed by the driver on the same death; the respawn confines afresh.
if (got.isChildExit()) {
const dead = got.childProcessId();
const v = &(volume orelse return);
if (v.filesystem_pid != dead) return;
const v = volumeByPid(dead) orelse return;
v.filesystem_pid = 0;
const reason = process.exitReason(dead) orelse .fault;
if (reason == .exited) {
std.log.info("filesystem for volume {d} exited cleanly; not restarting", .{v.id});
return;
}
const alive = time.clock() -| fs_spawn_ns;
fs_restarts = if (alive < fast_death_ns) fs_restarts + 1 else 1;
if (fs_restarts >= crash_loop_cap) {
fs_failed = true;
const alive = time.clock() -| v.spawn_ns;
v.restarts = if (alive < fast_death_ns) v.restarts + 1 else 1;
if (v.restarts >= crash_loop_cap) {
v.failed = true;
std.log.info("filesystem for volume {d} is failing repeatedly; giving up", .{v.id});
return;
}
std.log.info("filesystem for volume {d} died ({s}); restarting", .{ v.id, @tagName(reason) });
armRestart();
armRestart(v);
}
}
+65
View File
@@ -816,6 +816,47 @@ CASES = [
"expect": r"volume-manager: volume 0x0*12345678 -> \S+ \(pid \d+\), lba \d+, \d+ blocks"
r"[\s\S]*volume-manager: handed volume \d+ to pid \d+",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# S3 multi-volume: a SECOND usb-storage device (a generated data volume, serial
# da7a0001, an empty FAT with no /system) plugged in beside the boot volume.
# Proves the volume manager adopts BOTH devices and spawns a confined fat per
# volume, each mounted at its own CONTENT id-path (/volumes/fat-<serial>); and
# that boot-volume detection is by content — only the volume that carries
# /system backs /system/configuration, while the data volume mounts at its
# id-path alone. Against the pre-S3 one-device/one-volume manager the data
# volume never mounts, so the da7a0001 lookaheads fail (toggle-demonstrated by
# checking out the step-2 volume-manager.zig).
{"name": "two-volumes",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"data_volume": {"serial": "DA7A0001", "label": "DATAVOL", "size_mib": 64},
"expect": r"(?s)(?=.*volume-manager: volume 0x0*12345678 -> )"
r"(?=.*volume-manager: volume 0x0*da7a0001 -> )"
r"(?=.*fat: mounted /volumes/fat-12345678)"
r"(?=.*fat: mounted /volumes/fat-da7a0001)"
r"(?=.*carries the system tree)"
r"(?=.*data volume; mounted at /volumes/fat-da7a0001)",
"fail": r"data volume; mounted at /volumes/fat-12345678|DANOS-TEST-RESULT: FAIL"},
# S3 shared-channel multi-volume: ONE usb-storage device carrying an MBR with
# TWO FAT partitions (da7a0001 at lba 2048, da7a0002 at lba 83968). allVolumes
# walks the table and the manager spawns a confined fat per partition on the
# SAME block channel, each clamped to its own LBA range (usb-storage's
# per-badge range table) — the path a pair of single-volume sticks (the
# two-volumes case) does NOT exercise. The two mount lines sit at two DISTINCT
# non-zero base_lbas on one device. Against the pre-uncap allVolumes (S3 step
# 2, capped to one partition) only da7a0001 mounts, so the da7a0002 lookaheads
# fail.
{"name": "partitioned-volume",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"data_volume": {"partitions": [{"serial": "DA7A0001", "size_mib": 40},
{"serial": "DA7A0002", "size_mib": 40}]},
"expect": r"(?s)(?=.*volume 0x0*da7a0001 -> \S+ \(pid \d+\), lba 2048, )"
r"(?=.*volume 0x0*da7a0002 -> \S+ \(pid \d+\), lba 83968, )"
r"(?=.*fat: mounted /volumes/fat-da7a0001)"
r"(?=.*fat: mounted /volumes/fat-da7a0002)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the
# fat-test client, after listing, makes a directory, writes+reads a file inside
# it, then removes the file, exercising the whole VFS -> fat mutation path.
@@ -1422,6 +1463,30 @@ def run_case(arch, case):
cmd[cmd.index("-m") + 1] = case["mem"]
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
cmd += case["qemu_extra"]
# A multi-volume case attaches a second usb-storage device backed by a freshly
# GENERATED data volume: a distinct-serial FAT32 with no /system tree, so the
# volume manager mounts it at its own id-path and the fat process marks it a
# data volume (never a system volume). Regenerated per run — no image is
# committed to the tree (the user keeps the boot files copyable, not baked in).
if case.get("data_volume"):
dv = case["data_volume"]
data_img = os.path.join(WORK, "data-volume.img")
if dv.get("partitions"):
# One device, an MBR with several FAT partitions: several volumes share
# ONE block channel, each confined to its own LBA range.
gen = [sys.executable, os.path.join(REPO, "tools", "make-partitioned-image.py"), data_img]
for part in dv["partitions"]:
gen += [part["serial"], str(part.get("size_mib", 40))]
else:
# One device, one bare FAT volume.
gen = [sys.executable, os.path.join(REPO, "tools", "make-fat-image.py"),
"--serial", dv["serial"], "--label", dv.get("label", "DATAVOL"),
data_img, str(dv.get("size_mib", 64))]
subprocess.run(gen, check=True, stdout=subprocess.DEVNULL)
cmd += [
"-drive", f"if=none,id=datausb,format=raw,file={data_img}",
"-device", "usb-storage,bus=xhci.0,port=4,drive=datausb,removable=on,id=datastorage",
]
# A QMP control socket, always present (additive): how a case's `qmp_after`
# hook injects host-side events into the guest mid-run. Kept under a short temp
# dir, not WORK: a unix socket path is capped at ~104 bytes (sun_path), and a
-1
View File
@@ -194,7 +194,6 @@ system/kernel/process.zig:write_buffer
system/kernel/scheduler.zig:ipc_maximum_handles
system/kernel/scheduler.zig:maximum_space_mappings
system/kernel/vfs.zig:maximum_directories
system/kernel/vfs.zig:maximum_mounts
system/kernel/vfs.zig:maximum_prefix
system/kernel/vfs.zig:maximum_rewrite
system/services/acpi/acpi.zig:blocks
+31 -7
View File
@@ -47,8 +47,14 @@ def fat32_geometry(total_sectors):
class Fat32Image:
def __init__(self, total_sectors):
def __init__(self, total_sectors, volume_id=0x12345678, label="DANOS"):
self.total_sectors = total_sectors
# The FAT volume serial (its content identity — the /volumes/fat-<id>
# mount path danos derives from it) and the display label. A second image
# needs a distinct serial so its id-path does not collide with the boot
# volume's.
self.volume_id = volume_id & 0xFFFFFFFF
self.label = label
self.fat_size, self.cluster_count = fat32_geometry(total_sectors)
if self.cluster_count < 65525:
sys.exit(f"error: image too small for FAT32 ({self.cluster_count} clusters "
@@ -146,8 +152,8 @@ class Fat32Image:
0x80, # drive number
0, # reserved
0x29, # extended boot signature
0x12345678, # volume id
b"DANOS ", # volume label
self.volume_id, # volume id
self.label.encode("ascii", "replace")[:11].ljust(11, b" "), # volume label
b"FAT32 ", # filesystem type
)
sector[510] = 0x55
@@ -276,9 +282,9 @@ def build_tree(pairs):
return root
def build(out_path, size_mib, pairs):
def build(out_path, size_mib, pairs, volume_id=0x12345678, label="DANOS"):
total_sectors = size_mib * 1024 * 1024 // SECTOR
image = Fat32Image(total_sectors)
image = Fat32Image(total_sectors, volume_id, label)
tree = build_tree(pairs)
write_directory(image, 2, tree, 0, True)
with open(out_path, "wb") as handle:
@@ -350,14 +356,32 @@ def main(argv):
if len(argv) == 3 and argv[1] == "--verify":
verify(argv[2])
return 0
# Optional flags ahead of the positionals: --serial <hex> sets the FAT volume
# id (the /volumes/fat-<id> content identity), --label <name> its display
# label. A second FAT image passes a distinct --serial so its id-path cannot
# collide with the boot volume's.
argv = list(argv)
volume_id = 0x12345678
label = "DANOS"
i = 1
while i < len(argv):
if argv[i] == "--serial" and i + 1 < len(argv):
volume_id = int(argv[i + 1], 16)
del argv[i:i + 2]
elif argv[i] == "--label" and i + 1 < len(argv):
label = argv[i + 1]
del argv[i:i + 2]
else:
i += 1
if len(argv) < 3 or (len(argv) - 3) % 2 != 0:
sys.exit("usage: make-fat-image.py <out.img> <size-MiB> [<dest> <host>]...\n"
sys.exit("usage: make-fat-image.py [--serial <hex>] [--label <name>] "
"<out.img> <size-MiB> [<dest> <host>]...\n"
" make-fat-image.py --verify <out.img>")
out_path = argv[1]
size_mib = int(argv[2])
rest = argv[3:]
pairs = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)]
build(out_path, size_mib, pairs)
build(out_path, size_mib, pairs, volume_id, label)
return 0
+88
View File
@@ -0,0 +1,88 @@
#!/usr/bin/env python3
"""Assemble an MBR-partitioned disk image from N FAT32 partitions — the danos
multi-volume test disk.
Pure Python 3 stdlib (no mtools / parted). Each partition is a real FAT32
filesystem produced by make-fat-image.py, laid out behind a classic MBR so the
danos partition prober (partition.allVolumes) walks the table and the volume
manager spawns one confined filesystem per partition — several volumes sharing
ONE block channel, each clamped to its own LBA range. That shared-channel,
per-partition path is what a single stick with two partitions exercises and a
pair of single-volume sticks does not.
make-partitioned-image.py <out.img> [<serial-hex> <size-MiB>]...
Each partition is an empty FAT32 with the given volume serial (its /volumes/
fat-<serial> content id). Partitions are 1-MiB aligned; the MBR marks each
type 0x0C (FAT32 LBA). At most four (an MBR holds four primaries).
"""
import os
import struct
import subprocess
import sys
import tempfile
SECTOR = 512
ALIGN = 2048 # sectors (1 MiB) — standard partition alignment, and the gap the MBR sits in
MBR_TYPE_FAT32_LBA = 0x0C
MAX_PRIMARY_PARTITIONS = 4
HERE = os.path.dirname(os.path.abspath(__file__))
def align_up(sectors, to=ALIGN):
return (sectors + to - 1) // to * to
def main(argv):
if len(argv) < 4 or (len(argv) - 2) % 2 != 0:
sys.exit("usage: make-partitioned-image.py <out.img> [<serial-hex> <size-MiB>]...")
out_path = argv[1]
specs = [(argv[i], int(argv[i + 1])) for i in range(2, len(argv), 2)]
if len(specs) > MAX_PRIMARY_PARTITIONS:
sys.exit(f"error: an MBR holds at most {MAX_PRIMARY_PARTITIONS} primary partitions")
# Generate each partition's FAT32 image, then place it at its aligned start.
partitions = [] # (start_sector, sector_count, bytes)
cursor = ALIGN # leave the first 1 MiB for the MBR + alignment gap
with tempfile.TemporaryDirectory() as tmp:
for idx, (serial, size_mib) in enumerate(specs):
part_path = os.path.join(tmp, f"p{idx}.img")
subprocess.run(
[sys.executable, os.path.join(HERE, "make-fat-image.py"),
"--serial", serial, "--label", f"DATA{idx}",
part_path, str(size_mib)],
check=True, stdout=subprocess.DEVNULL)
with open(part_path, "rb") as handle:
data = handle.read()
count = len(data) // SECTOR
partitions.append((cursor, count, data))
cursor = align_up(cursor + count)
total_sectors = cursor
disk = bytearray(total_sectors * SECTOR)
# The MBR: a disk signature, one partition entry per FAT partition, 0x55AA.
# No boot code (this disk is data, never booted); danos's mount() sees the
# signature but no BPB at LBA 0 and takes the MBR-walk path.
struct.pack_into("<I", disk, 440, 0x0D05DA05) # arbitrary but fixed disk signature
for idx, (start, count, data) in enumerate(partitions):
entry = 446 + idx * 16
disk[entry + 0] = 0x00 # not bootable
disk[entry + 1:entry + 4] = b"\xFE\xFF\xFF" # CHS start (LBA-aware tools ignore)
disk[entry + 4] = MBR_TYPE_FAT32_LBA
disk[entry + 5:entry + 8] = b"\xFE\xFF\xFF" # CHS end
struct.pack_into("<I", disk, entry + 8, start) # start LBA
struct.pack_into("<I", disk, entry + 12, count) # sector count
disk[start * SECTOR:start * SECTOR + len(data)] = data
disk[510] = 0x55
disk[511] = 0xAA
with open(out_path, "wb") as handle:
handle.write(disk)
print(f"make-partitioned-image: wrote {out_path} "
f"({total_sectors * SECTOR // (1024 * 1024)} MiB, {len(partitions)} partitions)")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))