Compare commits
8
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0b25cd2c94 | ||
|
|
d59279422e | ||
|
|
bf9f8560c6 | ||
|
|
da7dcce64e | ||
|
|
9750db14da | ||
|
|
b2a5a0a3c6 | ||
|
|
7efe7b72d8 | ||
|
|
d4b544d66b |
@@ -11,15 +11,21 @@
|
||||
> identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:
|
||||
> `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional
|
||||
> override), a volume's mount path IS its content id (`/volumes/<id>`), with the
|
||||
> label as display metadata a `volumes` query returns. **Still pending**: the
|
||||
> `filesystem UUID` rung (needs a non-FAT engine), multi-volume (one FAT volume
|
||||
> today; fat's boot rewrites are unconditional until S3 makes them
|
||||
> content-conditional), the volume manager *consuming* `medium_changed` (removal
|
||||
> is detected by device-presence polling; the event is published but only a
|
||||
> card-reader medium change needs the subscription), and the remount-on-replug
|
||||
> label as display metadata a `volumes` query returns. Multi-volume is **built**:
|
||||
> the manager adopts every storage device, probes each device's whole partition
|
||||
> table, and spawns one range-confined FAT per volume — several volumes across
|
||||
> several devices, or several partitions sharing one device's channel — each at
|
||||
> its own `/volumes/<id>` path with its own supervision. The boot volume is
|
||||
> identified by **content** (a volume backs `/system/configuration` + `/system/logs`
|
||||
> only when it resolves `/system/configuration` on its own media), so it works as
|
||||
> any partition of any device. **Still pending**: the `filesystem UUID` rung and
|
||||
> a second engine (exFAT, S4); the volume manager *consuming* `medium_changed`
|
||||
> (removal is detected by device-presence polling; the event is published but only
|
||||
> a card-reader medium change needs the subscription); the remount-on-replug
|
||||
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller
|
||||
> device, so it is bench-verified). A few markers below are left where a duty is
|
||||
> still pending.
|
||||
> device, so it is bench-verified); and arbitration when two volumes both resolve
|
||||
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
|
||||
> markers below are left where a duty is still pending.
|
||||
|
||||
## The model
|
||||
|
||||
@@ -121,12 +127,16 @@ its own mounts with the kernel; its write cache lives inside the process, so a
|
||||
write error is observed by the code that owns the volume and surfaces on the
|
||||
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
||||
*(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the
|
||||
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there, plus
|
||||
the two `/system` hierarchy rewrites it installs in place (unconditional this
|
||||
increment; S3 makes them content-conditional across volumes). It no longer
|
||||
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there. It
|
||||
installs the two `/system` hierarchy rewrites (`/system/configuration`,
|
||||
`/system/logs`) only when it is the boot volume — decided by **content**: it
|
||||
resolves `/system/configuration` on its own media at mount, so a data volume
|
||||
mounts at its id-path alone and never shadows the running system. It no longer
|
||||
self-acquires a volume — the V3b flip made it receive its volume id and block
|
||||
channel from the volume manager, consistent with "it never discovers devices"
|
||||
above.
|
||||
above. Because several volumes now serve at once, no filesystem binds a shared
|
||||
service name; clients reach each through the kernel mount table (`fs_resolve`
|
||||
routes by prefix to the backing endpoint).
|
||||
|
||||
**Kernel** (mechanism only): the mount table routes paths to backend
|
||||
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
||||
|
||||
@@ -145,9 +145,13 @@ matrix-proven shape; genuinely open.
|
||||
|
||||
**The pressure points, honestly:**
|
||||
|
||||
1. **Multi-volume providers are reserved, not implemented.** The volume
|
||||
manager flow assumes one provider, one volume; NVMe namespaces make
|
||||
endpoint-per-volume real work with hardware demanding it.
|
||||
1. **Multi-volume is built; multi-namespace-per-provider is untried.** The
|
||||
volume manager adopts every device and spawns one range-confined FAT per
|
||||
partition — several volumes across several devices, or several partitions
|
||||
sharing one device's channel, both proven on USB. What is untried is a single
|
||||
provider exposing several volumes as *namespaces* (NVMe): the endpoint and
|
||||
per-badge range machinery generalizes, but no such driver exists yet to
|
||||
exercise it.
|
||||
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
|
||||
one operation in flight, one bounce buffer — fine for a USB2 stick,
|
||||
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
|
||||
@@ -233,13 +237,15 @@ matrix-proven shape; genuinely open.
|
||||
Consequences, each mechanical once identity keys the map: **moving a drive
|
||||
to a different port changes nothing** — same identity, same mount point,
|
||||
whether USB port, hub depth, SATA port, or a stick that left as USB and
|
||||
returned in a SATA dock; **replug remounts at the same path** (a map lookup
|
||||
once the map exists; today's single volume re-probes and remounts at the
|
||||
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the
|
||||
recorded identity of the volume carrying `/system/configuration`, findable
|
||||
on any port; and **duplicate identity is a policy case, not a surprise** —
|
||||
two cloned sticks at once: first keeps the mapped name, second mounts
|
||||
suffixed and is logged loudly, never silently shadowed. Unknown identities
|
||||
returned in a SATA dock; **replug remounts at the same path** (the id-path is
|
||||
content-derived, so a volume returns to `/volumes/<id>` wherever it reappears;
|
||||
remount-on-replug end-to-end is bench-verified, not QEMU-tested, because QEMU
|
||||
can't re-present the boot-controller device); **the boot volume** is the volume
|
||||
that resolves `/system/configuration` on its own media, findable on any port or
|
||||
partition; and **duplicate identity is a known S4 gap** — two cloned sticks
|
||||
share one content id, so today they collide on `/volumes/<id>` (the kernel
|
||||
remount-replaces; the last wins) and each boot-volume claim is logged loudly.
|
||||
Distinguishing them with a suffix is arbitration, deferred to S4. Unknown identities
|
||||
mount under a derived name (sanitized label, else generated) at
|
||||
`/volumes/<name>` — the hierarchy's documented home for attached media,
|
||||
which stands: `/system` is what danos IS; attached media is what it isn't.
|
||||
|
||||
@@ -61,10 +61,14 @@ pub fn Server(comptime Engine: type) type {
|
||||
/// caller does the filesystem-specific bring-up (find the block
|
||||
/// device, set up DMA, mount the engine) and returns a `Volume`.
|
||||
bringUp: *const fn (endpoint: ipc.Handle) ?Volume,
|
||||
/// The vfs contract name to bind. A filesystem serving one volume
|
||||
/// binds "vfs" today; the volume-manager era hands each per-volume
|
||||
/// process its own establishment and this fades.
|
||||
service_name: ?[]const u8 = "vfs",
|
||||
/// A contract name to bind under /protocol, or null to bind none. In
|
||||
/// the volume-manager era every filesystem is a per-volume process and
|
||||
/// clients reach it through the kernel mount table — fs_resolve routes
|
||||
/// a path to its backing endpoint by prefix — so no filesystem binds a
|
||||
/// shared name. Two volumes would collide on one: the second's bind is
|
||||
/// refused and service.run would exit, so its volume never mounts. The
|
||||
/// endpoint still serves as the mount backend without a name.
|
||||
service_name: ?[]const u8 = null,
|
||||
};
|
||||
|
||||
// --- the harness's own state, one set per instantiation ---------------
|
||||
|
||||
+25
-6
@@ -55,7 +55,17 @@ fn tokenIndex(t: u64) u64 {
|
||||
|
||||
// --- the mount table ---------------------------------------------------------
|
||||
|
||||
pub const maximum_mounts = 8;
|
||||
/// bound: prefixes mounted in the kernel VFS table at once
|
||||
/// decided-by: ours
|
||||
/// protects: the `mounts` table below
|
||||
/// at-limit: refuse - installMount returns false and mountBackend propagates it;
|
||||
/// the mounting filesystem's harness logs "could not mount <prefix>" and the
|
||||
/// mount simply does not exist (no silent success). Budget: the initrd's
|
||||
/// top-level dirs (/system, /test) plus one id-path mount per volume and the
|
||||
/// system volume's two FHS rewrites — a few over the volume manager's
|
||||
/// maximum_volumes (16); 32 leaves headroom.
|
||||
/// observed-by: the harness "file-system: could not mount <prefix>" ring line
|
||||
pub const maximum_mounts = 32;
|
||||
const maximum_prefix = 64;
|
||||
const maximum_rewrite = 32;
|
||||
|
||||
@@ -156,7 +166,10 @@ pub fn setInitialRamdisk(image: []const u8) void {
|
||||
for (directories[0..directory_count], 0..) |*d, index| {
|
||||
const parent = parentOf(d.slice());
|
||||
d.parent = directoryIndex(parent) orelse index;
|
||||
if (parent.len == 1) installMount(d.slice(), .kernel_initrd, null, "");
|
||||
// Boot-time install of one mount per top-level initrd dir (/system, /test):
|
||||
// provably few, far under maximum_mounts, so a full table here is
|
||||
// impossible — but discard the result explicitly rather than assume it.
|
||||
if (parent.len == 1) _ = installMount(d.slice(), .kernel_initrd, null, "");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -167,7 +180,12 @@ fn directoryIndex(path: []const u8) ?usize {
|
||||
return null;
|
||||
}
|
||||
|
||||
fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, rewrite: []const u8) void {
|
||||
/// Install (or remount-replace) a prefix. Returns false when the table is full
|
||||
/// and no slot could be claimed — the caller must surface that, never report a
|
||||
/// dropped mount as success. A remount of an already-mounted prefix reuses its
|
||||
/// slot and always succeeds; a /protocol remount is refused-as-noop (returns
|
||||
/// true: the first mount stands, nothing is dropped).
|
||||
fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, rewrite: []const u8) bool {
|
||||
// Remount replaces: a restarted backend re-mounts its prefix.
|
||||
var slot: ?*Mount = null;
|
||||
for (&mounts) |*m| {
|
||||
@@ -176,19 +194,20 @@ fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, re
|
||||
// restarted FAT retakes /volumes/usb; letting it retake /protocol
|
||||
// would hand the whole naming layer to whoever asked second.
|
||||
// First mount wins, and init (PID 1) is always first.
|
||||
if (std.mem.eql(u8, prefix, protocol_root)) return;
|
||||
if (std.mem.eql(u8, prefix, protocol_root)) return true;
|
||||
if (m.backend) |old| ipc.dropRef(old);
|
||||
slot = m;
|
||||
break;
|
||||
}
|
||||
if (slot == null and !m.used) slot = m;
|
||||
}
|
||||
const m = slot orelse return;
|
||||
const m = slot orelse return false;
|
||||
m.* = .{ .used = true, .kind = kind, .backend = backend };
|
||||
@memcpy(m.prefix[0..prefix.len], prefix);
|
||||
m.prefix_len = prefix.len;
|
||||
@memcpy(m.rewrite[0..rewrite.len], rewrite);
|
||||
m.rewrite_len = rewrite.len;
|
||||
return true;
|
||||
}
|
||||
|
||||
// --- resolve -----------------------------------------------------------------
|
||||
@@ -406,7 +425,7 @@ pub fn mountBackend(prefix: []const u8, backend: *ipc.Endpoint, rewrite: []const
|
||||
if (!isInitrdCarveOut(prefix)) return false;
|
||||
}
|
||||
}
|
||||
installMount(prefix, .backend, backend, rewrite);
|
||||
if (!installMount(prefix, .backend, backend, rewrite)) return false; // table full
|
||||
for (&mounts) |*m| {
|
||||
if (m.used and std.mem.eql(u8, m.prefixSlice(), prefix)) m.owner = owner;
|
||||
}
|
||||
|
||||
@@ -173,16 +173,25 @@ fn fatBringUp(endpoint: ipc.Handle) ?Harness.Volume {
|
||||
};
|
||||
std.log.info("mounted FAT ({s}, {d} clusters, partition lba {d})", .{ @tagName(filesystem.geometry.fat_type), filesystem.geometry.cluster_count, filesystem.base_lba });
|
||||
|
||||
// The volume mounts at its id-path (argv[2]), plus the two FHS rewrites so
|
||||
// hierarchy paths (the logger's /system/logs) stay decoupled from which
|
||||
// volume backs them. This single-volume increment's one volume IS the boot
|
||||
// volume, so it installs both unconditionally; S3 (multi-volume) makes the
|
||||
// rewrites content-conditional — installed only by whichever volume carries
|
||||
// the system, decided by content, not order.
|
||||
// Every volume mounts at its own id-path (argv[2]). The boot/system volume —
|
||||
// the one carrying the /system tree — ADDITIONALLY installs the two FHS
|
||||
// rewrites, so hierarchy paths (config reads, the logger's persistent
|
||||
// /system/logs) stay decoupled from which volume backs them. Detection is by
|
||||
// CONTENT, not spawn order: a volume is the system volume iff /system/
|
||||
// configuration resolves on its own media. A data volume has no /system, so it
|
||||
// mounts only at its id-path and never shadows the running system's config or
|
||||
// logs with a dead mount.
|
||||
mount_specs[0] = .{ .prefix = volume_mount_prefix };
|
||||
var mount_count: usize = 1;
|
||||
if (filesystem.resolve("/system/configuration") != null) {
|
||||
std.log.info("volume {d} carries the system tree; backing /system/configuration and /system/logs", .{my_volume_id});
|
||||
mount_specs[1] = .{ .prefix = "/system/configuration", .rewrite = "/system/configuration" };
|
||||
mount_specs[2] = .{ .prefix = "/system/logs", .rewrite = "/system/logs" };
|
||||
return .{ .engine = &filesystem, .mounts = mount_specs[0..3], .flush = flushIfDirty };
|
||||
mount_count = 3;
|
||||
} else {
|
||||
std.log.info("volume {d} is a data volume; mounted at {s}", .{ my_volume_id, volume_mount_prefix });
|
||||
}
|
||||
return .{ .engine = &filesystem, .mounts = mount_specs[0..mount_count], .flush = flushIfDirty };
|
||||
}
|
||||
|
||||
pub fn main(init: process.Init) void {
|
||||
|
||||
@@ -172,39 +172,41 @@ fn setLabelFromUtf16(id: *Identity, name_bytes: []const u8) void {
|
||||
id.label_len = @intCast(out);
|
||||
}
|
||||
|
||||
/// The first GPT volume, or null if LBA 1 is not a valid GPT header or no entry
|
||||
/// validates. The header CRC-32 and the per-entry overflow-safe range check are
|
||||
/// the confinement-safety guards the driver's clamp rests on — the invariant
|
||||
/// firstVolume documents for MBR, extended to untrusted GPT metadata. The
|
||||
/// entry-array CRC is deferred (correctness-only; the range check carries safety).
|
||||
fn gptFirstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
|
||||
/// Append every valid GPT volume to `out` (up to `out.len`), returning the count
|
||||
/// (0 if LBA 1 is not a valid GPT header). The header CRC-32 and the per-entry
|
||||
/// overflow-safe range check are the confinement-safety guards the driver's clamp
|
||||
/// rests on — the invariant documented for MBR, extended to untrusted GPT
|
||||
/// metadata. The entry-array CRC is deferred (correctness-only; the range check
|
||||
/// carries safety).
|
||||
fn gptAllVolumes(reader: SectorReader, device_blocks: u64, out: []Volume) usize {
|
||||
var header: [sector_bytes]u8 = undefined;
|
||||
if (!reader.read(1, &header)) return null;
|
||||
if (!std.mem.eql(u8, header[0..8], gpt_signature)) return null;
|
||||
if (!reader.read(1, &header)) return 0;
|
||||
if (!std.mem.eql(u8, header[0..8], gpt_signature)) return 0;
|
||||
const header_size = std.mem.readInt(u32, header[12..16], .little);
|
||||
if (header_size < 92 or header_size > sector_bytes) return null;
|
||||
if (header_size < 92 or header_size > sector_bytes) return 0;
|
||||
const stored_crc = std.mem.readInt(u32, header[16..20], .little);
|
||||
var check: [sector_bytes]u8 = undefined;
|
||||
@memcpy(check[0..header_size], header[0..header_size]);
|
||||
@memset(check[16..20], 0);
|
||||
if (crc32(check[0..header_size]) != stored_crc) return null;
|
||||
if (crc32(check[0..header_size]) != stored_crc) return 0;
|
||||
|
||||
const entry_lba = std.mem.readInt(u64, header[72..80], .little);
|
||||
const num_entries = std.mem.readInt(u32, header[80..84], .little);
|
||||
const entry_size = std.mem.readInt(u32, header[84..88], .little);
|
||||
if (entry_size != 128 and entry_size != 256 and entry_size != 512) return null;
|
||||
if (entry_lba == 0 or entry_lba >= device_blocks) return null;
|
||||
if (entry_size != 128 and entry_size != 256 and entry_size != 512) return 0;
|
||||
if (entry_lba == 0 or entry_lba >= device_blocks) return 0;
|
||||
|
||||
const scan = @min(num_entries, gpt_entry_scan_maximum);
|
||||
var sector_buf: [sector_bytes]u8 = undefined;
|
||||
var loaded: u64 = std.math.maxInt(u64);
|
||||
var count: usize = 0;
|
||||
var i: u32 = 0;
|
||||
while (i < scan) : (i += 1) {
|
||||
while (i < scan and count < out.len) : (i += 1) {
|
||||
const abs = @as(u64, i) * entry_size;
|
||||
const lba = entry_lba + abs / sector_bytes;
|
||||
const off = @as(usize, @intCast(abs % sector_bytes));
|
||||
if (lba != loaded) {
|
||||
if (!reader.read(lba, §or_buf)) return null;
|
||||
if (!reader.read(lba, §or_buf)) break; // return what we have
|
||||
loaded = lba;
|
||||
}
|
||||
const entry = sector_buf[off..][0..128]; // the fields we read live in the first 128 bytes
|
||||
@@ -224,9 +226,10 @@ fn gptFirstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
|
||||
if (start == 0 or end < start or end >= device_blocks) continue;
|
||||
var id = Identity{ .rung = .gpt_guid, .key = std.mem.readInt(u128, entry[16..32], .little) };
|
||||
setLabelFromUtf16(&id, entry[56..128]);
|
||||
return .{ .base_lba = start, .block_count = end - start + 1, .identity = id };
|
||||
out[count] = .{ .base_lba = start, .block_count = end - start + 1, .identity = id };
|
||||
count += 1;
|
||||
}
|
||||
return null;
|
||||
return count;
|
||||
}
|
||||
|
||||
/// Trim trailing spaces (FAT labels are space-padded) and copy into the display
|
||||
@@ -259,18 +262,22 @@ fn fatIdentity(reader: SectorReader, start_lba: u64) ?Identity {
|
||||
return id;
|
||||
}
|
||||
|
||||
/// The first volume on the device `reader` addresses, whose whole-device size is
|
||||
/// `device_blocks`, or null if none is found. A GPT disk (protective MBR) is
|
||||
/// handled by GPT, authoritatively — its null is final. Otherwise an MBR with a
|
||||
/// non-empty entry yields that partition's [start, size); otherwise a boot
|
||||
/// signature with no partitions is treated as a bare FAT spanning the device.
|
||||
pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
|
||||
/// Append every volume on the device `reader` addresses, whose whole-device size
|
||||
/// is `device_blocks`, to `out` (up to `out.len`), returning the count. A GPT
|
||||
/// disk (protective MBR) is enumerated by GPT, authoritatively — a zero count is
|
||||
/// final. Otherwise every fitting MBR entry is a volume; a boot signature with no
|
||||
/// partition entries is a bare FAT spanning the whole device. Each volume's
|
||||
/// [start, count) is validated overflow-safe (the confinement invariant the
|
||||
/// driver's clamp rests on), and each prefers its FAT serial identity over the
|
||||
/// disk signature.
|
||||
pub fn allVolumes(reader: SectorReader, device_blocks: u64, out: []Volume) usize {
|
||||
var block0: [sector_bytes]u8 = undefined;
|
||||
if (!reader.read(0, &block0)) return null;
|
||||
if (!hasBootSignature(&block0)) return null;
|
||||
if (isProtectiveMbr(&block0)) return gptFirstVolume(reader, device_blocks);
|
||||
if (!reader.read(0, &block0)) return 0;
|
||||
if (!hasBootSignature(&block0)) return 0;
|
||||
if (isProtectiveMbr(&block0)) return gptAllVolumes(reader, device_blocks, out);
|
||||
var count: usize = 0;
|
||||
var index: u8 = 0;
|
||||
while (index < 4) : (index += 1) {
|
||||
while (index < 4 and count < out.len) : (index += 1) {
|
||||
const entry = block0[446 + @as(usize, index) * 16 ..][0..16];
|
||||
const kind = entry[4];
|
||||
const start = std.mem.readInt(u32, entry[8..12], .little);
|
||||
@@ -283,10 +290,25 @@ pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
|
||||
// device (usb-storage.zig resolveTransfer), which only holds because the
|
||||
// range handed down is validated here. The subtraction cannot overflow.
|
||||
if (start > device_blocks or device_blocks - start < size) continue;
|
||||
return .{ .base_lba = start, .block_count = size, .identity = fatIdentity(reader, start) orelse mbrIdentity(&block0, index) };
|
||||
out[count] = .{ .base_lba = start, .block_count = size, .identity = fatIdentity(reader, start) orelse mbrIdentity(&block0, index) };
|
||||
count += 1;
|
||||
}
|
||||
if (count == 0 and out.len > 0) {
|
||||
// No partition entries: a bare FAT spanning the device.
|
||||
return .{ .base_lba = 0, .block_count = device_blocks, .identity = fatIdentity(reader, 0) orelse mbrIdentity(&block0, 0) };
|
||||
out[0] = .{ .base_lba = 0, .block_count = device_blocks, .identity = fatIdentity(reader, 0) orelse mbrIdentity(&block0, 0) };
|
||||
return 1;
|
||||
}
|
||||
return count;
|
||||
}
|
||||
|
||||
/// firstVolume is allVolumes into a one-element buffer.
|
||||
const one_volume_slot = 1;
|
||||
|
||||
/// The first volume on the device, or null — the single-volume case of
|
||||
/// `allVolumes`, kept for callers that want just one.
|
||||
pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
|
||||
var one: [one_volume_slot]Volume = undefined;
|
||||
return if (allVolumes(reader, device_blocks, &one) > 0) one[0] else null;
|
||||
}
|
||||
|
||||
/// A read-only RAM disk over a byte slice of sectors, for the host tests.
|
||||
@@ -507,3 +529,33 @@ test "GPT with 256-byte entries reads the non-128 offset arithmetic correctly" {
|
||||
try std.testing.expectEqual(Rung.gpt_guid, v.identity.rung);
|
||||
try std.testing.expectEqual(@as(u128, 0xF00D), v.identity.key);
|
||||
}
|
||||
|
||||
// A fixture-sized volume buffer for the multi-volume tests, named so the bounds
|
||||
// gate (which flags literal array lengths) stays quiet: a test input.
|
||||
const test_volume_slots = 4;
|
||||
|
||||
test "allVolumes returns every fitting MBR partition with distinct identities" {
|
||||
var block0 = [_]u8{0} ** 512;
|
||||
block0[510] = 0x55;
|
||||
block0[511] = 0xAA;
|
||||
std.mem.writeInt(u32, block0[440..444], 0xDEADBEEF, .little);
|
||||
// partition 0: start 2048, size 1000
|
||||
block0[446 + 4] = 0x0c;
|
||||
std.mem.writeInt(u32, block0[446 + 8 ..][0..4], 2048, .little);
|
||||
std.mem.writeInt(u32, block0[446 + 12 ..][0..4], 1000, .little);
|
||||
// partition 1: start 4096, size 2000
|
||||
block0[462 + 4] = 0x0c;
|
||||
std.mem.writeInt(u32, block0[462 + 8 ..][0..4], 4096, .little);
|
||||
std.mem.writeInt(u32, block0[462 + 12 ..][0..4], 2000, .little);
|
||||
const disk = RamDisk{ .sectors = &block0 };
|
||||
var vols: [test_volume_slots]Volume = undefined;
|
||||
const n = allVolumes(disk.reader(), 200000, &vols);
|
||||
try std.testing.expectEqual(@as(usize, 2), n); // both partitions, not just the first
|
||||
try std.testing.expectEqual(@as(u64, 2048), vols[0].base_lba);
|
||||
try std.testing.expectEqual(@as(u64, 4096), vols[1].base_lba);
|
||||
// distinct rung-4 identities (no FAT VBR at those LBAs): index 0 vs 1.
|
||||
try std.testing.expectEqual((@as(u128, 0xDEADBEEF) << 8) | 0, vols[0].identity.key);
|
||||
try std.testing.expectEqual((@as(u128, 0xDEADBEEF) << 8) | 1, vols[1].identity.key);
|
||||
// firstVolume (the 1-buffer case) still returns just the first.
|
||||
try std.testing.expectEqual(@as(u64, 2048), firstVolume(disk.reader(), 200000).?.base_lba);
|
||||
}
|
||||
|
||||
@@ -8,10 +8,12 @@
|
||||
//! supervises the filesystems it spawns, exactly as the device manager
|
||||
//! supervises drivers.
|
||||
//!
|
||||
//! This increment (V3b) is the flip: the FAT service stops acquiring its own
|
||||
//! volume and is spawned here instead, confined to its partition, and handed
|
||||
//! its channel over the volume-manager protocol. Single volume for now; the
|
||||
//! mount map (volumes.csv) and multi-volume land next.
|
||||
//! The manager holds a table of adopted storage DEVICES and a table of the
|
||||
//! VOLUMES on them: it adopts every storage device the device-manager tree
|
||||
//! carries, probes each one's whole partition table, and spawns one filesystem
|
||||
//! process per volume — each confined to its partition's badge-scoped block
|
||||
//! range, each supervised with its own budget. A device leaving the tree takes
|
||||
//! its volumes with it.
|
||||
|
||||
const std = @import("std");
|
||||
const channel = @import("channel");
|
||||
@@ -35,22 +37,37 @@ const Serve = volume_manager_protocol.Protocol.Provider(void);
|
||||
const Invocation = envelope.Invocation;
|
||||
const Answer = envelope.Answer;
|
||||
|
||||
/// The single volume this increment handles: its provider channel, its block
|
||||
/// sub-range, its identity, the id it is addressed by, and the filesystem
|
||||
/// process serving it (0 until spawned; reset on death for respawn).
|
||||
const Volume = struct {
|
||||
storage: block.Device,
|
||||
storage_device_id: u64, // the device-manager id this volume's provider serves
|
||||
base_lba: u64,
|
||||
block_count: u64,
|
||||
identity: partition.Identity,
|
||||
id: u64,
|
||||
binary: []const u8, // the service binary, from filesystems.csv by signature
|
||||
mount_prefix: []const u8, // the volume-root mount path (its id-path, or a volumes.csv override)
|
||||
filesystem_pid: u32 = 0,
|
||||
/// One adopted storage device: the block channel to its provider (opened once and
|
||||
/// shared — refcounted per confined filesystem via the hello reply) and the
|
||||
/// device-manager id it serves. A device leaving the tree takes its volumes.
|
||||
const StorageDevice = struct {
|
||||
used: bool = false,
|
||||
device_id: u64 = 0,
|
||||
channel: block.Device = undefined,
|
||||
};
|
||||
|
||||
const volume_id: u64 = 1;
|
||||
/// One volume: which device serves it, its block sub-range, its content
|
||||
/// identity, the id it is addressed by, the service binary + mount path it was
|
||||
/// spawned with, the filesystem process serving it, and its own supervision
|
||||
/// budget (so one volume's crash loop never touches another's).
|
||||
const Volume = struct {
|
||||
used: bool = false,
|
||||
device_id: u64 = 0,
|
||||
base_lba: u64 = 0,
|
||||
block_count: u64 = 0,
|
||||
identity: partition.Identity = .{ .rung = .anonymous },
|
||||
id: u64 = 0,
|
||||
binary: []const u8 = "",
|
||||
mount_prefix: []const u8 = "",
|
||||
filesystem_pid: u32 = 0,
|
||||
// Per-volume supervision, mirroring the device manager's: a clean exit is not
|
||||
// restarted, a fault restarts with backoff, a fast crash loop gives up.
|
||||
restarts: u32 = 0,
|
||||
spawn_ns: u64 = 0,
|
||||
failed: bool = false,
|
||||
restart_pending: bool = false,
|
||||
restart_due_ns: u64 = 0,
|
||||
};
|
||||
|
||||
// The mount map, read from configuration at boot (the policy home, storage-
|
||||
// architecture.md): filesystems.csv (content signature -> service binary) and
|
||||
@@ -80,47 +97,77 @@ var filesystem_rules: [maximum_filesystem_rules]filesystem_map.Rule = undefined;
|
||||
var filesystem_rule_count: usize = 0;
|
||||
var volume_rules: [maximum_volume_rules]volume_map.Override = undefined;
|
||||
var volume_rule_count: usize = 0;
|
||||
/// The composed default mount path (/volumes/<id>) for the current volume; a
|
||||
/// volumes.csv override is used in place and needs no buffer (it is already a
|
||||
/// slice into volumes_source). One buffer suffices while the manager serves one
|
||||
/// volume (multi-volume gives each its own in S3).
|
||||
/// bound: bytes of a composed /volumes/<id> mount path
|
||||
/// decided-by: ours
|
||||
/// protects: the mount_prefix_buf below
|
||||
/// protects: the per-volume mount_prefix buffers below
|
||||
/// at-limit: truncate - bufPrint fails; the volume mounts at a fallback path (logged)
|
||||
/// observed-by: the fallback path in the log
|
||||
const mount_path_maximum = 64;
|
||||
var mount_prefix_buf: [mount_path_maximum]u8 = undefined;
|
||||
|
||||
/// bound: volumes the manager serves at once
|
||||
/// decided-by: ours
|
||||
/// protects: the volumes table and its per-volume mount-path buffers
|
||||
/// at-limit: truncate - a further partition is left unserved and logged (real
|
||||
/// machines carry a handful of volumes, far under this)
|
||||
/// observed-by: the "volume table full" log line
|
||||
const maximum_volumes = 16;
|
||||
/// bound: storage devices the manager adopts at once
|
||||
/// decided-by: ours
|
||||
/// protects: the devices table
|
||||
/// at-limit: truncate - a further device is left unadopted and logged
|
||||
/// observed-by: the "device table full" log line
|
||||
const maximum_devices = 8;
|
||||
var devices = [_]StorageDevice{.{}} ** maximum_devices;
|
||||
var volumes = [_]Volume{.{}} ** maximum_volumes;
|
||||
/// Each volume's composed default mount path lives in its slot's buffer; a
|
||||
/// volumes.csv override is used in place (a slice into volumes_source, no buffer).
|
||||
var mount_prefix_bufs: [maximum_volumes][mount_path_maximum]u8 = undefined;
|
||||
var next_volume_id: u64 = 1; // monotonic — never reused, so a stale id can't address the wrong child
|
||||
|
||||
var service_endpoint: ipc.Handle = 0;
|
||||
var manager_handle: ?ipc.Handle = null;
|
||||
var bounce: memory.DmaRegion = undefined;
|
||||
var bounce_ready = false;
|
||||
/// The currently-mounted volume, or null while no storage is present. The whole
|
||||
/// removal lifecycle is this field going null and back: the poll sees the
|
||||
/// storage provider leave the device tree (a pulled stick), kills the filesystem
|
||||
/// and clears this; when it returns, the poll re-acquires and re-mounts.
|
||||
var volume: ?Volume = null;
|
||||
var logged_no_volume = false;
|
||||
/// How often the poll checks whether the storage provider is present. Fast
|
||||
/// enough that an unplug unmounts promptly; the poll is a bare device-manager
|
||||
/// enumerate, no channel work, so it is cheap to run continuously.
|
||||
/// How often the poll checks device presence and fires due restarts. Fast enough
|
||||
/// that an unplug unmounts promptly; the poll is a bare device-manager enumerate,
|
||||
/// no channel work, so it is cheap to run continuously.
|
||||
const poll_interval_ms = 500;
|
||||
|
||||
// Filesystem supervision, mirroring the device manager's (device-manager.zig):
|
||||
// a clean exit is not restarted, a fault restarts with backoff, and a fast
|
||||
// crash loop gives up rather than spinning. Without this a faulting filesystem
|
||||
// respawns in a zero-delay loop.
|
||||
// Filesystem supervision, mirroring the device manager's (device-manager.zig).
|
||||
const fast_death_ns: u64 = 2_000_000_000;
|
||||
const crash_loop_cap: u32 = 3;
|
||||
const backoff_base_ms: u64 = 300;
|
||||
var fs_restarts: u32 = 0;
|
||||
var fs_spawn_ns: u64 = 0;
|
||||
var fs_failed = false;
|
||||
/// A fat restart is due at `restart_due_ns`; the poll loop performs it once the
|
||||
/// backoff has elapsed (one timer, folded into the poll — no second timer).
|
||||
var restart_pending = false;
|
||||
var restart_due_ns: u64 = 0;
|
||||
/// bytes to format a u64 volume id as decimal (20 digits fit)
|
||||
const id_decimal_bytes = 24;
|
||||
|
||||
// --- table lookups -----------------------------------------------------------
|
||||
|
||||
fn deviceById(id: u64) ?*StorageDevice {
|
||||
for (&devices) |*d| if (d.used and d.device_id == id) return d;
|
||||
return null;
|
||||
}
|
||||
fn claimDevice() ?*StorageDevice {
|
||||
for (&devices) |*d| if (!d.used) return d;
|
||||
return null;
|
||||
}
|
||||
fn volumeById(id: u64) ?*Volume {
|
||||
for (&volumes) |*v| if (v.used and v.id == id) return v;
|
||||
return null;
|
||||
}
|
||||
fn volumeByPid(pid: u32) ?*Volume {
|
||||
for (&volumes) |*v| if (v.used and v.filesystem_pid == pid) return v;
|
||||
return null;
|
||||
}
|
||||
fn firstUsedVolume() ?*Volume {
|
||||
for (&volumes) |*v| if (v.used) return v;
|
||||
return null;
|
||||
}
|
||||
fn claimVolumeIndex() ?usize {
|
||||
for (&volumes, 0..) |*v, i| if (!v.used) return i;
|
||||
return null;
|
||||
}
|
||||
|
||||
// --- device-manager plumbing -------------------------------------------------
|
||||
|
||||
fn deviceManager() ?ipc.Handle {
|
||||
if (manager_handle) |h| return h;
|
||||
@@ -131,13 +178,12 @@ fn deviceManager() ?ipc.Handle {
|
||||
|
||||
const OpenedStorage = struct { device_id: u64, device: block.Device };
|
||||
|
||||
/// The first mass-storage provider whose block channel actually opens, with its
|
||||
/// device id. A device-manager tree can carry more than one entry of the
|
||||
/// mass-storage identity — a phantom that no driver is bound to answers a
|
||||
/// consumer hello with NO channel — so this tries each and takes the first that
|
||||
/// yields a channel, exactly as a filesystem's own acquisition loop does.
|
||||
/// Called only when there is no volume (an insertion), so the hellos it makes
|
||||
/// are not per-poll churn.
|
||||
/// The first mass-storage provider whose block channel opens and is NOT already
|
||||
/// adopted, with its device id. A device-manager tree can carry more than one
|
||||
/// entry of the mass-storage identity — a phantom that no driver is bound to
|
||||
/// answers a consumer hello with NO channel — so this tries each and takes the
|
||||
/// first that yields a channel. Skips already-adopted devices so a re-poll does
|
||||
/// not re-open a device it already serves.
|
||||
fn openAnyStorage() ?OpenedStorage {
|
||||
const manager = deviceManager() orelse return null;
|
||||
const Entry = device_manager_protocol.ChildEntry;
|
||||
@@ -157,6 +203,7 @@ fn openAnyStorage() ?OpenedStorage {
|
||||
const entry = std.mem.bytesToValue(Entry, tail[index * @sizeOf(Entry) ..][0..@sizeOf(Entry)]);
|
||||
if (entry.device_id == device_manager_protocol.no_device) continue;
|
||||
if ((entry.identity >> 16) & 0xff != 0x08 or (entry.identity >> 8) & 0xff != 0x06) continue;
|
||||
if (deviceById(entry.device_id) != null) continue; // already adopted
|
||||
const exchanged = driver.helloOn(manager, .consumer, entry.device_id, null, true) orelse continue;
|
||||
const provider = exchanged.channel orelse continue; // a phantom / not-yet-bound entry
|
||||
return .{ .device_id = entry.device_id, .device = .{ .endpoint = provider } };
|
||||
@@ -167,7 +214,7 @@ fn openAnyStorage() ?OpenedStorage {
|
||||
|
||||
/// Whether `device_id` is still in the device-manager tree — a bare enumerate,
|
||||
/// no consumer-hello, so it is cheap to call every poll. This is how removal is
|
||||
/// detected: the specific device the mounted volume sits on disappears.
|
||||
/// detected: the specific device a mounted volume sits on disappears.
|
||||
fn isDevicePresent(device_id: u64) bool {
|
||||
const manager = deviceManager() orelse return false;
|
||||
const Entry = device_manager_protocol.ChildEntry;
|
||||
@@ -191,59 +238,88 @@ fn isDevicePresent(device_id: u64) bool {
|
||||
}
|
||||
}
|
||||
|
||||
/// Spawn the filesystem for `v`, confine it to the volume's range, and record
|
||||
/// its pid. The confinement is defined for the fresh pid BEFORE the filesystem
|
||||
/// runs, so its first read is already bounded; the volume manager is the
|
||||
/// confinement controller (it defines the first range on the device).
|
||||
// --- lifecycle ---------------------------------------------------------------
|
||||
|
||||
/// Spawn the filesystem for `v`, confine it to the volume's range on its device's
|
||||
/// channel, and record its pid. The confinement is defined for the fresh pid
|
||||
/// BEFORE the filesystem runs, so its first read is already bounded; the volume
|
||||
/// manager is the confinement controller (it defines the first range on the
|
||||
/// device).
|
||||
fn spawnFilesystem(v: *Volume) void {
|
||||
if (fs_failed) return;
|
||||
const pid = process.spawnSupervised(v.binary, &.{ "1", v.mount_prefix }, service_endpoint) orelse {
|
||||
if (v.failed) return;
|
||||
const dev = deviceById(v.device_id) orelse return; // its device left — poll will clean up
|
||||
var id_str_buf: [id_decimal_bytes]u8 = undefined;
|
||||
const id_str = std.fmt.bufPrint(&id_str_buf, "{d}", .{v.id}) catch "1";
|
||||
const pid = process.spawnSupervised(v.binary, &.{ id_str, v.mount_prefix }, service_endpoint) orelse {
|
||||
_ = logging.write("volume-manager: could not spawn the filesystem; retrying\n");
|
||||
armRestart();
|
||||
armRestart(v);
|
||||
return;
|
||||
};
|
||||
if (!v.storage.defineRange(pid, v.base_lba, v.block_count)) {
|
||||
if (!dev.channel.defineRange(pid, v.base_lba, v.block_count)) {
|
||||
_ = logging.write("volume-manager: could not confine the filesystem to its volume; retrying\n");
|
||||
_ = process.kill(pid);
|
||||
armRestart();
|
||||
armRestart(v);
|
||||
return;
|
||||
}
|
||||
v.filesystem_pid = pid;
|
||||
fs_spawn_ns = time.clock();
|
||||
v.spawn_ns = time.clock();
|
||||
std.log.info("volume 0x{x} -> {s} (pid {d}), lba {d}, {d} blocks", .{ v.identity.key, v.binary, pid, v.base_lba, v.block_count });
|
||||
}
|
||||
|
||||
/// Schedule a fat restart after backoff; the poll loop performs it once due.
|
||||
fn armRestart() void {
|
||||
const delay = if (fs_restarts == 0) backoff_base_ms else backoff_base_ms << @intCast(@min(fs_restarts - 1, 5));
|
||||
restart_due_ns = time.clock() + delay * 1_000_000;
|
||||
restart_pending = true;
|
||||
/// Schedule a restart for `v` after backoff; the poll loop performs it once due.
|
||||
fn armRestart(v: *Volume) void {
|
||||
const delay = if (v.restarts == 0) backoff_base_ms else backoff_base_ms << @intCast(@min(v.restarts - 1, 5));
|
||||
v.restart_due_ns = time.clock() + delay * 1_000_000;
|
||||
v.restart_pending = true;
|
||||
}
|
||||
|
||||
/// A storage provider just appeared: open its channel, read block 0, parse the
|
||||
/// volume, and spawn its filesystem. On any failure the channel is closed (so a
|
||||
/// present-but-unreadable device does not leak a handle every poll) and `volume`
|
||||
/// stays null — the next poll retries. A fresh medium gets a fresh supervision
|
||||
/// budget.
|
||||
fn bringUpVolume() void {
|
||||
/// Compose a volume's mount path (its id-path `/volumes/<id>`, or a volumes.csv
|
||||
/// override) into its slot's buffer, and return the slice.
|
||||
fn composeMountPrefix(slot: usize, identity: partition.Identity) []const u8 {
|
||||
var id_buf: [volume_map.id_maximum]u8 = undefined;
|
||||
const id = volume_map.idString(identity, &id_buf);
|
||||
return volume_map.overrideFor(volume_rules[0..volume_rule_count], id) orelse
|
||||
(std.fmt.bufPrint(&mount_prefix_bufs[slot], "/volumes/{s}", .{id}) catch "/volumes/unknown");
|
||||
}
|
||||
|
||||
/// Adopt the next present, not-yet-adopted storage device: take its channel,
|
||||
/// probe its whole partition table, and spawn a filesystem per volume it carries.
|
||||
/// Returns true when it consumed a device (so the caller can loop to adopt every
|
||||
/// present device in one tick), false when none remain or the device table is full.
|
||||
///
|
||||
/// A device is adopted exactly once and kept until it leaves the tree — even when
|
||||
/// it carries no volume we can serve, or its geometry cannot be read. Keeping the
|
||||
/// empty/unreadable device adopted (rather than dropping and re-probing) is what
|
||||
/// lets openAnyStorage advance PAST it to the devices behind it; dropping it would
|
||||
/// make openAnyStorage hand back the same unservable device every tick and starve
|
||||
/// the rest. A genuine removal frees the slot (removeDevice); a re-insert gets a
|
||||
/// fresh device id and is probed anew.
|
||||
fn bringUpVolume() bool {
|
||||
if (!bounce_ready) {
|
||||
bounce = memory.dmaAlloc(512, memory.dma_coherent | memory.dma_shareable) orelse return;
|
||||
bounce = memory.dmaAlloc(512, memory.dma_coherent | memory.dma_shareable) orelse return false;
|
||||
bounce_ready = true;
|
||||
}
|
||||
const opened = openAnyStorage() orelse return;
|
||||
const opened = openAnyStorage() orelse return false;
|
||||
const dev = claimDevice() orelse {
|
||||
_ = logging.write("volume-manager: device table full; a storage device is left unadopted\n");
|
||||
_ = ipc.close(opened.device.endpoint);
|
||||
return false;
|
||||
};
|
||||
dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device };
|
||||
const device = opened.device;
|
||||
// Attach the read buffer to THIS device (a no-op without an enforcing IOMMU).
|
||||
// The handle is kept, not closed, so it can be re-attached to the next
|
||||
// device after a replug.
|
||||
// The handle is kept, not closed, so it can be re-attached after a replug. A
|
||||
// failed attach or geometry read leaves the device adopted but empty — we just
|
||||
// cannot read it, and the slot still watches it for removal.
|
||||
if (bounce.handle) |handle| {
|
||||
if (!device.attach(handle)) {
|
||||
_ = ipc.close(device.endpoint);
|
||||
return;
|
||||
_ = logging.write("volume-manager: could not attach the read buffer to a storage device; no volume served\n");
|
||||
return true;
|
||||
}
|
||||
}
|
||||
const geometry = device.geometry() orelse {
|
||||
_ = ipc.close(device.endpoint);
|
||||
return;
|
||||
_ = logging.write("volume-manager: could not read a storage device's geometry; no volume served\n");
|
||||
return true;
|
||||
};
|
||||
const ProbeReader = struct {
|
||||
device: block.Device,
|
||||
@@ -257,77 +333,88 @@ fn bringUpVolume() void {
|
||||
};
|
||||
var probe = ProbeReader{ .device = device };
|
||||
const reader = partition.SectorReader{ .context = &probe, .readFn = ProbeReader.readSector };
|
||||
const found = partition.firstVolume(reader, geometry.block_count) orelse {
|
||||
if (!logged_no_volume) {
|
||||
_ = logging.write("volume-manager: storage present but no recognizable volume\n");
|
||||
logged_no_volume = true;
|
||||
var found: [maximum_volumes]partition.Volume = undefined;
|
||||
const n = partition.allVolumes(reader, geometry.block_count, found[0..]);
|
||||
if (n == 0) {
|
||||
std.log.info("device {d} present but carries no recognizable volume", .{dev.device_id});
|
||||
return true;
|
||||
}
|
||||
_ = ipc.close(device.endpoint);
|
||||
return;
|
||||
};
|
||||
for (found[0..n]) |fv| {
|
||||
// Pick the service binary from the volume's content signature. A signature
|
||||
// no filesystems.csv row serves goes unserved (logged), like an unbound
|
||||
// device — the manager does not guess.
|
||||
const binary = filesystem_map.match(filesystem_rules[0..filesystem_rule_count], found.signature) orelse {
|
||||
if (!logged_no_volume) {
|
||||
const binary = filesystem_map.match(filesystem_rules[0..filesystem_rule_count], fv.signature) orelse {
|
||||
_ = logging.write("volume-manager: no filesystem serves this volume's content; unserved\n");
|
||||
logged_no_volume = true;
|
||||
}
|
||||
_ = ipc.close(device.endpoint);
|
||||
return;
|
||||
continue;
|
||||
};
|
||||
// The mount path is the volume's identity id (/volumes/<id>), or a
|
||||
// volumes.csv override pinning it to a chosen path. The id is content-derived,
|
||||
// so the path is stable and never a port or a label.
|
||||
var id_buf: [volume_map.id_maximum]u8 = undefined;
|
||||
const id = volume_map.idString(found.identity, &id_buf);
|
||||
const mount_prefix = volume_map.overrideFor(volume_rules[0..volume_rule_count], id) orelse
|
||||
(std.fmt.bufPrint(&mount_prefix_buf, "/volumes/{s}", .{id}) catch "/volumes/unknown");
|
||||
logged_no_volume = false;
|
||||
fs_restarts = 0;
|
||||
fs_failed = false;
|
||||
restart_pending = false;
|
||||
volume = .{ .storage = device, .storage_device_id = opened.device_id, .base_lba = found.base_lba, .block_count = found.block_count, .identity = found.identity, .id = volume_id, .binary = binary, .mount_prefix = mount_prefix };
|
||||
spawnFilesystem(&volume.?);
|
||||
const slot = claimVolumeIndex() orelse {
|
||||
_ = logging.write("volume-manager: volume table full; a volume is left unserved\n");
|
||||
break;
|
||||
};
|
||||
volumes[slot] = .{
|
||||
.used = true,
|
||||
.device_id = dev.device_id,
|
||||
.base_lba = fv.base_lba,
|
||||
.block_count = fv.block_count,
|
||||
.identity = fv.identity,
|
||||
.id = next_volume_id,
|
||||
.binary = binary,
|
||||
.mount_prefix = composeMountPrefix(slot, fv.identity),
|
||||
};
|
||||
next_volume_id += 1;
|
||||
spawnFilesystem(&volumes[slot]);
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/// The storage provider left the device tree (a pulled stick): kill the
|
||||
/// filesystem so its mounts are retired. Retirement is lazy, not an eager
|
||||
/// death-time sweep — killing the process marks the filesystem's backend
|
||||
/// endpoint dead, and the VFS router drops each mount that endpoint backed on
|
||||
/// the next path resolution under it (that resolve frees the slot and returns
|
||||
/// not_found). Then drop the now-dead channel and clear the volume; the next
|
||||
/// poll that sees storage return re-mounts.
|
||||
fn removeVolume() void {
|
||||
const v = volume orelse return;
|
||||
/// Close a device's channel and free its slot. No volumes are touched (the caller
|
||||
/// ensures none remain, or there never were any).
|
||||
fn dropDevice(dev: *StorageDevice) void {
|
||||
_ = ipc.close(dev.channel.endpoint);
|
||||
dev.* = .{};
|
||||
}
|
||||
|
||||
/// Retire one volume: kill its filesystem so its mounts are retired. Retirement
|
||||
/// is lazy, not an eager death-time sweep — killing the process marks the
|
||||
/// filesystem's backend endpoint dead, and the VFS router drops each mount that
|
||||
/// endpoint backed on the next path resolution under it (that resolve frees the
|
||||
/// slot and returns not_found). Then free the volume slot.
|
||||
fn removeVolumeState(v: *Volume) void {
|
||||
std.log.info("storage for volume {d} removed; unmounting", .{v.id});
|
||||
if (v.filesystem_pid != 0) _ = process.kill(v.filesystem_pid);
|
||||
_ = ipc.close(v.storage.endpoint);
|
||||
volume = null;
|
||||
restart_pending = false;
|
||||
fs_restarts = 0;
|
||||
fs_failed = false;
|
||||
v.* = .{};
|
||||
}
|
||||
|
||||
/// One poll tick. Removal is checked FIRST and supersedes a pending restart: if
|
||||
/// the device is gone there is nothing to restart fat onto, and respawning it
|
||||
/// against the dead channel would just churn until the crash cap. Only once the
|
||||
/// device is confirmed present does a due restart fire.
|
||||
/// A storage device left the tree (a pulled stick): retire every volume it served
|
||||
/// and drop its channel. One removal path, whether the device is pulled cleanly
|
||||
/// or vanishes.
|
||||
fn removeDevice(dev: *StorageDevice) void {
|
||||
for (&volumes) |*v| {
|
||||
if (v.used and v.device_id == dev.device_id) removeVolumeState(v);
|
||||
}
|
||||
dropDevice(dev);
|
||||
}
|
||||
|
||||
/// One poll tick. Device removal is reconciled FIRST and supersedes a pending
|
||||
/// restart: a volume whose device left is retired before its restart could fire,
|
||||
/// so nothing respawns against a dead channel. Then due restarts fire for present
|
||||
/// volumes; then, if no device is adopted, a present device is brought up.
|
||||
fn pollTick() void {
|
||||
if (volume) |v| {
|
||||
// Serving: watch for the specific device leaving (a pulled stick).
|
||||
if (!isDevicePresent(v.storage_device_id)) {
|
||||
removeVolume();
|
||||
return;
|
||||
for (&devices) |*dev| {
|
||||
if (dev.used and !isDevicePresent(dev.device_id)) removeDevice(dev);
|
||||
}
|
||||
if (restart_pending and time.clock() >= restart_due_ns) {
|
||||
restart_pending = false;
|
||||
spawnFilesystem(&volume.?);
|
||||
for (&volumes) |*v| {
|
||||
if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) {
|
||||
v.restart_pending = false;
|
||||
spawnFilesystem(v);
|
||||
}
|
||||
} else {
|
||||
// Idle: try to bring a present storage device up.
|
||||
bringUpVolume();
|
||||
}
|
||||
// Adopt every present, not-yet-adopted storage device. Each call consumes at
|
||||
// most one device (openAnyStorage skips the adopted), so the loop terminates
|
||||
// once none remain; the maximum_devices guard is insurance against a logic
|
||||
// slip, never the normal exit.
|
||||
var adopted: usize = 0;
|
||||
while (adopted < maximum_devices and bringUpVolume()) : (adopted += 1) {}
|
||||
}
|
||||
|
||||
/// A filesystem announces itself for the volume it was spawned to serve. Reply
|
||||
@@ -335,26 +422,26 @@ fn pollTick() void {
|
||||
/// badge) as the call's returned capability. No channel means the volume is not
|
||||
/// ready — the filesystem retries.
|
||||
fn onHello(_: void, invocation: Invocation(volume_manager_protocol.Hello), _: Answer(void)) isize {
|
||||
const v = volume orelse return 0; // not probed yet — retryable, no cap
|
||||
if (invocation.target != v.id) return 0; // unknown volume — retryable
|
||||
const v = volumeById(invocation.target) orelse return 0; // not probed yet — retryable, no cap
|
||||
if (invocation.sender != v.filesystem_pid) {
|
||||
// Not the filesystem we spawned for this volume. Refuse: only the
|
||||
// confined filesystem gets the channel.
|
||||
// Not the filesystem we spawned for this volume. Refuse: only the confined
|
||||
// filesystem gets the channel.
|
||||
std.log.info("refused hello for volume {d} from process {d}", .{ invocation.target, invocation.sender });
|
||||
return -envelope.EPERM;
|
||||
}
|
||||
service.replyWithCapability(v.storage.endpoint);
|
||||
const dev = deviceById(v.device_id) orelse return 0; // its device left — retryable
|
||||
service.replyWithCapability(dev.channel.endpoint);
|
||||
std.log.info("handed volume {d} to pid {d}", .{ v.id, invocation.sender });
|
||||
return 0;
|
||||
}
|
||||
|
||||
/// Answer a `volumes` query with the mounted volume's descriptor — its id (its
|
||||
/// mount path is /volumes/<id> unless overridden), its actual mount path, and
|
||||
/// its display label. This is how a shell or file manager reads a volume's
|
||||
/// friendly name: software keys on the id, a UI shows the label. An empty reply
|
||||
/// Answer a `volumes` query with a mounted volume's descriptor — its id (its
|
||||
/// mount path is /volumes/<id> unless overridden), its actual mount path, and its
|
||||
/// display label. Software keys on the id; a UI shows the label. Returns the first
|
||||
/// mounted volume for now; a full enumerate is a later refinement. Empty reply
|
||||
/// means no volume is mounted.
|
||||
fn onVolumes(_: void, _: Invocation(volume_manager_protocol.Volumes), answer: Answer(void)) isize {
|
||||
const v = volume orelse return 0;
|
||||
const v = firstUsedVolume() orelse return 0;
|
||||
var id_buf: [volume_map.id_maximum]u8 = undefined;
|
||||
const info = volume_manager_protocol.VolumeInfo{
|
||||
.id = volume_map.idString(v.identity, &id_buf),
|
||||
@@ -381,9 +468,9 @@ fn readConfig(path: []const u8, buf: []u8) usize {
|
||||
defer file.close();
|
||||
var used: usize = 0;
|
||||
while (used < buf.len) {
|
||||
const n = file.read(buf[used..]) orelse break;
|
||||
if (n == 0) break;
|
||||
used += n;
|
||||
const nn = file.read(buf[used..]) orelse break;
|
||||
if (nn == 0) break;
|
||||
used += nn;
|
||||
}
|
||||
return used;
|
||||
}
|
||||
@@ -427,23 +514,22 @@ fn onNotification(badge: u64) void {
|
||||
// reclaimed by the driver on the same death; the respawn confines afresh.
|
||||
if (got.isChildExit()) {
|
||||
const dead = got.childProcessId();
|
||||
const v = &(volume orelse return);
|
||||
if (v.filesystem_pid != dead) return;
|
||||
const v = volumeByPid(dead) orelse return;
|
||||
v.filesystem_pid = 0;
|
||||
const reason = process.exitReason(dead) orelse .fault;
|
||||
if (reason == .exited) {
|
||||
std.log.info("filesystem for volume {d} exited cleanly; not restarting", .{v.id});
|
||||
return;
|
||||
}
|
||||
const alive = time.clock() -| fs_spawn_ns;
|
||||
fs_restarts = if (alive < fast_death_ns) fs_restarts + 1 else 1;
|
||||
if (fs_restarts >= crash_loop_cap) {
|
||||
fs_failed = true;
|
||||
const alive = time.clock() -| v.spawn_ns;
|
||||
v.restarts = if (alive < fast_death_ns) v.restarts + 1 else 1;
|
||||
if (v.restarts >= crash_loop_cap) {
|
||||
v.failed = true;
|
||||
std.log.info("filesystem for volume {d} is failing repeatedly; giving up", .{v.id});
|
||||
return;
|
||||
}
|
||||
std.log.info("filesystem for volume {d} died ({s}); restarting", .{ v.id, @tagName(reason) });
|
||||
armRestart();
|
||||
armRestart(v);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -816,6 +816,47 @@ CASES = [
|
||||
"expect": r"volume-manager: volume 0x0*12345678 -> \S+ \(pid \d+\), lba \d+, \d+ blocks"
|
||||
r"[\s\S]*volume-manager: handed volume \d+ to pid \d+",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# S3 multi-volume: a SECOND usb-storage device (a generated data volume, serial
|
||||
# da7a0001, an empty FAT with no /system) plugged in beside the boot volume.
|
||||
# Proves the volume manager adopts BOTH devices and spawns a confined fat per
|
||||
# volume, each mounted at its own CONTENT id-path (/volumes/fat-<serial>); and
|
||||
# that boot-volume detection is by content — only the volume that carries
|
||||
# /system backs /system/configuration, while the data volume mounts at its
|
||||
# id-path alone. Against the pre-S3 one-device/one-volume manager the data
|
||||
# volume never mounts, so the da7a0001 lookaheads fail (toggle-demonstrated by
|
||||
# checking out the step-2 volume-manager.zig).
|
||||
{"name": "two-volumes",
|
||||
"build_case": "fat-mount",
|
||||
"smp": 4,
|
||||
"timeout": 150,
|
||||
"data_volume": {"serial": "DA7A0001", "label": "DATAVOL", "size_mib": 64},
|
||||
"expect": r"(?s)(?=.*volume-manager: volume 0x0*12345678 -> )"
|
||||
r"(?=.*volume-manager: volume 0x0*da7a0001 -> )"
|
||||
r"(?=.*fat: mounted /volumes/fat-12345678)"
|
||||
r"(?=.*fat: mounted /volumes/fat-da7a0001)"
|
||||
r"(?=.*carries the system tree)"
|
||||
r"(?=.*data volume; mounted at /volumes/fat-da7a0001)",
|
||||
"fail": r"data volume; mounted at /volumes/fat-12345678|DANOS-TEST-RESULT: FAIL"},
|
||||
# S3 shared-channel multi-volume: ONE usb-storage device carrying an MBR with
|
||||
# TWO FAT partitions (da7a0001 at lba 2048, da7a0002 at lba 83968). allVolumes
|
||||
# walks the table and the manager spawns a confined fat per partition on the
|
||||
# SAME block channel, each clamped to its own LBA range (usb-storage's
|
||||
# per-badge range table) — the path a pair of single-volume sticks (the
|
||||
# two-volumes case) does NOT exercise. The two mount lines sit at two DISTINCT
|
||||
# non-zero base_lbas on one device. Against the pre-uncap allVolumes (S3 step
|
||||
# 2, capped to one partition) only da7a0001 mounts, so the da7a0002 lookaheads
|
||||
# fail.
|
||||
{"name": "partitioned-volume",
|
||||
"build_case": "fat-mount",
|
||||
"smp": 4,
|
||||
"timeout": 150,
|
||||
"data_volume": {"partitions": [{"serial": "DA7A0001", "size_mib": 40},
|
||||
{"serial": "DA7A0002", "size_mib": 40}]},
|
||||
"expect": r"(?s)(?=.*volume 0x0*da7a0001 -> \S+ \(pid \d+\), lba 2048, )"
|
||||
r"(?=.*volume 0x0*da7a0002 -> \S+ \(pid \d+\), lba 83968, )"
|
||||
r"(?=.*fat: mounted /volumes/fat-da7a0001)"
|
||||
r"(?=.*fat: mounted /volumes/fat-da7a0002)",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the
|
||||
# fat-test client, after listing, makes a directory, writes+reads a file inside
|
||||
# it, then removes the file, exercising the whole VFS -> fat mutation path.
|
||||
@@ -1422,6 +1463,30 @@ def run_case(arch, case):
|
||||
cmd[cmd.index("-m") + 1] = case["mem"]
|
||||
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
|
||||
cmd += case["qemu_extra"]
|
||||
# A multi-volume case attaches a second usb-storage device backed by a freshly
|
||||
# GENERATED data volume: a distinct-serial FAT32 with no /system tree, so the
|
||||
# volume manager mounts it at its own id-path and the fat process marks it a
|
||||
# data volume (never a system volume). Regenerated per run — no image is
|
||||
# committed to the tree (the user keeps the boot files copyable, not baked in).
|
||||
if case.get("data_volume"):
|
||||
dv = case["data_volume"]
|
||||
data_img = os.path.join(WORK, "data-volume.img")
|
||||
if dv.get("partitions"):
|
||||
# One device, an MBR with several FAT partitions: several volumes share
|
||||
# ONE block channel, each confined to its own LBA range.
|
||||
gen = [sys.executable, os.path.join(REPO, "tools", "make-partitioned-image.py"), data_img]
|
||||
for part in dv["partitions"]:
|
||||
gen += [part["serial"], str(part.get("size_mib", 40))]
|
||||
else:
|
||||
# One device, one bare FAT volume.
|
||||
gen = [sys.executable, os.path.join(REPO, "tools", "make-fat-image.py"),
|
||||
"--serial", dv["serial"], "--label", dv.get("label", "DATAVOL"),
|
||||
data_img, str(dv.get("size_mib", 64))]
|
||||
subprocess.run(gen, check=True, stdout=subprocess.DEVNULL)
|
||||
cmd += [
|
||||
"-drive", f"if=none,id=datausb,format=raw,file={data_img}",
|
||||
"-device", "usb-storage,bus=xhci.0,port=4,drive=datausb,removable=on,id=datastorage",
|
||||
]
|
||||
# A QMP control socket, always present (additive): how a case's `qmp_after`
|
||||
# hook injects host-side events into the guest mid-run. Kept under a short temp
|
||||
# dir, not WORK: a unix socket path is capped at ~104 bytes (sun_path), and a
|
||||
|
||||
@@ -194,7 +194,6 @@ system/kernel/process.zig:write_buffer
|
||||
system/kernel/scheduler.zig:ipc_maximum_handles
|
||||
system/kernel/scheduler.zig:maximum_space_mappings
|
||||
system/kernel/vfs.zig:maximum_directories
|
||||
system/kernel/vfs.zig:maximum_mounts
|
||||
system/kernel/vfs.zig:maximum_prefix
|
||||
system/kernel/vfs.zig:maximum_rewrite
|
||||
system/services/acpi/acpi.zig:blocks
|
||||
|
||||
+31
-7
@@ -47,8 +47,14 @@ def fat32_geometry(total_sectors):
|
||||
|
||||
|
||||
class Fat32Image:
|
||||
def __init__(self, total_sectors):
|
||||
def __init__(self, total_sectors, volume_id=0x12345678, label="DANOS"):
|
||||
self.total_sectors = total_sectors
|
||||
# The FAT volume serial (its content identity — the /volumes/fat-<id>
|
||||
# mount path danos derives from it) and the display label. A second image
|
||||
# needs a distinct serial so its id-path does not collide with the boot
|
||||
# volume's.
|
||||
self.volume_id = volume_id & 0xFFFFFFFF
|
||||
self.label = label
|
||||
self.fat_size, self.cluster_count = fat32_geometry(total_sectors)
|
||||
if self.cluster_count < 65525:
|
||||
sys.exit(f"error: image too small for FAT32 ({self.cluster_count} clusters "
|
||||
@@ -146,8 +152,8 @@ class Fat32Image:
|
||||
0x80, # drive number
|
||||
0, # reserved
|
||||
0x29, # extended boot signature
|
||||
0x12345678, # volume id
|
||||
b"DANOS ", # volume label
|
||||
self.volume_id, # volume id
|
||||
self.label.encode("ascii", "replace")[:11].ljust(11, b" "), # volume label
|
||||
b"FAT32 ", # filesystem type
|
||||
)
|
||||
sector[510] = 0x55
|
||||
@@ -276,9 +282,9 @@ def build_tree(pairs):
|
||||
return root
|
||||
|
||||
|
||||
def build(out_path, size_mib, pairs):
|
||||
def build(out_path, size_mib, pairs, volume_id=0x12345678, label="DANOS"):
|
||||
total_sectors = size_mib * 1024 * 1024 // SECTOR
|
||||
image = Fat32Image(total_sectors)
|
||||
image = Fat32Image(total_sectors, volume_id, label)
|
||||
tree = build_tree(pairs)
|
||||
write_directory(image, 2, tree, 0, True)
|
||||
with open(out_path, "wb") as handle:
|
||||
@@ -350,14 +356,32 @@ def main(argv):
|
||||
if len(argv) == 3 and argv[1] == "--verify":
|
||||
verify(argv[2])
|
||||
return 0
|
||||
# Optional flags ahead of the positionals: --serial <hex> sets the FAT volume
|
||||
# id (the /volumes/fat-<id> content identity), --label <name> its display
|
||||
# label. A second FAT image passes a distinct --serial so its id-path cannot
|
||||
# collide with the boot volume's.
|
||||
argv = list(argv)
|
||||
volume_id = 0x12345678
|
||||
label = "DANOS"
|
||||
i = 1
|
||||
while i < len(argv):
|
||||
if argv[i] == "--serial" and i + 1 < len(argv):
|
||||
volume_id = int(argv[i + 1], 16)
|
||||
del argv[i:i + 2]
|
||||
elif argv[i] == "--label" and i + 1 < len(argv):
|
||||
label = argv[i + 1]
|
||||
del argv[i:i + 2]
|
||||
else:
|
||||
i += 1
|
||||
if len(argv) < 3 or (len(argv) - 3) % 2 != 0:
|
||||
sys.exit("usage: make-fat-image.py <out.img> <size-MiB> [<dest> <host>]...\n"
|
||||
sys.exit("usage: make-fat-image.py [--serial <hex>] [--label <name>] "
|
||||
"<out.img> <size-MiB> [<dest> <host>]...\n"
|
||||
" make-fat-image.py --verify <out.img>")
|
||||
out_path = argv[1]
|
||||
size_mib = int(argv[2])
|
||||
rest = argv[3:]
|
||||
pairs = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)]
|
||||
build(out_path, size_mib, pairs)
|
||||
build(out_path, size_mib, pairs, volume_id, label)
|
||||
return 0
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Assemble an MBR-partitioned disk image from N FAT32 partitions — the danos
|
||||
multi-volume test disk.
|
||||
|
||||
Pure Python 3 stdlib (no mtools / parted). Each partition is a real FAT32
|
||||
filesystem produced by make-fat-image.py, laid out behind a classic MBR so the
|
||||
danos partition prober (partition.allVolumes) walks the table and the volume
|
||||
manager spawns one confined filesystem per partition — several volumes sharing
|
||||
ONE block channel, each clamped to its own LBA range. That shared-channel,
|
||||
per-partition path is what a single stick with two partitions exercises and a
|
||||
pair of single-volume sticks does not.
|
||||
|
||||
make-partitioned-image.py <out.img> [<serial-hex> <size-MiB>]...
|
||||
|
||||
Each partition is an empty FAT32 with the given volume serial (its /volumes/
|
||||
fat-<serial> content id). Partitions are 1-MiB aligned; the MBR marks each
|
||||
type 0x0C (FAT32 LBA). At most four (an MBR holds four primaries).
|
||||
"""
|
||||
|
||||
import os
|
||||
import struct
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
|
||||
SECTOR = 512
|
||||
ALIGN = 2048 # sectors (1 MiB) — standard partition alignment, and the gap the MBR sits in
|
||||
MBR_TYPE_FAT32_LBA = 0x0C
|
||||
MAX_PRIMARY_PARTITIONS = 4
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
|
||||
|
||||
def align_up(sectors, to=ALIGN):
|
||||
return (sectors + to - 1) // to * to
|
||||
|
||||
|
||||
def main(argv):
|
||||
if len(argv) < 4 or (len(argv) - 2) % 2 != 0:
|
||||
sys.exit("usage: make-partitioned-image.py <out.img> [<serial-hex> <size-MiB>]...")
|
||||
out_path = argv[1]
|
||||
specs = [(argv[i], int(argv[i + 1])) for i in range(2, len(argv), 2)]
|
||||
if len(specs) > MAX_PRIMARY_PARTITIONS:
|
||||
sys.exit(f"error: an MBR holds at most {MAX_PRIMARY_PARTITIONS} primary partitions")
|
||||
|
||||
# Generate each partition's FAT32 image, then place it at its aligned start.
|
||||
partitions = [] # (start_sector, sector_count, bytes)
|
||||
cursor = ALIGN # leave the first 1 MiB for the MBR + alignment gap
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
for idx, (serial, size_mib) in enumerate(specs):
|
||||
part_path = os.path.join(tmp, f"p{idx}.img")
|
||||
subprocess.run(
|
||||
[sys.executable, os.path.join(HERE, "make-fat-image.py"),
|
||||
"--serial", serial, "--label", f"DATA{idx}",
|
||||
part_path, str(size_mib)],
|
||||
check=True, stdout=subprocess.DEVNULL)
|
||||
with open(part_path, "rb") as handle:
|
||||
data = handle.read()
|
||||
count = len(data) // SECTOR
|
||||
partitions.append((cursor, count, data))
|
||||
cursor = align_up(cursor + count)
|
||||
|
||||
total_sectors = cursor
|
||||
disk = bytearray(total_sectors * SECTOR)
|
||||
# The MBR: a disk signature, one partition entry per FAT partition, 0x55AA.
|
||||
# No boot code (this disk is data, never booted); danos's mount() sees the
|
||||
# signature but no BPB at LBA 0 and takes the MBR-walk path.
|
||||
struct.pack_into("<I", disk, 440, 0x0D05DA05) # arbitrary but fixed disk signature
|
||||
for idx, (start, count, data) in enumerate(partitions):
|
||||
entry = 446 + idx * 16
|
||||
disk[entry + 0] = 0x00 # not bootable
|
||||
disk[entry + 1:entry + 4] = b"\xFE\xFF\xFF" # CHS start (LBA-aware tools ignore)
|
||||
disk[entry + 4] = MBR_TYPE_FAT32_LBA
|
||||
disk[entry + 5:entry + 8] = b"\xFE\xFF\xFF" # CHS end
|
||||
struct.pack_into("<I", disk, entry + 8, start) # start LBA
|
||||
struct.pack_into("<I", disk, entry + 12, count) # sector count
|
||||
disk[start * SECTOR:start * SECTOR + len(data)] = data
|
||||
disk[510] = 0x55
|
||||
disk[511] = 0xAA
|
||||
|
||||
with open(out_path, "wb") as handle:
|
||||
handle.write(disk)
|
||||
print(f"make-partitioned-image: wrote {out_path} "
|
||||
f"({total_sectors * SECTOR // (1024 * 1024)} MiB, {len(partitions)} partitions)")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main(sys.argv))
|
||||
Reference in New Issue
Block a user