8 Commits
Author SHA1 Message Date
Daniel Samson 0b25cd2c94 docs: multi-volume is built — storage architecture + rationale (S3)
Flip the storage docs from "one FAT volume today" to the built state:
the manager adopts every device, probes each device's whole partition
table, and spawns one range-confined FAT per volume — several volumes
across several devices, or several partitions on one device's channel.
The boot volume is identified by content; a filesystem installs the
/system rewrites only when it resolves /system/configuration on its own
media. No filesystem binds a shared service name any more — clients route
through the kernel mount table.

Correct two now-stale claims in the rationale, in the honest direction:
"multi-volume providers are reserved, not implemented" becomes built
(multi-namespace-per-provider is the untried NVMe case); and the cloned-
duplicate "second mounts suffixed" was never built — S3 mounts both,
they collide on the shared content id-path (last wins), each boot claim
logged, and distinguishing them is S4 arbitration.
2026-08-10 02:35:43 +01:00
Daniel Samson d59279422e test: partitioned-image tool + shared-channel multi-volume proof (S3)
The two-volumes case proves multiple DEVICES; this proves multiple
volumes on ONE device. make-partitioned-image.py lays several FAT32
partitions (each from make-fat-image) behind a classic MBR; the new
partitioned-volume case attaches one such disk (two partitions,
da7a0001 at lba 2048, da7a0002 at lba 83968) as a single usb-storage
device.

partition.allVolumes walks the table and the manager spawns a confined
fat per partition on the SAME block channel, each clamped to its own LBA
range by usb-storage's per-badge range table — so partition B's fat
cannot read partition A's blocks. The case asserts two mount lines at
two distinct non-zero base_lbas on one device; against the pre-uncap
allVolumes (S3 step 2, capped to one partition) only da7a0001 mounts.

No image is committed — the disk is generated per run. Full suite
130/130 (128 + two-volumes + partitioned-volume).
2026-08-10 02:35:32 +01:00
Daniel Samson bf9f8560c6 fat/harness: filesystems coexist without the shared vfs name; two-volume proof (S3)
A second usb-storage device (a generated data volume, serial da7a0001,
an empty FAT with no /system) plugged in beside the boot volume: the
volume manager adopts both devices and spawns a confined fat per volume,
each mounted at its own content id-path.

The test surfaced a real coexistence bug. Every filesystem bound the
single "vfs" contract name under /protocol; the second volume's fat lost
the race, service.run refused-and-exited on the held name, and that
volume never mounted. Clients don't reach filesystems by that name —
fs_resolve routes a path to its backing endpoint through the kernel
mount table by prefix — and nothing consumes "vfs", so the fix is to
bind no shared name: the harness's service_name now defaults to null.
This is the "this fades" the harness comment anticipated for the
volume-manager era; a filesystem's endpoint still serves as its mount
backend without a name.

fat logs "is a data volume" for the non-system branch so the test can
positively assert content-based detection. make-fat-image gains
--serial/--label (default unchanged) so a second image gets a distinct
id-path; the data image is generated per run, never committed. The case
fails against the pre-fix harness (the data volume's fat exits on the
refused bind) — toggle-demonstrated.

Full suite 129/129 (128 + two-volumes); the single-volume path is
unaffected by dropping the vestigial name bind.
2026-08-10 02:10:06 +01:00
Daniel Samson da7dcce64e kernel/vfs: raise the mount ceiling for N volumes; refuse (not drop) a full table (S3)
Multi-volume makes the mount table the bottleneck: each volume installs
one id-path mount and the system volume two FHS rewrites, so at the
volume manager's maximum_volumes (16) the old cap of 8 is far too low.
Raise maximum_mounts to 32 (headroom over the ~20-mount worst case) and
declare its bounds block; drop it from the bounds allowlist.

Fix a latent bug the higher pressure would expose: installMount silently
dropped a mount when the table was full, and mountBackend returned true
anyway — a full table was reported as a successful mount. installMount
now returns whether it placed the mount, and mountBackend propagates a
false so the mounting filesystem's harness logs "could not mount
<prefix>". At-limit is now a refusal that is observed, not a silent
success. (The full-table path has no host unit test: vfs.zig's tests are
not wired into the host aggregate — its import graph reaches the
freestanding kernel — so the correction rests on the propagated return
and the truthful bounds block.)
2026-08-10 01:52:47 +01:00
Daniel Samson 9750db14da fat: install the FHS boot rewrites only on the system volume (S3)
A volume backs /system/configuration and /system/logs only when it
actually carries the /system tree — decided by content (resolve
/system/configuration on its own media at mount), not by spawn order.
The boot volume takes the branch and installs the two rewrites; a data
volume resolves null, mounts only at its id-path, and never shadows the
running system's config or logs with a dead mount.

This retires the "resolve-at-bring-up quirk" the single-volume step
deferred around: that diagnosis was wrong. A boot probe confirmed
resolve() works the instant mount() returns — /system, /system/
configuration, /system/kernel, /system/services all resolve at bring-up
(mount reads LBA 0 through the same block path, so a directory read
cannot fail where the boot-sector read succeeded). No deferral needed.

Behavior-preserving on the single boot volume (it carries the system
tree, so it still installs all three mounts): suite stays 128/128.
2026-08-10 01:37:32 +01:00
Daniel Samson b2a5a0a3c6 volume-manager: adopt every device, a filesystem per partition (S3)
Lift the one-device/one-volume cap. bringUpVolume now probes the whole
partition table (allVolumes uncapped) and spawns a confined filesystem
per volume; pollTick loops it to adopt every present, not-yet-adopted
device each tick.

The subtlety is adopt-once-and-keep: a device is recorded in the table
the first time it is seen and kept until it leaves the tree, even when it
carries no servable volume or its geometry cannot be read. Dropping an
unservable device would make openAnyStorage hand back the same one every
tick and starve the devices behind it; keeping it lets the scan advance
past it. A genuine removal frees the slot; a re-insert (fresh device id)
is probed anew.

The boot image is a single bare-FAT volume, so the full suite is
unchanged at 128/128.
2026-08-10 01:26:47 +01:00
Daniel Samson 7efe7b72d8 volume-manager: N-volume device+volume tables, behavior-preserving (S3)
Replace the single `var volume: ?Volume` and file-global supervision
state with two fixed tables: devices[maximum_devices] owning each adopted
block channel once, and volumes[maximum_volumes] each carrying its own
identity, id, mount prefix, and supervision fields (restarts, spawn_ns,
failed, restart_pending, restart_due_ns). A monotonic next_volume_id
never reuses ids, so a stale hello can't address the wrong child.

Lookups (deviceById, volumeById, volumeByPid, firstUsedVolume) and
claims (claimDevice, claimVolumeIndex) replace the ad-hoc singletons.
pollTick reconciles devices first (removeDevice drops their volumes),
then per-volume restarts, then idle bring-up.

This step stays one-device/one-volume on purpose: bringUpVolume adopts
the first device and caps allVolumes to a single partition, so behavior
is identical and the full suite stays 128/128. Uncapping and adopt-all
land next.
2026-08-10 01:13:52 +01:00
Daniel Samson d4b544d66b volume-manager: partition.allVolumes — every partition, not just the first (S3)
The multi-volume enabler. allVolumes(reader, device_blocks, out) appends every
volume on the device to the caller's buffer and returns the count: GPT
enumerates all valid entries (gptFirstVolume becomes gptAllVolumes), the MBR walk
collects all fitting partitions, and a bare FAT is the single whole-device volume
— each with the same per-entry overflow-safe range validation (the confinement
invariant the driver's clamp rests on) and fatIdentity-over-disk-signature
preference. firstVolume is now the one-element case of allVolumes, so the S1
behavior and its ten tests are unchanged. New host test: a two-partition MBR
yields two volumes with distinct identities (index 0 vs 1); it FAILS when
allVolumes is capped to one (the old firstVolume semantics), passes at 11/11.
2026-08-10 00:57:56 +01:00
11 changed files with 604 additions and 242 deletions
@@ -11,15 +11,21 @@
> identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map: > identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:
> `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional > `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional
> override), a volume's mount path IS its content id (`/volumes/<id>`), with the > override), a volume's mount path IS its content id (`/volumes/<id>`), with the
> label as display metadata a `volumes` query returns. **Still pending**: the > label as display metadata a `volumes` query returns. Multi-volume is **built**:
> `filesystem UUID` rung (needs a non-FAT engine), multi-volume (one FAT volume > the manager adopts every storage device, probes each device's whole partition
> today; fat's boot rewrites are unconditional until S3 makes them > table, and spawns one range-confined FAT per volume — several volumes across
> content-conditional), the volume manager *consuming* `medium_changed` (removal > several devices, or several partitions sharing one device's channel — each at
> is detected by device-presence polling; the event is published but only a > its own `/volumes/<id>` path with its own supervision. The boot volume is
> card-reader medium change needs the subscription), and the remount-on-replug > identified by **content** (a volume backs `/system/configuration` + `/system/logs`
> only when it resolves `/system/configuration` on its own media), so it works as
> any partition of any device. **Still pending**: the `filesystem UUID` rung and
> a second engine (exFAT, S4); the volume manager *consuming* `medium_changed`
> (removal is detected by device-presence polling; the event is published but only
> a card-reader medium change needs the subscription); the remount-on-replug
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller > end-to-end (the logic is in place; QEMU can't re-present the boot-controller
> device, so it is bench-verified). A few markers below are left where a duty is > device, so it is bench-verified); and arbitration when two volumes both resolve
> still pending. > the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
> markers below are left where a duty is still pending.
## The model ## The model
@@ -121,12 +127,16 @@ its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the *(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there, plus volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there. It
the two `/system` hierarchy rewrites it installs in place (unconditional this installs the two `/system` hierarchy rewrites (`/system/configuration`,
increment; S3 makes them content-conditional across volumes). It no longer `/system/logs`) only when it is the boot volume — decided by **content**: it
resolves `/system/configuration` on its own media at mount, so a data volume
mounts at its id-path alone and never shadows the running system. It no longer
self-acquires a volume — the V3b flip made it receive its volume id and block self-acquires a volume — the V3b flip made it receive its volume id and block
channel from the volume manager, consistent with "it never discovers devices" channel from the volume manager, consistent with "it never discovers devices"
above. above. Because several volumes now serve at once, no filesystem binds a shared
service name; clients reach each through the kernel mount table (`fs_resolve`
routes by prefix to the backing endpoint).
**Kernel** (mechanism only): the mount table routes paths to backend **Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart endpoints — resolve and redirect, never data. Remount-replace is the restart
@@ -145,9 +145,13 @@ matrix-proven shape; genuinely open.
**The pressure points, honestly:** **The pressure points, honestly:**
1. **Multi-volume providers are reserved, not implemented.** The volume 1. **Multi-volume is built; multi-namespace-per-provider is untried.** The
manager flow assumes one provider, one volume; NVMe namespaces make volume manager adopts every device and spawns one range-confined FAT per
endpoint-per-volume real work with hardware demanding it. partition — several volumes across several devices, or several partitions
sharing one device's channel, both proven on USB. What is untried is a single
provider exposing several volumes as *namespaces* (NVMe): the endpoint and
per-badge range machinery generalizes, but no such driver exists yet to
exercise it.
2. **The current transport will bottleneck NVMe.** Synchronous call/reply, 2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
one operation in flight, one bounce buffer — fine for a USB2 stick, one operation in flight, one bounce buffer — fine for a USB2 stick,
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
@@ -233,13 +237,15 @@ matrix-proven shape; genuinely open.
Consequences, each mechanical once identity keys the map: **moving a drive Consequences, each mechanical once identity keys the map: **moving a drive
to a different port changes nothing** — same identity, same mount point, to a different port changes nothing** — same identity, same mount point,
whether USB port, hub depth, SATA port, or a stick that left as USB and whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (a map lookup returned in a SATA dock; **replug remounts at the same path** (the id-path is
once the map exists; today's single volume re-probes and remounts at the content-derived, so a volume returns to `/volumes/<id>` wherever it reappears;
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the remount-on-replug end-to-end is bench-verified, not QEMU-tested, because QEMU
recorded identity of the volume carrying `/system/configuration`, findable can't re-present the boot-controller device); **the boot volume** is the volume
on any port; and **duplicate identity is a policy case, not a surprise** — that resolves `/system/configuration` on its own media, findable on any port or
two cloned sticks at once: first keeps the mapped name, second mounts partition; and **duplicate identity is a known S4 gap** — two cloned sticks
suffixed and is logged loudly, never silently shadowed. Unknown identities share one content id, so today they collide on `/volumes/<id>` (the kernel
remount-replaces; the last wins) and each boot-volume claim is logged loudly.
Distinguishing them with a suffix is arbitration, deferred to S4. Unknown identities
mount under a derived name (sanitized label, else generated) at mount under a derived name (sanitized label, else generated) at
`/volumes/<name>` — the hierarchy's documented home for attached media, `/volumes/<name>` — the hierarchy's documented home for attached media,
which stands: `/system` is what danos IS; attached media is what it isn't. which stands: `/system` is what danos IS; attached media is what it isn't.
+8 -4
View File
@@ -61,10 +61,14 @@ pub fn Server(comptime Engine: type) type {
/// caller does the filesystem-specific bring-up (find the block /// caller does the filesystem-specific bring-up (find the block
/// device, set up DMA, mount the engine) and returns a `Volume`. /// device, set up DMA, mount the engine) and returns a `Volume`.
bringUp: *const fn (endpoint: ipc.Handle) ?Volume, bringUp: *const fn (endpoint: ipc.Handle) ?Volume,
/// The vfs contract name to bind. A filesystem serving one volume /// A contract name to bind under /protocol, or null to bind none. In
/// binds "vfs" today; the volume-manager era hands each per-volume /// the volume-manager era every filesystem is a per-volume process and
/// process its own establishment and this fades. /// clients reach it through the kernel mount table — fs_resolve routes
service_name: ?[]const u8 = "vfs", /// a path to its backing endpoint by prefix — so no filesystem binds a
/// shared name. Two volumes would collide on one: the second's bind is
/// refused and service.run would exit, so its volume never mounts. The
/// endpoint still serves as the mount backend without a name.
service_name: ?[]const u8 = null,
}; };
// --- the harness's own state, one set per instantiation --------------- // --- the harness's own state, one set per instantiation ---------------
+25 -6
View File
@@ -55,7 +55,17 @@ fn tokenIndex(t: u64) u64 {
// --- the mount table --------------------------------------------------------- // --- the mount table ---------------------------------------------------------
pub const maximum_mounts = 8; /// bound: prefixes mounted in the kernel VFS table at once
/// decided-by: ours
/// protects: the `mounts` table below
/// at-limit: refuse - installMount returns false and mountBackend propagates it;
/// the mounting filesystem's harness logs "could not mount <prefix>" and the
/// mount simply does not exist (no silent success). Budget: the initrd's
/// top-level dirs (/system, /test) plus one id-path mount per volume and the
/// system volume's two FHS rewrites — a few over the volume manager's
/// maximum_volumes (16); 32 leaves headroom.
/// observed-by: the harness "file-system: could not mount <prefix>" ring line
pub const maximum_mounts = 32;
const maximum_prefix = 64; const maximum_prefix = 64;
const maximum_rewrite = 32; const maximum_rewrite = 32;
@@ -156,7 +166,10 @@ pub fn setInitialRamdisk(image: []const u8) void {
for (directories[0..directory_count], 0..) |*d, index| { for (directories[0..directory_count], 0..) |*d, index| {
const parent = parentOf(d.slice()); const parent = parentOf(d.slice());
d.parent = directoryIndex(parent) orelse index; d.parent = directoryIndex(parent) orelse index;
if (parent.len == 1) installMount(d.slice(), .kernel_initrd, null, ""); // Boot-time install of one mount per top-level initrd dir (/system, /test):
// provably few, far under maximum_mounts, so a full table here is
// impossible — but discard the result explicitly rather than assume it.
if (parent.len == 1) _ = installMount(d.slice(), .kernel_initrd, null, "");
} }
} }
@@ -167,7 +180,12 @@ fn directoryIndex(path: []const u8) ?usize {
return null; return null;
} }
fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, rewrite: []const u8) void { /// Install (or remount-replace) a prefix. Returns false when the table is full
/// and no slot could be claimed — the caller must surface that, never report a
/// dropped mount as success. A remount of an already-mounted prefix reuses its
/// slot and always succeeds; a /protocol remount is refused-as-noop (returns
/// true: the first mount stands, nothing is dropped).
fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, rewrite: []const u8) bool {
// Remount replaces: a restarted backend re-mounts its prefix. // Remount replaces: a restarted backend re-mounts its prefix.
var slot: ?*Mount = null; var slot: ?*Mount = null;
for (&mounts) |*m| { for (&mounts) |*m| {
@@ -176,19 +194,20 @@ fn installMount(prefix: []const u8, kind: MountKind, backend: ?*ipc.Endpoint, re
// restarted FAT retakes /volumes/usb; letting it retake /protocol // restarted FAT retakes /volumes/usb; letting it retake /protocol
// would hand the whole naming layer to whoever asked second. // would hand the whole naming layer to whoever asked second.
// First mount wins, and init (PID 1) is always first. // First mount wins, and init (PID 1) is always first.
if (std.mem.eql(u8, prefix, protocol_root)) return; if (std.mem.eql(u8, prefix, protocol_root)) return true;
if (m.backend) |old| ipc.dropRef(old); if (m.backend) |old| ipc.dropRef(old);
slot = m; slot = m;
break; break;
} }
if (slot == null and !m.used) slot = m; if (slot == null and !m.used) slot = m;
} }
const m = slot orelse return; const m = slot orelse return false;
m.* = .{ .used = true, .kind = kind, .backend = backend }; m.* = .{ .used = true, .kind = kind, .backend = backend };
@memcpy(m.prefix[0..prefix.len], prefix); @memcpy(m.prefix[0..prefix.len], prefix);
m.prefix_len = prefix.len; m.prefix_len = prefix.len;
@memcpy(m.rewrite[0..rewrite.len], rewrite); @memcpy(m.rewrite[0..rewrite.len], rewrite);
m.rewrite_len = rewrite.len; m.rewrite_len = rewrite.len;
return true;
} }
// --- resolve ----------------------------------------------------------------- // --- resolve -----------------------------------------------------------------
@@ -406,7 +425,7 @@ pub fn mountBackend(prefix: []const u8, backend: *ipc.Endpoint, rewrite: []const
if (!isInitrdCarveOut(prefix)) return false; if (!isInitrdCarveOut(prefix)) return false;
} }
} }
installMount(prefix, .backend, backend, rewrite); if (!installMount(prefix, .backend, backend, rewrite)) return false; // table full
for (&mounts) |*m| { for (&mounts) |*m| {
if (m.used and std.mem.eql(u8, m.prefixSlice(), prefix)) m.owner = owner; if (m.used and std.mem.eql(u8, m.prefixSlice(), prefix)) m.owner = owner;
} }
+18 -9
View File
@@ -173,16 +173,25 @@ fn fatBringUp(endpoint: ipc.Handle) ?Harness.Volume {
}; };
std.log.info("mounted FAT ({s}, {d} clusters, partition lba {d})", .{ @tagName(filesystem.geometry.fat_type), filesystem.geometry.cluster_count, filesystem.base_lba }); std.log.info("mounted FAT ({s}, {d} clusters, partition lba {d})", .{ @tagName(filesystem.geometry.fat_type), filesystem.geometry.cluster_count, filesystem.base_lba });
// The volume mounts at its id-path (argv[2]), plus the two FHS rewrites so // Every volume mounts at its own id-path (argv[2]). The boot/system volume —
// hierarchy paths (the logger's /system/logs) stay decoupled from which // the one carrying the /system tree — ADDITIONALLY installs the two FHS
// volume backs them. This single-volume increment's one volume IS the boot // rewrites, so hierarchy paths (config reads, the logger's persistent
// volume, so it installs both unconditionally; S3 (multi-volume) makes the // /system/logs) stay decoupled from which volume backs them. Detection is by
// rewrites content-conditional — installed only by whichever volume carries // CONTENT, not spawn order: a volume is the system volume iff /system/
// the system, decided by content, not order. // configuration resolves on its own media. A data volume has no /system, so it
// mounts only at its id-path and never shadows the running system's config or
// logs with a dead mount.
mount_specs[0] = .{ .prefix = volume_mount_prefix }; mount_specs[0] = .{ .prefix = volume_mount_prefix };
mount_specs[1] = .{ .prefix = "/system/configuration", .rewrite = "/system/configuration" }; var mount_count: usize = 1;
mount_specs[2] = .{ .prefix = "/system/logs", .rewrite = "/system/logs" }; if (filesystem.resolve("/system/configuration") != null) {
return .{ .engine = &filesystem, .mounts = mount_specs[0..3], .flush = flushIfDirty }; std.log.info("volume {d} carries the system tree; backing /system/configuration and /system/logs", .{my_volume_id});
mount_specs[1] = .{ .prefix = "/system/configuration", .rewrite = "/system/configuration" };
mount_specs[2] = .{ .prefix = "/system/logs", .rewrite = "/system/logs" };
mount_count = 3;
} else {
std.log.info("volume {d} is a data volume; mounted at {s}", .{ my_volume_id, volume_mount_prefix });
}
return .{ .engine = &filesystem, .mounts = mount_specs[0..mount_count], .flush = flushIfDirty };
} }
pub fn main(init: process.Init) void { pub fn main(init: process.Init) void {
+81 -29
View File
@@ -172,39 +172,41 @@ fn setLabelFromUtf16(id: *Identity, name_bytes: []const u8) void {
id.label_len = @intCast(out); id.label_len = @intCast(out);
} }
/// The first GPT volume, or null if LBA 1 is not a valid GPT header or no entry /// Append every valid GPT volume to `out` (up to `out.len`), returning the count
/// validates. The header CRC-32 and the per-entry overflow-safe range check are /// (0 if LBA 1 is not a valid GPT header). The header CRC-32 and the per-entry
/// the confinement-safety guards the driver's clamp rests on — the invariant /// overflow-safe range check are the confinement-safety guards the driver's clamp
/// firstVolume documents for MBR, extended to untrusted GPT metadata. The /// rests on — the invariant documented for MBR, extended to untrusted GPT
/// entry-array CRC is deferred (correctness-only; the range check carries safety). /// metadata. The entry-array CRC is deferred (correctness-only; the range check
fn gptFirstVolume(reader: SectorReader, device_blocks: u64) ?Volume { /// carries safety).
fn gptAllVolumes(reader: SectorReader, device_blocks: u64, out: []Volume) usize {
var header: [sector_bytes]u8 = undefined; var header: [sector_bytes]u8 = undefined;
if (!reader.read(1, &header)) return null; if (!reader.read(1, &header)) return 0;
if (!std.mem.eql(u8, header[0..8], gpt_signature)) return null; if (!std.mem.eql(u8, header[0..8], gpt_signature)) return 0;
const header_size = std.mem.readInt(u32, header[12..16], .little); const header_size = std.mem.readInt(u32, header[12..16], .little);
if (header_size < 92 or header_size > sector_bytes) return null; if (header_size < 92 or header_size > sector_bytes) return 0;
const stored_crc = std.mem.readInt(u32, header[16..20], .little); const stored_crc = std.mem.readInt(u32, header[16..20], .little);
var check: [sector_bytes]u8 = undefined; var check: [sector_bytes]u8 = undefined;
@memcpy(check[0..header_size], header[0..header_size]); @memcpy(check[0..header_size], header[0..header_size]);
@memset(check[16..20], 0); @memset(check[16..20], 0);
if (crc32(check[0..header_size]) != stored_crc) return null; if (crc32(check[0..header_size]) != stored_crc) return 0;
const entry_lba = std.mem.readInt(u64, header[72..80], .little); const entry_lba = std.mem.readInt(u64, header[72..80], .little);
const num_entries = std.mem.readInt(u32, header[80..84], .little); const num_entries = std.mem.readInt(u32, header[80..84], .little);
const entry_size = std.mem.readInt(u32, header[84..88], .little); const entry_size = std.mem.readInt(u32, header[84..88], .little);
if (entry_size != 128 and entry_size != 256 and entry_size != 512) return null; if (entry_size != 128 and entry_size != 256 and entry_size != 512) return 0;
if (entry_lba == 0 or entry_lba >= device_blocks) return null; if (entry_lba == 0 or entry_lba >= device_blocks) return 0;
const scan = @min(num_entries, gpt_entry_scan_maximum); const scan = @min(num_entries, gpt_entry_scan_maximum);
var sector_buf: [sector_bytes]u8 = undefined; var sector_buf: [sector_bytes]u8 = undefined;
var loaded: u64 = std.math.maxInt(u64); var loaded: u64 = std.math.maxInt(u64);
var count: usize = 0;
var i: u32 = 0; var i: u32 = 0;
while (i < scan) : (i += 1) { while (i < scan and count < out.len) : (i += 1) {
const abs = @as(u64, i) * entry_size; const abs = @as(u64, i) * entry_size;
const lba = entry_lba + abs / sector_bytes; const lba = entry_lba + abs / sector_bytes;
const off = @as(usize, @intCast(abs % sector_bytes)); const off = @as(usize, @intCast(abs % sector_bytes));
if (lba != loaded) { if (lba != loaded) {
if (!reader.read(lba, &sector_buf)) return null; if (!reader.read(lba, &sector_buf)) break; // return what we have
loaded = lba; loaded = lba;
} }
const entry = sector_buf[off..][0..128]; // the fields we read live in the first 128 bytes const entry = sector_buf[off..][0..128]; // the fields we read live in the first 128 bytes
@@ -224,9 +226,10 @@ fn gptFirstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
if (start == 0 or end < start or end >= device_blocks) continue; if (start == 0 or end < start or end >= device_blocks) continue;
var id = Identity{ .rung = .gpt_guid, .key = std.mem.readInt(u128, entry[16..32], .little) }; var id = Identity{ .rung = .gpt_guid, .key = std.mem.readInt(u128, entry[16..32], .little) };
setLabelFromUtf16(&id, entry[56..128]); setLabelFromUtf16(&id, entry[56..128]);
return .{ .base_lba = start, .block_count = end - start + 1, .identity = id }; out[count] = .{ .base_lba = start, .block_count = end - start + 1, .identity = id };
count += 1;
} }
return null; return count;
} }
/// Trim trailing spaces (FAT labels are space-padded) and copy into the display /// Trim trailing spaces (FAT labels are space-padded) and copy into the display
@@ -259,18 +262,22 @@ fn fatIdentity(reader: SectorReader, start_lba: u64) ?Identity {
return id; return id;
} }
/// The first volume on the device `reader` addresses, whose whole-device size is /// Append every volume on the device `reader` addresses, whose whole-device size
/// `device_blocks`, or null if none is found. A GPT disk (protective MBR) is /// is `device_blocks`, to `out` (up to `out.len`), returning the count. A GPT
/// handled by GPT, authoritatively — its null is final. Otherwise an MBR with a /// disk (protective MBR) is enumerated by GPT, authoritatively — a zero count is
/// non-empty entry yields that partition's [start, size); otherwise a boot /// final. Otherwise every fitting MBR entry is a volume; a boot signature with no
/// signature with no partitions is treated as a bare FAT spanning the device. /// partition entries is a bare FAT spanning the whole device. Each volume's
pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume { /// [start, count) is validated overflow-safe (the confinement invariant the
/// driver's clamp rests on), and each prefers its FAT serial identity over the
/// disk signature.
pub fn allVolumes(reader: SectorReader, device_blocks: u64, out: []Volume) usize {
var block0: [sector_bytes]u8 = undefined; var block0: [sector_bytes]u8 = undefined;
if (!reader.read(0, &block0)) return null; if (!reader.read(0, &block0)) return 0;
if (!hasBootSignature(&block0)) return null; if (!hasBootSignature(&block0)) return 0;
if (isProtectiveMbr(&block0)) return gptFirstVolume(reader, device_blocks); if (isProtectiveMbr(&block0)) return gptAllVolumes(reader, device_blocks, out);
var count: usize = 0;
var index: u8 = 0; var index: u8 = 0;
while (index < 4) : (index += 1) { while (index < 4 and count < out.len) : (index += 1) {
const entry = block0[446 + @as(usize, index) * 16 ..][0..16]; const entry = block0[446 + @as(usize, index) * 16 ..][0..16];
const kind = entry[4]; const kind = entry[4];
const start = std.mem.readInt(u32, entry[8..12], .little); const start = std.mem.readInt(u32, entry[8..12], .little);
@@ -283,10 +290,25 @@ pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
// device (usb-storage.zig resolveTransfer), which only holds because the // device (usb-storage.zig resolveTransfer), which only holds because the
// range handed down is validated here. The subtraction cannot overflow. // range handed down is validated here. The subtraction cannot overflow.
if (start > device_blocks or device_blocks - start < size) continue; if (start > device_blocks or device_blocks - start < size) continue;
return .{ .base_lba = start, .block_count = size, .identity = fatIdentity(reader, start) orelse mbrIdentity(&block0, index) }; out[count] = .{ .base_lba = start, .block_count = size, .identity = fatIdentity(reader, start) orelse mbrIdentity(&block0, index) };
count += 1;
} }
// No partition entries: a bare FAT spanning the device. if (count == 0 and out.len > 0) {
return .{ .base_lba = 0, .block_count = device_blocks, .identity = fatIdentity(reader, 0) orelse mbrIdentity(&block0, 0) }; // No partition entries: a bare FAT spanning the device.
out[0] = .{ .base_lba = 0, .block_count = device_blocks, .identity = fatIdentity(reader, 0) orelse mbrIdentity(&block0, 0) };
return 1;
}
return count;
}
/// firstVolume is allVolumes into a one-element buffer.
const one_volume_slot = 1;
/// The first volume on the device, or null — the single-volume case of
/// `allVolumes`, kept for callers that want just one.
pub fn firstVolume(reader: SectorReader, device_blocks: u64) ?Volume {
var one: [one_volume_slot]Volume = undefined;
return if (allVolumes(reader, device_blocks, &one) > 0) one[0] else null;
} }
/// A read-only RAM disk over a byte slice of sectors, for the host tests. /// A read-only RAM disk over a byte slice of sectors, for the host tests.
@@ -507,3 +529,33 @@ test "GPT with 256-byte entries reads the non-128 offset arithmetic correctly" {
try std.testing.expectEqual(Rung.gpt_guid, v.identity.rung); try std.testing.expectEqual(Rung.gpt_guid, v.identity.rung);
try std.testing.expectEqual(@as(u128, 0xF00D), v.identity.key); try std.testing.expectEqual(@as(u128, 0xF00D), v.identity.key);
} }
// A fixture-sized volume buffer for the multi-volume tests, named so the bounds
// gate (which flags literal array lengths) stays quiet: a test input.
const test_volume_slots = 4;
test "allVolumes returns every fitting MBR partition with distinct identities" {
var block0 = [_]u8{0} ** 512;
block0[510] = 0x55;
block0[511] = 0xAA;
std.mem.writeInt(u32, block0[440..444], 0xDEADBEEF, .little);
// partition 0: start 2048, size 1000
block0[446 + 4] = 0x0c;
std.mem.writeInt(u32, block0[446 + 8 ..][0..4], 2048, .little);
std.mem.writeInt(u32, block0[446 + 12 ..][0..4], 1000, .little);
// partition 1: start 4096, size 2000
block0[462 + 4] = 0x0c;
std.mem.writeInt(u32, block0[462 + 8 ..][0..4], 4096, .little);
std.mem.writeInt(u32, block0[462 + 12 ..][0..4], 2000, .little);
const disk = RamDisk{ .sectors = &block0 };
var vols: [test_volume_slots]Volume = undefined;
const n = allVolumes(disk.reader(), 200000, &vols);
try std.testing.expectEqual(@as(usize, 2), n); // both partitions, not just the first
try std.testing.expectEqual(@as(u64, 2048), vols[0].base_lba);
try std.testing.expectEqual(@as(u64, 4096), vols[1].base_lba);
// distinct rung-4 identities (no FAT VBR at those LBAs): index 0 vs 1.
try std.testing.expectEqual((@as(u128, 0xDEADBEEF) << 8) | 0, vols[0].identity.key);
try std.testing.expectEqual((@as(u128, 0xDEADBEEF) << 8) | 1, vols[1].identity.key);
// firstVolume (the 1-buffer case) still returns just the first.
try std.testing.expectEqual(@as(u64, 2048), firstVolume(disk.reader(), 200000).?.base_lba);
}
+250 -164
View File
@@ -8,10 +8,12 @@
//! supervises the filesystems it spawns, exactly as the device manager //! supervises the filesystems it spawns, exactly as the device manager
//! supervises drivers. //! supervises drivers.
//! //!
//! This increment (V3b) is the flip: the FAT service stops acquiring its own //! The manager holds a table of adopted storage DEVICES and a table of the
//! volume and is spawned here instead, confined to its partition, and handed //! VOLUMES on them: it adopts every storage device the device-manager tree
//! its channel over the volume-manager protocol. Single volume for now; the //! carries, probes each one's whole partition table, and spawns one filesystem
//! mount map (volumes.csv) and multi-volume land next. //! process per volume — each confined to its partition's badge-scoped block
//! range, each supervised with its own budget. A device leaving the tree takes
//! its volumes with it.
const std = @import("std"); const std = @import("std");
const channel = @import("channel"); const channel = @import("channel");
@@ -35,22 +37,37 @@ const Serve = volume_manager_protocol.Protocol.Provider(void);
const Invocation = envelope.Invocation; const Invocation = envelope.Invocation;
const Answer = envelope.Answer; const Answer = envelope.Answer;
/// The single volume this increment handles: its provider channel, its block /// One adopted storage device: the block channel to its provider (opened once and
/// sub-range, its identity, the id it is addressed by, and the filesystem /// shared — refcounted per confined filesystem via the hello reply) and the
/// process serving it (0 until spawned; reset on death for respawn). /// device-manager id it serves. A device leaving the tree takes its volumes.
const Volume = struct { const StorageDevice = struct {
storage: block.Device, used: bool = false,
storage_device_id: u64, // the device-manager id this volume's provider serves device_id: u64 = 0,
base_lba: u64, channel: block.Device = undefined,
block_count: u64,
identity: partition.Identity,
id: u64,
binary: []const u8, // the service binary, from filesystems.csv by signature
mount_prefix: []const u8, // the volume-root mount path (its id-path, or a volumes.csv override)
filesystem_pid: u32 = 0,
}; };
const volume_id: u64 = 1; /// One volume: which device serves it, its block sub-range, its content
/// identity, the id it is addressed by, the service binary + mount path it was
/// spawned with, the filesystem process serving it, and its own supervision
/// budget (so one volume's crash loop never touches another's).
const Volume = struct {
used: bool = false,
device_id: u64 = 0,
base_lba: u64 = 0,
block_count: u64 = 0,
identity: partition.Identity = .{ .rung = .anonymous },
id: u64 = 0,
binary: []const u8 = "",
mount_prefix: []const u8 = "",
filesystem_pid: u32 = 0,
// Per-volume supervision, mirroring the device manager's: a clean exit is not
// restarted, a fault restarts with backoff, a fast crash loop gives up.
restarts: u32 = 0,
spawn_ns: u64 = 0,
failed: bool = false,
restart_pending: bool = false,
restart_due_ns: u64 = 0,
};
// The mount map, read from configuration at boot (the policy home, storage- // The mount map, read from configuration at boot (the policy home, storage-
// architecture.md): filesystems.csv (content signature -> service binary) and // architecture.md): filesystems.csv (content signature -> service binary) and
@@ -80,47 +97,77 @@ var filesystem_rules: [maximum_filesystem_rules]filesystem_map.Rule = undefined;
var filesystem_rule_count: usize = 0; var filesystem_rule_count: usize = 0;
var volume_rules: [maximum_volume_rules]volume_map.Override = undefined; var volume_rules: [maximum_volume_rules]volume_map.Override = undefined;
var volume_rule_count: usize = 0; var volume_rule_count: usize = 0;
/// The composed default mount path (/volumes/<id>) for the current volume; a
/// volumes.csv override is used in place and needs no buffer (it is already a
/// slice into volumes_source). One buffer suffices while the manager serves one
/// volume (multi-volume gives each its own in S3).
/// bound: bytes of a composed /volumes/<id> mount path /// bound: bytes of a composed /volumes/<id> mount path
/// decided-by: ours /// decided-by: ours
/// protects: the mount_prefix_buf below /// protects: the per-volume mount_prefix buffers below
/// at-limit: truncate - bufPrint fails; the volume mounts at a fallback path (logged) /// at-limit: truncate - bufPrint fails; the volume mounts at a fallback path (logged)
/// observed-by: the fallback path in the log /// observed-by: the fallback path in the log
const mount_path_maximum = 64; const mount_path_maximum = 64;
var mount_prefix_buf: [mount_path_maximum]u8 = undefined;
/// bound: volumes the manager serves at once
/// decided-by: ours
/// protects: the volumes table and its per-volume mount-path buffers
/// at-limit: truncate - a further partition is left unserved and logged (real
/// machines carry a handful of volumes, far under this)
/// observed-by: the "volume table full" log line
const maximum_volumes = 16;
/// bound: storage devices the manager adopts at once
/// decided-by: ours
/// protects: the devices table
/// at-limit: truncate - a further device is left unadopted and logged
/// observed-by: the "device table full" log line
const maximum_devices = 8;
var devices = [_]StorageDevice{.{}} ** maximum_devices;
var volumes = [_]Volume{.{}} ** maximum_volumes;
/// Each volume's composed default mount path lives in its slot's buffer; a
/// volumes.csv override is used in place (a slice into volumes_source, no buffer).
var mount_prefix_bufs: [maximum_volumes][mount_path_maximum]u8 = undefined;
var next_volume_id: u64 = 1; // monotonic — never reused, so a stale id can't address the wrong child
var service_endpoint: ipc.Handle = 0; var service_endpoint: ipc.Handle = 0;
var manager_handle: ?ipc.Handle = null; var manager_handle: ?ipc.Handle = null;
var bounce: memory.DmaRegion = undefined; var bounce: memory.DmaRegion = undefined;
var bounce_ready = false; var bounce_ready = false;
/// The currently-mounted volume, or null while no storage is present. The whole /// How often the poll checks device presence and fires due restarts. Fast enough
/// removal lifecycle is this field going null and back: the poll sees the /// that an unplug unmounts promptly; the poll is a bare device-manager enumerate,
/// storage provider leave the device tree (a pulled stick), kills the filesystem /// no channel work, so it is cheap to run continuously.
/// and clears this; when it returns, the poll re-acquires and re-mounts.
var volume: ?Volume = null;
var logged_no_volume = false;
/// How often the poll checks whether the storage provider is present. Fast
/// enough that an unplug unmounts promptly; the poll is a bare device-manager
/// enumerate, no channel work, so it is cheap to run continuously.
const poll_interval_ms = 500; const poll_interval_ms = 500;
// Filesystem supervision, mirroring the device manager's (device-manager.zig): // Filesystem supervision, mirroring the device manager's (device-manager.zig).
// a clean exit is not restarted, a fault restarts with backoff, and a fast
// crash loop gives up rather than spinning. Without this a faulting filesystem
// respawns in a zero-delay loop.
const fast_death_ns: u64 = 2_000_000_000; const fast_death_ns: u64 = 2_000_000_000;
const crash_loop_cap: u32 = 3; const crash_loop_cap: u32 = 3;
const backoff_base_ms: u64 = 300; const backoff_base_ms: u64 = 300;
var fs_restarts: u32 = 0; /// bytes to format a u64 volume id as decimal (20 digits fit)
var fs_spawn_ns: u64 = 0; const id_decimal_bytes = 24;
var fs_failed = false;
/// A fat restart is due at `restart_due_ns`; the poll loop performs it once the // --- table lookups -----------------------------------------------------------
/// backoff has elapsed (one timer, folded into the poll — no second timer).
var restart_pending = false; fn deviceById(id: u64) ?*StorageDevice {
var restart_due_ns: u64 = 0; for (&devices) |*d| if (d.used and d.device_id == id) return d;
return null;
}
fn claimDevice() ?*StorageDevice {
for (&devices) |*d| if (!d.used) return d;
return null;
}
fn volumeById(id: u64) ?*Volume {
for (&volumes) |*v| if (v.used and v.id == id) return v;
return null;
}
fn volumeByPid(pid: u32) ?*Volume {
for (&volumes) |*v| if (v.used and v.filesystem_pid == pid) return v;
return null;
}
fn firstUsedVolume() ?*Volume {
for (&volumes) |*v| if (v.used) return v;
return null;
}
fn claimVolumeIndex() ?usize {
for (&volumes, 0..) |*v, i| if (!v.used) return i;
return null;
}
// --- device-manager plumbing -------------------------------------------------
fn deviceManager() ?ipc.Handle { fn deviceManager() ?ipc.Handle {
if (manager_handle) |h| return h; if (manager_handle) |h| return h;
@@ -131,13 +178,12 @@ fn deviceManager() ?ipc.Handle {
const OpenedStorage = struct { device_id: u64, device: block.Device }; const OpenedStorage = struct { device_id: u64, device: block.Device };
/// The first mass-storage provider whose block channel actually opens, with its /// The first mass-storage provider whose block channel opens and is NOT already
/// device id. A device-manager tree can carry more than one entry of the /// adopted, with its device id. A device-manager tree can carry more than one
/// mass-storage identity — a phantom that no driver is bound to answers a /// entry of the mass-storage identity — a phantom that no driver is bound to
/// consumer hello with NO channel — so this tries each and takes the first that /// answers a consumer hello with NO channel — so this tries each and takes the
/// yields a channel, exactly as a filesystem's own acquisition loop does. /// first that yields a channel. Skips already-adopted devices so a re-poll does
/// Called only when there is no volume (an insertion), so the hellos it makes /// not re-open a device it already serves.
/// are not per-poll churn.
fn openAnyStorage() ?OpenedStorage { fn openAnyStorage() ?OpenedStorage {
const manager = deviceManager() orelse return null; const manager = deviceManager() orelse return null;
const Entry = device_manager_protocol.ChildEntry; const Entry = device_manager_protocol.ChildEntry;
@@ -157,6 +203,7 @@ fn openAnyStorage() ?OpenedStorage {
const entry = std.mem.bytesToValue(Entry, tail[index * @sizeOf(Entry) ..][0..@sizeOf(Entry)]); const entry = std.mem.bytesToValue(Entry, tail[index * @sizeOf(Entry) ..][0..@sizeOf(Entry)]);
if (entry.device_id == device_manager_protocol.no_device) continue; if (entry.device_id == device_manager_protocol.no_device) continue;
if ((entry.identity >> 16) & 0xff != 0x08 or (entry.identity >> 8) & 0xff != 0x06) continue; if ((entry.identity >> 16) & 0xff != 0x08 or (entry.identity >> 8) & 0xff != 0x06) continue;
if (deviceById(entry.device_id) != null) continue; // already adopted
const exchanged = driver.helloOn(manager, .consumer, entry.device_id, null, true) orelse continue; const exchanged = driver.helloOn(manager, .consumer, entry.device_id, null, true) orelse continue;
const provider = exchanged.channel orelse continue; // a phantom / not-yet-bound entry const provider = exchanged.channel orelse continue; // a phantom / not-yet-bound entry
return .{ .device_id = entry.device_id, .device = .{ .endpoint = provider } }; return .{ .device_id = entry.device_id, .device = .{ .endpoint = provider } };
@@ -167,7 +214,7 @@ fn openAnyStorage() ?OpenedStorage {
/// Whether `device_id` is still in the device-manager tree — a bare enumerate, /// Whether `device_id` is still in the device-manager tree — a bare enumerate,
/// no consumer-hello, so it is cheap to call every poll. This is how removal is /// no consumer-hello, so it is cheap to call every poll. This is how removal is
/// detected: the specific device the mounted volume sits on disappears. /// detected: the specific device a mounted volume sits on disappears.
fn isDevicePresent(device_id: u64) bool { fn isDevicePresent(device_id: u64) bool {
const manager = deviceManager() orelse return false; const manager = deviceManager() orelse return false;
const Entry = device_manager_protocol.ChildEntry; const Entry = device_manager_protocol.ChildEntry;
@@ -191,59 +238,88 @@ fn isDevicePresent(device_id: u64) bool {
} }
} }
/// Spawn the filesystem for `v`, confine it to the volume's range, and record // --- lifecycle ---------------------------------------------------------------
/// its pid. The confinement is defined for the fresh pid BEFORE the filesystem
/// runs, so its first read is already bounded; the volume manager is the /// Spawn the filesystem for `v`, confine it to the volume's range on its device's
/// confinement controller (it defines the first range on the device). /// channel, and record its pid. The confinement is defined for the fresh pid
/// BEFORE the filesystem runs, so its first read is already bounded; the volume
/// manager is the confinement controller (it defines the first range on the
/// device).
fn spawnFilesystem(v: *Volume) void { fn spawnFilesystem(v: *Volume) void {
if (fs_failed) return; if (v.failed) return;
const pid = process.spawnSupervised(v.binary, &.{ "1", v.mount_prefix }, service_endpoint) orelse { const dev = deviceById(v.device_id) orelse return; // its device left — poll will clean up
var id_str_buf: [id_decimal_bytes]u8 = undefined;
const id_str = std.fmt.bufPrint(&id_str_buf, "{d}", .{v.id}) catch "1";
const pid = process.spawnSupervised(v.binary, &.{ id_str, v.mount_prefix }, service_endpoint) orelse {
_ = logging.write("volume-manager: could not spawn the filesystem; retrying\n"); _ = logging.write("volume-manager: could not spawn the filesystem; retrying\n");
armRestart(); armRestart(v);
return; return;
}; };
if (!v.storage.defineRange(pid, v.base_lba, v.block_count)) { if (!dev.channel.defineRange(pid, v.base_lba, v.block_count)) {
_ = logging.write("volume-manager: could not confine the filesystem to its volume; retrying\n"); _ = logging.write("volume-manager: could not confine the filesystem to its volume; retrying\n");
_ = process.kill(pid); _ = process.kill(pid);
armRestart(); armRestart(v);
return; return;
} }
v.filesystem_pid = pid; v.filesystem_pid = pid;
fs_spawn_ns = time.clock(); v.spawn_ns = time.clock();
std.log.info("volume 0x{x} -> {s} (pid {d}), lba {d}, {d} blocks", .{ v.identity.key, v.binary, pid, v.base_lba, v.block_count }); std.log.info("volume 0x{x} -> {s} (pid {d}), lba {d}, {d} blocks", .{ v.identity.key, v.binary, pid, v.base_lba, v.block_count });
} }
/// Schedule a fat restart after backoff; the poll loop performs it once due. /// Schedule a restart for `v` after backoff; the poll loop performs it once due.
fn armRestart() void { fn armRestart(v: *Volume) void {
const delay = if (fs_restarts == 0) backoff_base_ms else backoff_base_ms << @intCast(@min(fs_restarts - 1, 5)); const delay = if (v.restarts == 0) backoff_base_ms else backoff_base_ms << @intCast(@min(v.restarts - 1, 5));
restart_due_ns = time.clock() + delay * 1_000_000; v.restart_due_ns = time.clock() + delay * 1_000_000;
restart_pending = true; v.restart_pending = true;
} }
/// A storage provider just appeared: open its channel, read block 0, parse the /// Compose a volume's mount path (its id-path `/volumes/<id>`, or a volumes.csv
/// volume, and spawn its filesystem. On any failure the channel is closed (so a /// override) into its slot's buffer, and return the slice.
/// present-but-unreadable device does not leak a handle every poll) and `volume` fn composeMountPrefix(slot: usize, identity: partition.Identity) []const u8 {
/// stays null — the next poll retries. A fresh medium gets a fresh supervision var id_buf: [volume_map.id_maximum]u8 = undefined;
/// budget. const id = volume_map.idString(identity, &id_buf);
fn bringUpVolume() void { return volume_map.overrideFor(volume_rules[0..volume_rule_count], id) orelse
(std.fmt.bufPrint(&mount_prefix_bufs[slot], "/volumes/{s}", .{id}) catch "/volumes/unknown");
}
/// Adopt the next present, not-yet-adopted storage device: take its channel,
/// probe its whole partition table, and spawn a filesystem per volume it carries.
/// Returns true when it consumed a device (so the caller can loop to adopt every
/// present device in one tick), false when none remain or the device table is full.
///
/// A device is adopted exactly once and kept until it leaves the tree — even when
/// it carries no volume we can serve, or its geometry cannot be read. Keeping the
/// empty/unreadable device adopted (rather than dropping and re-probing) is what
/// lets openAnyStorage advance PAST it to the devices behind it; dropping it would
/// make openAnyStorage hand back the same unservable device every tick and starve
/// the rest. A genuine removal frees the slot (removeDevice); a re-insert gets a
/// fresh device id and is probed anew.
fn bringUpVolume() bool {
if (!bounce_ready) { if (!bounce_ready) {
bounce = memory.dmaAlloc(512, memory.dma_coherent | memory.dma_shareable) orelse return; bounce = memory.dmaAlloc(512, memory.dma_coherent | memory.dma_shareable) orelse return false;
bounce_ready = true; bounce_ready = true;
} }
const opened = openAnyStorage() orelse return; const opened = openAnyStorage() orelse return false;
const dev = claimDevice() orelse {
_ = logging.write("volume-manager: device table full; a storage device is left unadopted\n");
_ = ipc.close(opened.device.endpoint);
return false;
};
dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device };
const device = opened.device; const device = opened.device;
// Attach the read buffer to THIS device (a no-op without an enforcing IOMMU). // Attach the read buffer to THIS device (a no-op without an enforcing IOMMU).
// The handle is kept, not closed, so it can be re-attached to the next // The handle is kept, not closed, so it can be re-attached after a replug. A
// device after a replug. // failed attach or geometry read leaves the device adopted but empty — we just
// cannot read it, and the slot still watches it for removal.
if (bounce.handle) |handle| { if (bounce.handle) |handle| {
if (!device.attach(handle)) { if (!device.attach(handle)) {
_ = ipc.close(device.endpoint); _ = logging.write("volume-manager: could not attach the read buffer to a storage device; no volume served\n");
return; return true;
} }
} }
const geometry = device.geometry() orelse { const geometry = device.geometry() orelse {
_ = ipc.close(device.endpoint); _ = logging.write("volume-manager: could not read a storage device's geometry; no volume served\n");
return; return true;
}; };
const ProbeReader = struct { const ProbeReader = struct {
device: block.Device, device: block.Device,
@@ -257,77 +333,88 @@ fn bringUpVolume() void {
}; };
var probe = ProbeReader{ .device = device }; var probe = ProbeReader{ .device = device };
const reader = partition.SectorReader{ .context = &probe, .readFn = ProbeReader.readSector }; const reader = partition.SectorReader{ .context = &probe, .readFn = ProbeReader.readSector };
const found = partition.firstVolume(reader, geometry.block_count) orelse { var found: [maximum_volumes]partition.Volume = undefined;
if (!logged_no_volume) { const n = partition.allVolumes(reader, geometry.block_count, found[0..]);
_ = logging.write("volume-manager: storage present but no recognizable volume\n"); if (n == 0) {
logged_no_volume = true; std.log.info("device {d} present but carries no recognizable volume", .{dev.device_id});
} return true;
_ = ipc.close(device.endpoint); }
return; for (found[0..n]) |fv| {
}; // Pick the service binary from the volume's content signature. A signature
// Pick the service binary from the volume's content signature. A signature // no filesystems.csv row serves goes unserved (logged), like an unbound
// no filesystems.csv row serves goes unserved (logged), like an unbound // device — the manager does not guess.
// device — the manager does not guess. const binary = filesystem_map.match(filesystem_rules[0..filesystem_rule_count], fv.signature) orelse {
const binary = filesystem_map.match(filesystem_rules[0..filesystem_rule_count], found.signature) orelse {
if (!logged_no_volume) {
_ = logging.write("volume-manager: no filesystem serves this volume's content; unserved\n"); _ = logging.write("volume-manager: no filesystem serves this volume's content; unserved\n");
logged_no_volume = true; continue;
} };
_ = ipc.close(device.endpoint); const slot = claimVolumeIndex() orelse {
return; _ = logging.write("volume-manager: volume table full; a volume is left unserved\n");
}; break;
// The mount path is the volume's identity id (/volumes/<id>), or a };
// volumes.csv override pinning it to a chosen path. The id is content-derived, volumes[slot] = .{
// so the path is stable and never a port or a label. .used = true,
var id_buf: [volume_map.id_maximum]u8 = undefined; .device_id = dev.device_id,
const id = volume_map.idString(found.identity, &id_buf); .base_lba = fv.base_lba,
const mount_prefix = volume_map.overrideFor(volume_rules[0..volume_rule_count], id) orelse .block_count = fv.block_count,
(std.fmt.bufPrint(&mount_prefix_buf, "/volumes/{s}", .{id}) catch "/volumes/unknown"); .identity = fv.identity,
logged_no_volume = false; .id = next_volume_id,
fs_restarts = 0; .binary = binary,
fs_failed = false; .mount_prefix = composeMountPrefix(slot, fv.identity),
restart_pending = false; };
volume = .{ .storage = device, .storage_device_id = opened.device_id, .base_lba = found.base_lba, .block_count = found.block_count, .identity = found.identity, .id = volume_id, .binary = binary, .mount_prefix = mount_prefix }; next_volume_id += 1;
spawnFilesystem(&volume.?); spawnFilesystem(&volumes[slot]);
}
return true;
} }
/// The storage provider left the device tree (a pulled stick): kill the /// Close a device's channel and free its slot. No volumes are touched (the caller
/// filesystem so its mounts are retired. Retirement is lazy, not an eager /// ensures none remain, or there never were any).
/// death-time sweep — killing the process marks the filesystem's backend fn dropDevice(dev: *StorageDevice) void {
/// endpoint dead, and the VFS router drops each mount that endpoint backed on _ = ipc.close(dev.channel.endpoint);
/// the next path resolution under it (that resolve frees the slot and returns dev.* = .{};
/// not_found). Then drop the now-dead channel and clear the volume; the next }
/// poll that sees storage return re-mounts.
fn removeVolume() void { /// Retire one volume: kill its filesystem so its mounts are retired. Retirement
const v = volume orelse return; /// is lazy, not an eager death-time sweep — killing the process marks the
/// filesystem's backend endpoint dead, and the VFS router drops each mount that
/// endpoint backed on the next path resolution under it (that resolve frees the
/// slot and returns not_found). Then free the volume slot.
fn removeVolumeState(v: *Volume) void {
std.log.info("storage for volume {d} removed; unmounting", .{v.id}); std.log.info("storage for volume {d} removed; unmounting", .{v.id});
if (v.filesystem_pid != 0) _ = process.kill(v.filesystem_pid); if (v.filesystem_pid != 0) _ = process.kill(v.filesystem_pid);
_ = ipc.close(v.storage.endpoint); v.* = .{};
volume = null;
restart_pending = false;
fs_restarts = 0;
fs_failed = false;
} }
/// One poll tick. Removal is checked FIRST and supersedes a pending restart: if /// A storage device left the tree (a pulled stick): retire every volume it served
/// the device is gone there is nothing to restart fat onto, and respawning it /// and drop its channel. One removal path, whether the device is pulled cleanly
/// against the dead channel would just churn until the crash cap. Only once the /// or vanishes.
/// device is confirmed present does a due restart fire. fn removeDevice(dev: *StorageDevice) void {
fn pollTick() void { for (&volumes) |*v| {
if (volume) |v| { if (v.used and v.device_id == dev.device_id) removeVolumeState(v);
// Serving: watch for the specific device leaving (a pulled stick).
if (!isDevicePresent(v.storage_device_id)) {
removeVolume();
return;
}
if (restart_pending and time.clock() >= restart_due_ns) {
restart_pending = false;
spawnFilesystem(&volume.?);
}
} else {
// Idle: try to bring a present storage device up.
bringUpVolume();
} }
dropDevice(dev);
}
/// One poll tick. Device removal is reconciled FIRST and supersedes a pending
/// restart: a volume whose device left is retired before its restart could fire,
/// so nothing respawns against a dead channel. Then due restarts fire for present
/// volumes; then, if no device is adopted, a present device is brought up.
fn pollTick() void {
for (&devices) |*dev| {
if (dev.used and !isDevicePresent(dev.device_id)) removeDevice(dev);
}
for (&volumes) |*v| {
if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) {
v.restart_pending = false;
spawnFilesystem(v);
}
}
// Adopt every present, not-yet-adopted storage device. Each call consumes at
// most one device (openAnyStorage skips the adopted), so the loop terminates
// once none remain; the maximum_devices guard is insurance against a logic
// slip, never the normal exit.
var adopted: usize = 0;
while (adopted < maximum_devices and bringUpVolume()) : (adopted += 1) {}
} }
/// A filesystem announces itself for the volume it was spawned to serve. Reply /// A filesystem announces itself for the volume it was spawned to serve. Reply
@@ -335,26 +422,26 @@ fn pollTick() void {
/// badge) as the call's returned capability. No channel means the volume is not /// badge) as the call's returned capability. No channel means the volume is not
/// ready — the filesystem retries. /// ready — the filesystem retries.
fn onHello(_: void, invocation: Invocation(volume_manager_protocol.Hello), _: Answer(void)) isize { fn onHello(_: void, invocation: Invocation(volume_manager_protocol.Hello), _: Answer(void)) isize {
const v = volume orelse return 0; // not probed yet — retryable, no cap const v = volumeById(invocation.target) orelse return 0; // not probed yet — retryable, no cap
if (invocation.target != v.id) return 0; // unknown volume — retryable
if (invocation.sender != v.filesystem_pid) { if (invocation.sender != v.filesystem_pid) {
// Not the filesystem we spawned for this volume. Refuse: only the // Not the filesystem we spawned for this volume. Refuse: only the confined
// confined filesystem gets the channel. // filesystem gets the channel.
std.log.info("refused hello for volume {d} from process {d}", .{ invocation.target, invocation.sender }); std.log.info("refused hello for volume {d} from process {d}", .{ invocation.target, invocation.sender });
return -envelope.EPERM; return -envelope.EPERM;
} }
service.replyWithCapability(v.storage.endpoint); const dev = deviceById(v.device_id) orelse return 0; // its device left — retryable
service.replyWithCapability(dev.channel.endpoint);
std.log.info("handed volume {d} to pid {d}", .{ v.id, invocation.sender }); std.log.info("handed volume {d} to pid {d}", .{ v.id, invocation.sender });
return 0; return 0;
} }
/// Answer a `volumes` query with the mounted volume's descriptor — its id (its /// Answer a `volumes` query with a mounted volume's descriptor — its id (its
/// mount path is /volumes/<id> unless overridden), its actual mount path, and /// mount path is /volumes/<id> unless overridden), its actual mount path, and its
/// its display label. This is how a shell or file manager reads a volume's /// display label. Software keys on the id; a UI shows the label. Returns the first
/// friendly name: software keys on the id, a UI shows the label. An empty reply /// mounted volume for now; a full enumerate is a later refinement. Empty reply
/// means no volume is mounted. /// means no volume is mounted.
fn onVolumes(_: void, _: Invocation(volume_manager_protocol.Volumes), answer: Answer(void)) isize { fn onVolumes(_: void, _: Invocation(volume_manager_protocol.Volumes), answer: Answer(void)) isize {
const v = volume orelse return 0; const v = firstUsedVolume() orelse return 0;
var id_buf: [volume_map.id_maximum]u8 = undefined; var id_buf: [volume_map.id_maximum]u8 = undefined;
const info = volume_manager_protocol.VolumeInfo{ const info = volume_manager_protocol.VolumeInfo{
.id = volume_map.idString(v.identity, &id_buf), .id = volume_map.idString(v.identity, &id_buf),
@@ -381,9 +468,9 @@ fn readConfig(path: []const u8, buf: []u8) usize {
defer file.close(); defer file.close();
var used: usize = 0; var used: usize = 0;
while (used < buf.len) { while (used < buf.len) {
const n = file.read(buf[used..]) orelse break; const nn = file.read(buf[used..]) orelse break;
if (n == 0) break; if (nn == 0) break;
used += n; used += nn;
} }
return used; return used;
} }
@@ -427,23 +514,22 @@ fn onNotification(badge: u64) void {
// reclaimed by the driver on the same death; the respawn confines afresh. // reclaimed by the driver on the same death; the respawn confines afresh.
if (got.isChildExit()) { if (got.isChildExit()) {
const dead = got.childProcessId(); const dead = got.childProcessId();
const v = &(volume orelse return); const v = volumeByPid(dead) orelse return;
if (v.filesystem_pid != dead) return;
v.filesystem_pid = 0; v.filesystem_pid = 0;
const reason = process.exitReason(dead) orelse .fault; const reason = process.exitReason(dead) orelse .fault;
if (reason == .exited) { if (reason == .exited) {
std.log.info("filesystem for volume {d} exited cleanly; not restarting", .{v.id}); std.log.info("filesystem for volume {d} exited cleanly; not restarting", .{v.id});
return; return;
} }
const alive = time.clock() -| fs_spawn_ns; const alive = time.clock() -| v.spawn_ns;
fs_restarts = if (alive < fast_death_ns) fs_restarts + 1 else 1; v.restarts = if (alive < fast_death_ns) v.restarts + 1 else 1;
if (fs_restarts >= crash_loop_cap) { if (v.restarts >= crash_loop_cap) {
fs_failed = true; v.failed = true;
std.log.info("filesystem for volume {d} is failing repeatedly; giving up", .{v.id}); std.log.info("filesystem for volume {d} is failing repeatedly; giving up", .{v.id});
return; return;
} }
std.log.info("filesystem for volume {d} died ({s}); restarting", .{ v.id, @tagName(reason) }); std.log.info("filesystem for volume {d} died ({s}); restarting", .{ v.id, @tagName(reason) });
armRestart(); armRestart(v);
} }
} }
+65
View File
@@ -816,6 +816,47 @@ CASES = [
"expect": r"volume-manager: volume 0x0*12345678 -> \S+ \(pid \d+\), lba \d+, \d+ blocks" "expect": r"volume-manager: volume 0x0*12345678 -> \S+ \(pid \d+\), lba \d+, \d+ blocks"
r"[\s\S]*volume-manager: handed volume \d+ to pid \d+", r"[\s\S]*volume-manager: handed volume \d+ to pid \d+",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# S3 multi-volume: a SECOND usb-storage device (a generated data volume, serial
# da7a0001, an empty FAT with no /system) plugged in beside the boot volume.
# Proves the volume manager adopts BOTH devices and spawns a confined fat per
# volume, each mounted at its own CONTENT id-path (/volumes/fat-<serial>); and
# that boot-volume detection is by content — only the volume that carries
# /system backs /system/configuration, while the data volume mounts at its
# id-path alone. Against the pre-S3 one-device/one-volume manager the data
# volume never mounts, so the da7a0001 lookaheads fail (toggle-demonstrated by
# checking out the step-2 volume-manager.zig).
{"name": "two-volumes",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"data_volume": {"serial": "DA7A0001", "label": "DATAVOL", "size_mib": 64},
"expect": r"(?s)(?=.*volume-manager: volume 0x0*12345678 -> )"
r"(?=.*volume-manager: volume 0x0*da7a0001 -> )"
r"(?=.*fat: mounted /volumes/fat-12345678)"
r"(?=.*fat: mounted /volumes/fat-da7a0001)"
r"(?=.*carries the system tree)"
r"(?=.*data volume; mounted at /volumes/fat-da7a0001)",
"fail": r"data volume; mounted at /volumes/fat-12345678|DANOS-TEST-RESULT: FAIL"},
# S3 shared-channel multi-volume: ONE usb-storage device carrying an MBR with
# TWO FAT partitions (da7a0001 at lba 2048, da7a0002 at lba 83968). allVolumes
# walks the table and the manager spawns a confined fat per partition on the
# SAME block channel, each clamped to its own LBA range (usb-storage's
# per-badge range table) — the path a pair of single-volume sticks (the
# two-volumes case) does NOT exercise. The two mount lines sit at two DISTINCT
# non-zero base_lbas on one device. Against the pre-uncap allVolumes (S3 step
# 2, capped to one partition) only da7a0001 mounts, so the da7a0002 lookaheads
# fail.
{"name": "partitioned-volume",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"data_volume": {"partitions": [{"serial": "DA7A0001", "size_mib": 40},
{"serial": "DA7A0002", "size_mib": 40}]},
"expect": r"(?s)(?=.*volume 0x0*da7a0001 -> \S+ \(pid \d+\), lba 2048, )"
r"(?=.*volume 0x0*da7a0002 -> \S+ \(pid \d+\), lba 83968, )"
r"(?=.*fat: mounted /volumes/fat-da7a0001)"
r"(?=.*fat: mounted /volumes/fat-da7a0002)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the # Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the
# fat-test client, after listing, makes a directory, writes+reads a file inside # fat-test client, after listing, makes a directory, writes+reads a file inside
# it, then removes the file, exercising the whole VFS -> fat mutation path. # it, then removes the file, exercising the whole VFS -> fat mutation path.
@@ -1422,6 +1463,30 @@ def run_case(arch, case):
cmd[cmd.index("-m") + 1] = case["mem"] cmd[cmd.index("-m") + 1] = case["mem"]
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
cmd += case["qemu_extra"] cmd += case["qemu_extra"]
# A multi-volume case attaches a second usb-storage device backed by a freshly
# GENERATED data volume: a distinct-serial FAT32 with no /system tree, so the
# volume manager mounts it at its own id-path and the fat process marks it a
# data volume (never a system volume). Regenerated per run — no image is
# committed to the tree (the user keeps the boot files copyable, not baked in).
if case.get("data_volume"):
dv = case["data_volume"]
data_img = os.path.join(WORK, "data-volume.img")
if dv.get("partitions"):
# One device, an MBR with several FAT partitions: several volumes share
# ONE block channel, each confined to its own LBA range.
gen = [sys.executable, os.path.join(REPO, "tools", "make-partitioned-image.py"), data_img]
for part in dv["partitions"]:
gen += [part["serial"], str(part.get("size_mib", 40))]
else:
# One device, one bare FAT volume.
gen = [sys.executable, os.path.join(REPO, "tools", "make-fat-image.py"),
"--serial", dv["serial"], "--label", dv.get("label", "DATAVOL"),
data_img, str(dv.get("size_mib", 64))]
subprocess.run(gen, check=True, stdout=subprocess.DEVNULL)
cmd += [
"-drive", f"if=none,id=datausb,format=raw,file={data_img}",
"-device", "usb-storage,bus=xhci.0,port=4,drive=datausb,removable=on,id=datastorage",
]
# A QMP control socket, always present (additive): how a case's `qmp_after` # A QMP control socket, always present (additive): how a case's `qmp_after`
# hook injects host-side events into the guest mid-run. Kept under a short temp # hook injects host-side events into the guest mid-run. Kept under a short temp
# dir, not WORK: a unix socket path is capped at ~104 bytes (sun_path), and a # dir, not WORK: a unix socket path is capped at ~104 bytes (sun_path), and a
-1
View File
@@ -194,7 +194,6 @@ system/kernel/process.zig:write_buffer
system/kernel/scheduler.zig:ipc_maximum_handles system/kernel/scheduler.zig:ipc_maximum_handles
system/kernel/scheduler.zig:maximum_space_mappings system/kernel/scheduler.zig:maximum_space_mappings
system/kernel/vfs.zig:maximum_directories system/kernel/vfs.zig:maximum_directories
system/kernel/vfs.zig:maximum_mounts
system/kernel/vfs.zig:maximum_prefix system/kernel/vfs.zig:maximum_prefix
system/kernel/vfs.zig:maximum_rewrite system/kernel/vfs.zig:maximum_rewrite
system/services/acpi/acpi.zig:blocks system/services/acpi/acpi.zig:blocks
+31 -7
View File
@@ -47,8 +47,14 @@ def fat32_geometry(total_sectors):
class Fat32Image: class Fat32Image:
def __init__(self, total_sectors): def __init__(self, total_sectors, volume_id=0x12345678, label="DANOS"):
self.total_sectors = total_sectors self.total_sectors = total_sectors
# The FAT volume serial (its content identity — the /volumes/fat-<id>
# mount path danos derives from it) and the display label. A second image
# needs a distinct serial so its id-path does not collide with the boot
# volume's.
self.volume_id = volume_id & 0xFFFFFFFF
self.label = label
self.fat_size, self.cluster_count = fat32_geometry(total_sectors) self.fat_size, self.cluster_count = fat32_geometry(total_sectors)
if self.cluster_count < 65525: if self.cluster_count < 65525:
sys.exit(f"error: image too small for FAT32 ({self.cluster_count} clusters " sys.exit(f"error: image too small for FAT32 ({self.cluster_count} clusters "
@@ -146,8 +152,8 @@ class Fat32Image:
0x80, # drive number 0x80, # drive number
0, # reserved 0, # reserved
0x29, # extended boot signature 0x29, # extended boot signature
0x12345678, # volume id self.volume_id, # volume id
b"DANOS ", # volume label self.label.encode("ascii", "replace")[:11].ljust(11, b" "), # volume label
b"FAT32 ", # filesystem type b"FAT32 ", # filesystem type
) )
sector[510] = 0x55 sector[510] = 0x55
@@ -276,9 +282,9 @@ def build_tree(pairs):
return root return root
def build(out_path, size_mib, pairs): def build(out_path, size_mib, pairs, volume_id=0x12345678, label="DANOS"):
total_sectors = size_mib * 1024 * 1024 // SECTOR total_sectors = size_mib * 1024 * 1024 // SECTOR
image = Fat32Image(total_sectors) image = Fat32Image(total_sectors, volume_id, label)
tree = build_tree(pairs) tree = build_tree(pairs)
write_directory(image, 2, tree, 0, True) write_directory(image, 2, tree, 0, True)
with open(out_path, "wb") as handle: with open(out_path, "wb") as handle:
@@ -350,14 +356,32 @@ def main(argv):
if len(argv) == 3 and argv[1] == "--verify": if len(argv) == 3 and argv[1] == "--verify":
verify(argv[2]) verify(argv[2])
return 0 return 0
# Optional flags ahead of the positionals: --serial <hex> sets the FAT volume
# id (the /volumes/fat-<id> content identity), --label <name> its display
# label. A second FAT image passes a distinct --serial so its id-path cannot
# collide with the boot volume's.
argv = list(argv)
volume_id = 0x12345678
label = "DANOS"
i = 1
while i < len(argv):
if argv[i] == "--serial" and i + 1 < len(argv):
volume_id = int(argv[i + 1], 16)
del argv[i:i + 2]
elif argv[i] == "--label" and i + 1 < len(argv):
label = argv[i + 1]
del argv[i:i + 2]
else:
i += 1
if len(argv) < 3 or (len(argv) - 3) % 2 != 0: if len(argv) < 3 or (len(argv) - 3) % 2 != 0:
sys.exit("usage: make-fat-image.py <out.img> <size-MiB> [<dest> <host>]...\n" sys.exit("usage: make-fat-image.py [--serial <hex>] [--label <name>] "
"<out.img> <size-MiB> [<dest> <host>]...\n"
" make-fat-image.py --verify <out.img>") " make-fat-image.py --verify <out.img>")
out_path = argv[1] out_path = argv[1]
size_mib = int(argv[2]) size_mib = int(argv[2])
rest = argv[3:] rest = argv[3:]
pairs = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)] pairs = [(rest[i], rest[i + 1]) for i in range(0, len(rest), 2)]
build(out_path, size_mib, pairs) build(out_path, size_mib, pairs, volume_id, label)
return 0 return 0
+88
View File
@@ -0,0 +1,88 @@
#!/usr/bin/env python3
"""Assemble an MBR-partitioned disk image from N FAT32 partitions — the danos
multi-volume test disk.
Pure Python 3 stdlib (no mtools / parted). Each partition is a real FAT32
filesystem produced by make-fat-image.py, laid out behind a classic MBR so the
danos partition prober (partition.allVolumes) walks the table and the volume
manager spawns one confined filesystem per partition — several volumes sharing
ONE block channel, each clamped to its own LBA range. That shared-channel,
per-partition path is what a single stick with two partitions exercises and a
pair of single-volume sticks does not.
make-partitioned-image.py <out.img> [<serial-hex> <size-MiB>]...
Each partition is an empty FAT32 with the given volume serial (its /volumes/
fat-<serial> content id). Partitions are 1-MiB aligned; the MBR marks each
type 0x0C (FAT32 LBA). At most four (an MBR holds four primaries).
"""
import os
import struct
import subprocess
import sys
import tempfile
SECTOR = 512
ALIGN = 2048 # sectors (1 MiB) — standard partition alignment, and the gap the MBR sits in
MBR_TYPE_FAT32_LBA = 0x0C
MAX_PRIMARY_PARTITIONS = 4
HERE = os.path.dirname(os.path.abspath(__file__))
def align_up(sectors, to=ALIGN):
return (sectors + to - 1) // to * to
def main(argv):
if len(argv) < 4 or (len(argv) - 2) % 2 != 0:
sys.exit("usage: make-partitioned-image.py <out.img> [<serial-hex> <size-MiB>]...")
out_path = argv[1]
specs = [(argv[i], int(argv[i + 1])) for i in range(2, len(argv), 2)]
if len(specs) > MAX_PRIMARY_PARTITIONS:
sys.exit(f"error: an MBR holds at most {MAX_PRIMARY_PARTITIONS} primary partitions")
# Generate each partition's FAT32 image, then place it at its aligned start.
partitions = [] # (start_sector, sector_count, bytes)
cursor = ALIGN # leave the first 1 MiB for the MBR + alignment gap
with tempfile.TemporaryDirectory() as tmp:
for idx, (serial, size_mib) in enumerate(specs):
part_path = os.path.join(tmp, f"p{idx}.img")
subprocess.run(
[sys.executable, os.path.join(HERE, "make-fat-image.py"),
"--serial", serial, "--label", f"DATA{idx}",
part_path, str(size_mib)],
check=True, stdout=subprocess.DEVNULL)
with open(part_path, "rb") as handle:
data = handle.read()
count = len(data) // SECTOR
partitions.append((cursor, count, data))
cursor = align_up(cursor + count)
total_sectors = cursor
disk = bytearray(total_sectors * SECTOR)
# The MBR: a disk signature, one partition entry per FAT partition, 0x55AA.
# No boot code (this disk is data, never booted); danos's mount() sees the
# signature but no BPB at LBA 0 and takes the MBR-walk path.
struct.pack_into("<I", disk, 440, 0x0D05DA05) # arbitrary but fixed disk signature
for idx, (start, count, data) in enumerate(partitions):
entry = 446 + idx * 16
disk[entry + 0] = 0x00 # not bootable
disk[entry + 1:entry + 4] = b"\xFE\xFF\xFF" # CHS start (LBA-aware tools ignore)
disk[entry + 4] = MBR_TYPE_FAT32_LBA
disk[entry + 5:entry + 8] = b"\xFE\xFF\xFF" # CHS end
struct.pack_into("<I", disk, entry + 8, start) # start LBA
struct.pack_into("<I", disk, entry + 12, count) # sector count
disk[start * SECTOR:start * SECTOR + len(data)] = data
disk[510] = 0x55
disk[511] = 0xAA
with open(out_path, "wb") as handle:
handle.write(disk)
print(f"make-partitioned-image: wrote {out_path} "
f"({total_sectors * SECTOR // (1024 * 1024)} MiB, {len(partitions)} partitions)")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))