5 Commits
Author SHA1 Message Date
Daniel Samson 4dfb5012c0 docs: the removal lifecycle closes — three triggers, one path (S5)
Storage removal is now robust to all three ways a volume can leave, and the docs
say so. storage-architecture.md and storage-design-rationale.md move medium_changed
from "planned to be consumed" to consumed, and record the third trigger:

  - the DEVICE leaving the tree (a pulled stick)          — presence polling
  - the MEDIUM leaving while its device stays (a reader)  — the volume manager
    now consumes the pushed medium_changed event
  - the storage DRIVER crashing while its device stays    — a channel-liveness
    geometry() probe reaps the volume and rebuilds it on the restarted driver's
    fresh channel; presence polling alone cannot see this (the V4 open edge)

The re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a
physical unplug/replug is bench-verified (QEMU cannot re-present a usb-storage
device_add). The transport-native eject signal (SCSI UNIT ATTENTION, AHCI
PxSSTS, NVMe namespace-change AER) in place of the TEST UNIT READY poll stays
the documented future refinement.
2026-08-10 06:08:56 +01:00
Daniel Samson 061eb7c004 volume-manager: drop the medium-event dedup — it only ever misfired (S5 review)
The adversarial S5 review found a real, unrecoverable defect: onMediumEvent
deduped medium_changed events on a module-global last_medium_change compared by
equality against the event's change_count. But change_count is a PER-DRIVER
counter that restarts at 0 in every usb-storage instance — it is bumped +%=1 and
published only on a real medium transition, so within one instance every count
is unique and monotonic and an equality dedup can never legitimately fire.

The global was carried across a driver restart — S5's OWN crash-rebuild path —
so a fresh instance's first eject (count=1) collided with a stale last==1 and was
dropped. removeDevice never ran; the filesystem kept serving I/O against absent
media forever, and nothing else recovered it: onGeometry answers from a cached
block_count so channelAlive stays true, and isDevicePresent stays true (the
device never left the tree). medium_changed is the sole eject oracle there.

The dedup guarded a re-delivery the driver already makes impossible, and its
only observable effect was this bug. Remove it: react to each present-edge
directly. Both branches are idempotent (a freed device stops matching dev.used)
and the poll reconciles, so acting on every genuine edge is safe — and a fresh
driver instance's counter can no longer be mistaken for the previous one's.

Verified: build + bounds green; volume-removal, volume-medium-change and
volume-driver-restart all pass, so both removal paths survive the change. The
driver-restart-then-eject intersection that triggered the collision cannot be
staged in QEMU — the internal driver-kill cannot be ordered against a QMP eject,
and a usb-storage device_add is not re-presented — so the guarantee rests on the
driver's one-publish-per-transition-with-unique-count contract.
2026-08-10 05:56:06 +01:00
Daniel Samson 5dc966838a volume-manager: rebuild a volume when its storage driver dies (S5)
The V4 review's open edge: a storage driver that crashes while its device
stays in the tree left fat wedged on a dead channel — device-presence
polling (a device-manager enumerate) still reported the device present,
so nothing reaped it. pollTick now also probes channelAlive(dev), a
geometry() on the block channel that fails fast on the dead endpoint; a
present device with a dead channel is reaped like a pull, and the adopt
loop re-adopts it on the restarted driver's fresh channel — the rebuild.
The manager's own liveness probe makes fat self-detection unnecessary:
it rebuilds regardless of the wedged filesystem's state.

The drill: the device manager gains a test-storage-restart mode that
kills usb-storage once, ~2s after its hello (post-mount); a new
volume-driver-restart kernel case boots a manual tree with it, and the
QEMU case asserts a SECOND mount of the same id-path after the reap —
the rebuild. A pre-S5 manager, checking only device presence, never
reaps, so the second mount never appears. The manually-spawned volume
manager needed kernel-supervisor protocol grants (bind its name, open
the device manager), as the other manual-tree services already have.

Full suite 133/133.
2026-08-10 05:21:48 +01:00
Daniel Samson 700452dc4e volume-manager: consume medium_changed — the second removal trigger (S5)
The block client gains subscribeMedium / unsubscribeMedium /
decodeMediumChanged (the reserved subscribe/unsubscribe verbs plus the
MediumChanged decode), so a consumer never hand-rolls the wire format.
The volume manager subscribes to each device it adopts and consumes the
event through the new on_buffered_message seam — never the protocol
dispatch, whose op numbers collide with the manager's own hello.

A medium leaving while its device stays in the tree (a card reader, an
eject) now runs the SAME kill-retire-remount path as a pulled stick:
absent retires the volume, present re-probes it. That closes the "two
triggers, one lifecycle" the architecture specifies — device-presence
polling alone could never see a medium leave under a present device.
dropDevice unsubscribes before closing so the driver's bounded
subscriber table frees the slot; on a dead channel (a real pull) the
call fails fast, proven by volume-removal still passing.

New volume-medium-change case: eject the medium (not the device) ->
"medium absent" -> the manager unmounts. Fails against a pre-S5 manager
that never subscribed. Suite 132/132 (device-authority is the known
child-cleanup flake, green on rerun).
2026-08-10 04:58:54 +01:00
Daniel Samson 0faa0fd21b service: on_buffered_message for pushed events (S5)
An additive, behavior-neutral callback. A buffered async message
(Received.isMessage — a pushed event from a provider this service
subscribed to) carries a payload in the receive buffer; run() now hands
it to on_buffered_message before falling through to on_notification with
the badge, so a coalesced timer/exit riding the same wake is not lost. A
service that does not set the callback (all of them today) is unchanged:
the isMessage branch is a no-op and on_notification still runs, exactly
as before.

This is the seam the volume manager needs to consume block
medium_changed: the event's reserved op number collides with the volume
manager's own hello, so it must be decoded by hand here, never through
the protocol dispatch. Full suite neutral (the lone device-authority
miss is a known child-cleanup race that passes on rerun).
2026-08-10 04:43:27 +01:00
9 changed files with 252 additions and 37 deletions
@@ -22,14 +22,17 @@
> (`system/services/exfat`): full read + write, directories, rename, and on-disk > (`system/services/exfat`): full read + write, directories, rename, and on-disk
> up-case folding, reusing `library/kernel/file-system-harness` wholesale — the > up-case folding, reusing `library/kernel/file-system-harness` wholesale — the
> reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an > reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an
> `exfat-<serial>` id-path. **Still pending**: the `filesystem UUID` rung (ext- > `exfat-<serial>` id-path. Removal is robust to all three triggers now: a
> family superblocks, which need such an engine); the volume manager *consuming* > pulled device (presence polling), a medium that leaves while its device stays
> `medium_changed` > (the volume manager CONSUMES `medium_changed`), and a storage driver that
> (removal is detected by device-presence polling; the event is published but only > crashes while its device stays present (a channel-liveness `geometry()` probe
> a card-reader medium change needs the subscription); the remount-on-replug > reaps the volume and rebuilds it on the restarted driver's fresh channel). The
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller > re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a physical
> device, so it is bench-verified); and arbitration when two volumes both resolve > unplug/replug exercises the same path and is bench-verified (QEMU cannot
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few > re-present a usb-storage `device_add`). **Still pending**: the `filesystem UUID`
> rung (ext-family superblocks, which need such an engine); and arbitration when
> two volumes both resolve the boot markers (S3 mounts both and logs each claim;
> picking one is deferred). A few
> markers below are left where a duty is still pending. > markers below are left where a duty is still pending.
## The model ## The model
@@ -187,25 +190,26 @@ surprise-removal path — kill the filesystem process, retire its mounts,
respawn on return. No half-alive states, no `remount-ro`, no mounts that respawn on return. No half-alive states, no `remount-ro`, no mounts that
error forever (Plan 9's dead-server wart). error forever (Plan 9's dead-server wart).
The path has **two triggers, one lifecycle**: the *device* leaving (the The path folds **three triggers into one lifecycle**: the *device* leaving (a
storage driver dies — channel death, the table below), and the *medium* pulled stick — presence polling); the *medium* leaving while the device stays
leaving while the device stays (an SD card pulled from its reader, an ATAPI (an SD card pulled from its reader, an ATAPI tray opened, a USB card reader);
tray opened — including USB card readers today). The second trigger is the and a storage *driver crashing* while its device stays in the tree. The second
pushed `medium_changed` event on the block protocol — published today from a trigger is the pushed `medium_changed` event on the block protocol, published
TEST UNIT READY poll; still *planned* is the volume manager *consuming* it from a TEST UNIT READY poll — the volume manager now **consumes** it (subscribed
(today removal is driven only by device-presence polling) and translating the per device), running the same kill-retire path and re-probing on medium return,
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe so a swapped card is never served with the previous card's filesystem state. The
namespace-change AER) in place of the poll. On the event the volume manager third is caught by a channel-liveness `geometry()` probe: presence polling alone
runs the same kill-retire path, then re-probes on medium return exactly as on sees the device still present, but the channel is dead, so the manager reaps the
device return. Without it, a swapped card would be served with the previous volume and rebuilds it on the restarted driver's fresh channel. Still *planned*
card's filesystem state. is translating the transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS,
NVMe namespace-change AER) in place of the presence poll.
| Layer | Observes | Must do | Guarantees | | Layer | Observes | Must do | Guarantees |
|---|---|---|---| |---|---|---|---|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick | | Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds | | Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang | | Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* | | Volume manager *(built)* | a device leaving the tree (poll), a `medium_changed` event, or a dead channel under a still-present device (a crashed driver — `geometry()` liveness probe) | kill that volume's filesystem service (its mounts retire), then re-adopt + remount on return or on the restarted driver's fresh channel | one removal path for all three triggers; mounts never dangle; the manager never serves from behind a dead channel |
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel | | Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang | | Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media | | Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
@@ -209,18 +209,23 @@ matrix-proven shape; genuinely open.
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
precedent); format-level crash honesty (a Power-Safe-style journaling or COW precedent); format-level crash honesty (a Power-Safe-style journaling or COW
filesystem) once danos outgrows FAT; per-process namespaces. filesystem) once danos outgrows FAT; per-process namespaces.
7. **The media-presence event** (settled in principle; lands with the volume 7. **The media-presence event** (the consuming half is BUILT; the
manager): the block protocol gains a pushed event — `medium_changed`, with transport-native signal stays future): the block protocol carries a pushed
present/absent and a change counter — produced by the storage driver from event — `medium_changed`, with present/absent and a change counter —
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for produced today by the storage driver from a TEST UNIT READY poll (the
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and transport's native signal — SCSI UNIT ATTENTION, PxSSTS for AHCI,
consumed by the volume manager, which runs the SAME kill-retire-remount namespace-change AER for NVMe — is the future refinement in place of the
path it runs on channel death — one lifecycle, two triggers. The driver poll) and now **consumed** by the volume manager, which subscribes per
reports presence, never content; a pushed event carries no capability, device and runs the SAME kill-retire-remount path it runs on channel death.
which the kernel already guarantees. The device staying while its medium The driver reports presence, never content; a pushed event carries no
leaves is the one removable-media case the channel-death trigger cannot capability, which the kernel already guarantees. The device staying while
see; without this event a swapped SD card would be served with the old its medium leaves is the one removable-media case the channel-death trigger
card's filesystem state. cannot see; without this event a swapped SD card would be served with the
old card's filesystem state. A THIRD trigger closes the last gap — a
storage driver that *crashes* while its device stays present: channel death
there is invisible to presence polling, so the volume manager probes channel
liveness (`geometry()`) each tick and reaps-then-rebuilds the volume on the
restarted driver's fresh channel. One lifecycle, three triggers.
8. **Volume identity, and the mount map as danos's fstab** (settled). The 8. **Volume identity, and the mount map as danos's fstab** (settled). The
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
broke whenever a drive changed ports or enumeration order; `UUID=` entries broke whenever a drive changed ports or enumeration order; `UUID=` entries
+37
View File
@@ -7,12 +7,25 @@
//! `runtime.dma.alloc`), so whole sectors move without crossing the IPC size //! `runtime.dma.alloc`), so whole sectors move without crossing the IPC size
//! limit — the same handoff usb-storage uses toward the controller. //! limit — the same handoff usb-storage uses toward the controller.
const std = @import("std");
const envelope = @import("envelope"); const envelope = @import("envelope");
const ipc = @import("ipc"); const ipc = @import("ipc");
const block_protocol = @import("block-protocol"); const block_protocol = @import("block-protocol");
const Protocol = block_protocol.Protocol; const Protocol = block_protocol.Protocol;
/// The medium_changed event payload, re-exported so a consumer decodes it without
/// reaching into the wire-format module.
pub const MediumChanged = block_protocol.MediumChanged;
/// Decode a medium_changed event from a buffered-message payload a subscriber
/// received (a `Received.isMessage` wake). Null if the bytes are too short to be
/// one — a caller ignores anything that is not a well-formed event.
pub fn decodeMediumChanged(payload: []const u8) ?MediumChanged {
if (payload.len < envelope.prefix_size + @sizeOf(MediumChanged)) return null;
return std.mem.bytesToValue(MediumChanged, payload[envelope.prefix_size..][0..@sizeOf(MediumChanged)]);
}
pub const Geometry = struct { block_size: u32, block_count: u64 }; pub const Geometry = struct { block_size: u32, block_count: u64 };
pub const Device = struct { pub const Device = struct {
@@ -88,6 +101,30 @@ pub const Device = struct {
if (status.status != 0) return null; if (status.status != 0) return null;
return reply[0..answer.len]; return reply[0..answer.len];
} }
/// Subscribe `subscriber` (an endpoint) to this device's medium_changed
/// events: the reserved `subscribe` verb carries the subscriber's endpoint as
/// the capability, and the driver then ipc.sends each medium transition to it.
pub fn subscribeMedium(self: Device, subscriber: ipc.Handle) bool {
var packet: [block_protocol.message_maximum]u8 = undefined;
const framed = envelope.encodeSubscribe(0, &packet) orelse return false; // interest 0: every event (block has one)
var reply: [block_protocol.message_maximum]u8 = undefined;
const answer = ipc.callCap(self.endpoint, framed, &reply, subscriber) catch return false;
const status = envelope.statusOf(reply[0..answer.len]) orelse return false;
return status.status == 0;
}
/// Unsubscribe from this device's medium_changed events. Call before closing
/// the channel so the driver's bounded subscriber table frees the slot rather
/// than holding a dead endpoint until an exit sweep notices.
pub fn unsubscribeMedium(self: Device) bool {
var packet: [block_protocol.message_maximum]u8 = undefined;
const framed = envelope.encodeUnsubscribe(&packet) orelse return false;
var reply: [block_protocol.message_maximum]u8 = undefined;
const answer = ipc.callCap(self.endpoint, framed, &reply, null) catch return false;
const status = envelope.statusOf(reply[0..answer.len]) orelse return false;
return status.status == 0;
}
}; };
// There is deliberately no open-by-name here: `block` is not a registry name. // There is deliberately no open-by-name here: `block` is not a registry name.
+16
View File
@@ -67,6 +67,15 @@ pub const Callbacks = struct {
/// A notification that is not a signal — a subscribed exit event, a bound /// A notification that is not a signal — a subscribed exit event, a bound
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers. /// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
on_notification: ?*const fn (badge: u64) void = null, on_notification: ?*const fn (badge: u64) void = null,
/// A buffered async message (`Received.isMessage`): a pushed event from a
/// provider this service subscribed to, its payload in the receive buffer.
/// Unlike `on_message`, it never goes through the protocol dispatch — so an
/// event whose reserved op number collides with one of this service's own
/// verbs (a `block` `medium_changed` reaching the volume manager, whose own
/// protocol numbers `hello` the same) is decoded by hand here, not
/// mis-dispatched. Default null: the badge alone still reaches
/// `on_notification`, exactly as before this callback existed.
on_buffered_message: ?*const fn (message: []const u8) void = null,
/// The reload signal. Default: ignored. /// The reload signal. Default: ignored.
on_reload: ?*const fn () void = null, on_reload: ?*const fn () void = null,
/// The terminate signal, called before the loop returns. The clean exit is /// The terminate signal, called before the loop returns. The clean exit is
@@ -372,6 +381,13 @@ pub fn run(comptime maximum_message: usize, callbacks: Callbacks) void {
if (got.isChildExit()) { if (got.isChildExit()) {
if (callbacks.subscribers) |subscribers| subscribers.forget(got.childProcessId()); if (callbacks.subscribers) |subscribers| subscribers.forget(got.childProcessId());
} }
// A buffered async message (a pushed event) carries a payload; hand it
// to the service that asked for it. The badge still reaches
// on_notification below, so a coalesced timer/exit riding the same wake
// is not lost — and a service without this callback is unchanged.
if (got.isMessage()) {
if (callbacks.on_buffered_message) |onBuffered| onBuffered(receive[0..got.len]);
}
if (callbacks.on_notification) |onNotification| onNotification(got.badge); if (callbacks.on_notification) |onNotification| onNotification(got.badge);
continue; continue;
} }
+4
View File
@@ -79,6 +79,7 @@
# 'kernel' as the supervisor. Nothing else changes: the binary must still match. # 'kernel' as the supervisor. Nothing else changes: the binary must still match.
/system/services/input, kernel, bind, input /system/services/input, kernel, bind, input
/system/services/device-manager, kernel, bind, device-manager /system/services/device-manager, kernel, bind, device-manager
/system/services/volume-manager, kernel, bind, volume-manager
/system/services/fat, kernel, bind, vfs /system/services/fat, kernel, bind, vfs
/system/services/display, kernel, bind, display /system/services/display, kernel, bind, display
/system/services/discovery, kernel, bind, power /system/services/discovery, kernel, bind, power
@@ -107,6 +108,9 @@
# The volume manager reaches the device manager to be routed to each storage # The volume manager reaches the device manager to be routed to each storage
# provider's block channel, then confines a filesystem to each volume. # provider's block channel, then confines a filesystem to each volume.
/system/services/volume-manager, /system/services/init, open, device-manager /system/services/volume-manager, /system/services/init, open, device-manager
# ...and again under the kernel supervisor for the manual-tree drills (S5's
# volume-driver-restart spawns the volume manager directly, not via init).
/system/services/volume-manager, kernel, open, device-manager
/system/services/display, /system/services/init, open, scanout /system/services/display, /system/services/init, open, scanout
/system/services/display, /system/services/init, open, display /system/services/display, /system/services/init, open, display
/system/services/display, /system/services/init, open, input /system/services/display, /system/services/init, open, input
Can't render this file because it contains an unexpected character in line 12 and column 15.
+40
View File
@@ -229,6 +229,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
fatMountTest(boot_information); fatMountTest(boot_information);
} else if (eql(case, "exfat-volume")) { } else if (eql(case, "exfat-volume")) {
exfatVolumeTest(boot_information); exfatVolumeTest(boot_information);
} else if (eql(case, "volume-driver-restart")) {
volumeDriverRestartTest(boot_information);
} else if (eql(case, "device-list")) { } else if (eql(case, "device-list")) {
deviceListTest(boot_information); deviceListTest(boot_information);
} else if (eql(case, "pci-scan")) { } else if (eql(case, "pci-scan")) {
@@ -3685,6 +3687,44 @@ fn displayReattachTest(boot_information: *const BootInformation) void {
while (true) scheduler.yield(); while (true) scheduler.yield();
} }
/// Storage-driver-crash rebuild (S5): the device manager runs in
/// "test-storage-restart" mode and kills the usb-storage driver once, a moment
/// after its volume has mounted. The driver's device stays in the tree, so the
/// volume manager's presence poll alone would miss the death and leave fat wedged
/// on a dead channel; its channel-liveness probe must notice, reap the volume, and
/// rebuild on the restarted driver's fresh channel — a SECOND mount of the same
/// id-path is the proof. (A pre-S5 manager, checking only device presence, never
/// reaps, so the second mount never appears.)
fn volumeDriverRestartTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: volume-driver-restart\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
_ = spawnRegistry(rd);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ item.name, "test-storage-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned (test-storage-restart mode)", manager != 0);
check("volume-manager spawned", spawnNamed(rd, "volume-manager"));
check("fat-test client spawned", spawnNamed(rd, "fat-test"));
scheduler.setPriority(1); // below the tree, so it runs
while (true) scheduler.yield();
}
/// Process arguments, end to end: spawn args-echo bare (its argv[0] is the /// Process arguments, end to end: spawn args-echo bare (its argv[0] is the
/// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through /// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through
/// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument /// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument
@@ -170,6 +170,8 @@ var test_usb_killed = false;
var test_pci_restart_mode = false; var test_pci_restart_mode = false;
var test_scanout_restart_mode = false; var test_scanout_restart_mode = false;
var test_scanout_killed = false; var test_scanout_killed = false;
var test_storage_restart_mode = false;
var test_storage_killed = false;
var test_kill_pid: u32 = 0; var test_kill_pid: u32 = 0;
var test_kill_due_ns: u64 = 0; var test_kill_due_ns: u64 = 0;
@@ -600,6 +602,16 @@ fn onHello(_: void, invocation: Invocation(device_manager_protocol.Hello), _: An
test_kill_due_ns = time.clock() + 1_500_000_000; test_kill_due_ns = time.clock() + 1_500_000_000;
_ = time.timerOnce(manager_endpoint, 1600); _ = time.timerOnce(manager_endpoint, 1600);
} }
// Storage-driver-crash drill (S5): once, a moment after usb-storage hellos —
// long enough that its volume has mounted — kill it. The manager re-delegates
// the still-present device to a restarted driver on a fresh channel; the volume
// manager's channel-liveness probe must notice the dead channel and rebuild.
if (test_storage_restart_mode and !test_storage_killed and std.mem.eql(u8, driver.name(), "/system/drivers/usb-storage")) {
test_storage_killed = true;
test_kill_pid = invocation.sender;
test_kill_due_ns = time.clock() + 2_000_000_000; // after the ~0.6s mount
_ = time.timerOnce(manager_endpoint, 2100);
}
return 0; return 0;
} }
@@ -763,6 +775,7 @@ pub fn main(init: process.Init) void {
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart"); test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart"); test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart");
test_scanout_restart_mode = std.mem.eql(u8, mode, "test-scanout-restart"); test_scanout_restart_mode = std.mem.eql(u8, mode, "test-scanout-restart");
test_storage_restart_mode = std.mem.eql(u8, mode, "test-storage-restart");
} }
service.run(device_manager_protocol.message_maximum, .{ service.run(device_manager_protocol.message_maximum, .{
.service = "device-manager", .service = "device-manager",
@@ -306,6 +306,12 @@ fn bringUpVolume() bool {
return false; return false;
}; };
dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device }; dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device };
// Consume this device's medium_changed events (the second of the removal
// lifecycle's two triggers: the device stays in the tree while its medium
// leaves — a card reader, an eject). Best effort: a provider that never
// publishes the event simply never wakes us, and device-pull is still caught
// by the presence poll.
_ = opened.device.subscribeMedium(service_endpoint);
const device = opened.device; const device = opened.device;
// Attach the read buffer to THIS device (a no-op without an enforcing IOMMU). // Attach the read buffer to THIS device (a no-op without an enforcing IOMMU).
// The handle is kept, not closed, so it can be re-attached after a replug. A // The handle is kept, not closed, so it can be re-attached after a replug. A
@@ -370,6 +376,10 @@ fn bringUpVolume() bool {
/// Close a device's channel and free its slot. No volumes are touched (the caller /// Close a device's channel and free its slot. No volumes are touched (the caller
/// ensures none remain, or there never were any). /// ensures none remain, or there never were any).
fn dropDevice(dev: *StorageDevice) void { fn dropDevice(dev: *StorageDevice) void {
// Free the driver's subscriber slot before the channel closes. On a still-live
// channel (a medium eject) this frees the slot; on a dead one (a device pull)
// the call fails fast and the exit sweep frees it anyway.
_ = dev.channel.unsubscribeMedium();
_ = ipc.close(dev.channel.endpoint); _ = ipc.close(dev.channel.endpoint);
dev.* = .{}; dev.* = .{};
} }
@@ -395,13 +405,25 @@ fn removeDevice(dev: *StorageDevice) void {
dropDevice(dev); dropDevice(dev);
} }
/// Whether a device's block channel still answers — a geometry() probe. A storage
/// driver that DIED while its device stays in the tree (it crashed; the device
/// manager will re-delegate the device to a restarted driver on a FRESH channel)
/// leaves a dead channel here, even though isDevicePresent still reports the device
/// present. geometry() on the dead endpoint fails fast, so this catches the crash
/// that presence-polling alone cannot — the V4 review's open edge.
fn channelAlive(dev: *StorageDevice) bool {
return dev.channel.geometry() != null;
}
/// One poll tick. Device removal is reconciled FIRST and supersedes a pending /// One poll tick. Device removal is reconciled FIRST and supersedes a pending
/// restart: a volume whose device left is retired before its restart could fire, /// restart: a volume whose device left (a pull) OR whose driver died on a channel
/// so nothing respawns against a dead channel. Then due restarts fire for present /// that no longer answers is retired before its restart could fire, so nothing
/// volumes; then, if no device is adopted, a present device is brought up. /// respawns against a dead channel. Dropping the device frees its slot, so the
/// adopt loop below re-adopts the still-present device on the restarted driver's
/// fresh channel — the rebuild. Then due restarts fire for present volumes.
fn pollTick() void { fn pollTick() void {
for (&devices) |*dev| { for (&devices) |*dev| {
if (dev.used and !isDevicePresent(dev.device_id)) removeDevice(dev); if (dev.used and (!isDevicePresent(dev.device_id) or !channelAlive(dev))) removeDevice(dev);
} }
for (&volumes) |*v| { for (&volumes) |*v| {
if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) { if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) {
@@ -533,6 +555,43 @@ fn onNotification(badge: u64) void {
} }
} }
fn anyVolumeOn(device_id: u64) bool {
for (&volumes) |*v| if (v.used and v.device_id == device_id) return true;
return false;
}
/// A storage device published `medium_changed` — the second removal trigger: the
/// device stays in the tree while its medium leaves or returns (a card reader, an
/// eject). This arrives as a buffered async message, NOT a protocol request, so it
/// never reaches `Serve.dispatch` (its event op number collides with the manager's
/// own `hello`); it is decoded here by hand. Single-volume scope: the event names
/// no device, so `absent` retires every adopted device (its volumes unmount and
/// the poll re-adopts the still-present device with its now-empty medium), and
/// `present` frees any empty adopted device so the poll re-probes and remounts it.
///
/// We act on every edge and do NOT dedup on `change_count`. The driver publishes
/// exactly once per transition, each with a unique monotonic count, so a count is
/// never legitimately re-sent within one subscription — an equality dedup could
/// only ever fire spuriously, and it did: `change_count` restarts at 0 in each
/// driver instance (usb-storage.zig), so a global "last count" carried across a
/// driver restart (S5's own crash-rebuild) mistook the fresh instance's first
/// edge for a re-delivery and dropped a real eject, wedging a mount over absent
/// media. Both branches are idempotent (a freed device stops matching `dev.used`)
/// and the poll reconciles, so reacting to each genuine edge is safe.
fn onMediumEvent(payload: []const u8) void {
const event = block.decodeMediumChanged(payload) orelse return;
if (event.present == 0) {
std.log.info("medium left a storage device; unmounting its volume(s)", .{});
for (&devices) |*dev| {
if (dev.used) removeDevice(dev);
}
} else {
for (&devices) |*dev| {
if (dev.used and !anyVolumeOn(dev.device_id)) removeDevice(dev);
}
}
}
pub fn main(init: process.Init) void { pub fn main(init: process.Init) void {
_ = init; _ = init;
service.run(volume_manager_protocol.message_maximum, .{ service.run(volume_manager_protocol.message_maximum, .{
@@ -540,5 +599,6 @@ pub fn main(init: process.Init) void {
.init = initialise, .init = initialise,
.on_message = onMessage, .on_message = onMessage,
.on_notification = onNotification, .on_notification = onNotification,
.on_buffered_message = onMediumEvent,
}); });
} }
+36
View File
@@ -796,6 +796,42 @@ CASES = [
"expect": r"(?s)fat: mounted /volumes/fat-12345678" "expect": r"(?s)fat: mounted /volumes/fat-12345678"
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting", r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# S5 medium_changed: the SECOND removal trigger. QMP-eject the MEDIUM (the
# block backend, not the device) — the usb-storage device stays in the tree,
# but its TEST UNIT READY poll reports not-ready and publishes medium_changed
# (absent). The volume manager, now a subscriber, runs the same unmount path as
# a device pull. Discrimination: before S5 the manager never subscribed, so the
# event reached no one and the mount persisted (device-presence polling cannot
# see a medium leave while the device stays). One lifecycle, two triggers.
{"name": "volume-medium-change",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"qmp_sequence": [
{"delay": 8, "command": "eject", "arguments": {"device": "bootusb", "force": True}},
],
"expect": r"(?s)fat: mounted /volumes/fat-12345678"
r"[\s\S]*usb-storage: medium absent"
r"[\s\S]*volume-manager: medium left a storage device; unmounting"
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# S5 storage-driver-crash rebuild. The device manager (test-storage-restart
# mode) kills usb-storage once, ~2s in — after its volume mounted. The device
# stays in the tree, so device-presence polling alone would leave fat wedged on
# the dead channel; the volume manager's channel-liveness probe (a geometry()
# that fails on the dead endpoint) must notice, reap the volume, and rebuild on
# the restarted driver's fresh channel — a SECOND mount of the same id-path.
# Discrimination: a pre-S5 manager checks only isDevicePresent (still true), so
# it never reaps and the second mount never appears (it would restart fat on
# the stale channel and crash-loop).
{"name": "volume-driver-restart",
"build_case": "volume-driver-restart",
"smp": 4,
"timeout": 150,
"expect": r"(?s)fat: mounted /volumes/fat-12345678"
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting"
r"[\s\S]*fat: mounted /volumes/fat-12345678",
"fail": r"failing repeatedly; giving up|\[FAIL\]|DANOS-TEST-RESULT: FAIL"},
# Volume-manager discovery + probe (V3a, docs/volume-manager-plan.md). Reuses # Volume-manager discovery + probe (V3a, docs/volume-manager-plan.md). Reuses
# the fat-mount kernel build (the default boot now spawns the volume manager # the fat-mount kernel build (the default boot now spawns the volume manager
# from init.csv). It acquires the mass-storage block channel through the # from init.csv). It acquires the mass-storage block channel through the