Compare commits
6
Commits
c47215821c
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4037e746aa | ||
|
|
4dfb5012c0 | ||
|
|
061eb7c004 | ||
|
|
5dc966838a | ||
|
|
700452dc4e | ||
|
|
0faa0fd21b |
@@ -22,14 +22,17 @@
|
|||||||
> (`system/services/exfat`): full read + write, directories, rename, and on-disk
|
> (`system/services/exfat`): full read + write, directories, rename, and on-disk
|
||||||
> up-case folding, reusing `library/kernel/file-system-harness` wholesale — the
|
> up-case folding, reusing `library/kernel/file-system-harness` wholesale — the
|
||||||
> reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an
|
> reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an
|
||||||
> `exfat-<serial>` id-path. **Still pending**: the `filesystem UUID` rung (ext-
|
> `exfat-<serial>` id-path. Removal is robust to all three triggers now: a
|
||||||
> family superblocks, which need such an engine); the volume manager *consuming*
|
> pulled device (presence polling), a medium that leaves while its device stays
|
||||||
> `medium_changed`
|
> (the volume manager CONSUMES `medium_changed`), and a storage driver that
|
||||||
> (removal is detected by device-presence polling; the event is published but only
|
> crashes while its device stays present (a channel-liveness `geometry()` probe
|
||||||
> a card-reader medium change needs the subscription); the remount-on-replug
|
> reaps the volume and rebuilds it on the restarted driver's fresh channel). The
|
||||||
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller
|
> re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a physical
|
||||||
> device, so it is bench-verified); and arbitration when two volumes both resolve
|
> unplug/replug exercises the same path but is bench-pending (QEMU cannot
|
||||||
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
|
> re-present a usb-storage `device_add`). **Still pending**: the `filesystem UUID`
|
||||||
|
> rung (ext-family superblocks, which need such an engine); and arbitration when
|
||||||
|
> two volumes both resolve the boot markers (S3 mounts both and logs each claim;
|
||||||
|
> picking one is deferred). A few
|
||||||
> markers below are left where a duty is still pending.
|
> markers below are left where a duty is still pending.
|
||||||
|
|
||||||
## The model
|
## The model
|
||||||
@@ -187,25 +190,26 @@ surprise-removal path — kill the filesystem process, retire its mounts,
|
|||||||
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
||||||
error forever (Plan 9's dead-server wart).
|
error forever (Plan 9's dead-server wart).
|
||||||
|
|
||||||
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
The path folds **three triggers into one lifecycle**: the *device* leaving (a
|
||||||
storage driver dies — channel death, the table below), and the *medium*
|
pulled stick — presence polling); the *medium* leaving while the device stays
|
||||||
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
(an SD card pulled from its reader, an ATAPI tray opened, a USB card reader);
|
||||||
tray opened — including USB card readers today). The second trigger is the
|
and a storage *driver crashing* while its device stays in the tree. The second
|
||||||
pushed `medium_changed` event on the block protocol — published today from a
|
trigger is the pushed `medium_changed` event on the block protocol, published
|
||||||
TEST UNIT READY poll; still *planned* is the volume manager *consuming* it
|
from a TEST UNIT READY poll — the volume manager now **consumes** it (subscribed
|
||||||
(today removal is driven only by device-presence polling) and translating the
|
per device), running the same kill-retire path and re-probing on medium return,
|
||||||
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe
|
so a swapped card is never served with the previous card's filesystem state. The
|
||||||
namespace-change AER) in place of the poll. On the event the volume manager
|
third is caught by a channel-liveness `geometry()` probe: presence polling alone
|
||||||
runs the same kill-retire path, then re-probes on medium return exactly as on
|
sees the device still present, but the channel is dead, so the manager reaps the
|
||||||
device return. Without it, a swapped card would be served with the previous
|
volume and rebuilds it on the restarted driver's fresh channel. Still *planned*
|
||||||
card's filesystem state.
|
is translating the transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS,
|
||||||
|
NVMe namespace-change AER) in place of the presence poll.
|
||||||
|
|
||||||
| Layer | Observes | Must do | Guarantees |
|
| Layer | Observes | Must do | Guarantees |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
||||||
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||||||
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||||||
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
| Volume manager *(built)* | a device leaving the tree (poll), a `medium_changed` event, or a dead channel under a still-present device (a crashed driver — `geometry()` liveness probe) | kill that volume's filesystem service (its mounts retire), then re-adopt + remount on return or on the restarted driver's fresh channel | one removal path for all three triggers; mounts never dangle; the manager never serves from behind a dead channel |
|
||||||
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
|
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
|
||||||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
||||||
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||||||
|
|||||||
@@ -209,18 +209,23 @@ matrix-proven shape; genuinely open.
|
|||||||
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
||||||
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
||||||
filesystem) once danos outgrows FAT; per-process namespaces.
|
filesystem) once danos outgrows FAT; per-process namespaces.
|
||||||
7. **The media-presence event** (settled in principle; lands with the volume
|
7. **The media-presence event** (the consuming half is BUILT; the
|
||||||
manager): the block protocol gains a pushed event — `medium_changed`, with
|
transport-native signal stays future): the block protocol carries a pushed
|
||||||
present/absent and a change counter — produced by the storage driver from
|
event — `medium_changed`, with present/absent and a change counter —
|
||||||
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
|
produced today by the storage driver from a TEST UNIT READY poll (the
|
||||||
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
|
transport's native signal — SCSI UNIT ATTENTION, PxSSTS for AHCI,
|
||||||
consumed by the volume manager, which runs the SAME kill-retire-remount
|
namespace-change AER for NVMe — is the future refinement in place of the
|
||||||
path it runs on channel death — one lifecycle, two triggers. The driver
|
poll) and now **consumed** by the volume manager, which subscribes per
|
||||||
reports presence, never content; a pushed event carries no capability,
|
device and runs the SAME kill-retire-remount path it runs on channel death.
|
||||||
which the kernel already guarantees. The device staying while its medium
|
The driver reports presence, never content; a pushed event carries no
|
||||||
leaves is the one removable-media case the channel-death trigger cannot
|
capability, which the kernel already guarantees. The device staying while
|
||||||
see; without this event a swapped SD card would be served with the old
|
its medium leaves is the one removable-media case the channel-death trigger
|
||||||
card's filesystem state.
|
cannot see; without this event a swapped SD card would be served with the
|
||||||
|
old card's filesystem state. A THIRD trigger closes the last gap — a
|
||||||
|
storage driver that *crashes* while its device stays present: channel death
|
||||||
|
there is invisible to presence polling, so the volume manager probes channel
|
||||||
|
liveness (`geometry()`) each tick and reaps-then-rebuilds the volume on the
|
||||||
|
restarted driver's fresh channel. One lifecycle, three triggers.
|
||||||
8. **Volume identity, and the mount map as danos's fstab** (settled). The
|
8. **Volume identity, and the mount map as danos's fstab** (settled). The
|
||||||
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
|
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
|
||||||
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
||||||
@@ -246,8 +251,9 @@ matrix-proven shape; genuinely open.
|
|||||||
whether USB port, hub depth, SATA port, or a stick that left as USB and
|
whether USB port, hub depth, SATA port, or a stick that left as USB and
|
||||||
returned in a SATA dock; **replug remounts at the same path** (the id-path is
|
returned in a SATA dock; **replug remounts at the same path** (the id-path is
|
||||||
content-derived, so a volume returns to `/volumes/<id>` wherever it reappears;
|
content-derived, so a volume returns to `/volumes/<id>` wherever it reappears;
|
||||||
remount-on-replug end-to-end is bench-verified, not QEMU-tested, because QEMU
|
the re-adopt+remount code path is QEMU-proven by the driver-crash rebuild, but
|
||||||
can't re-present the boot-controller device); **the boot volume** is the volume
|
remount on a *physical* replug end-to-end is bench-pending, not QEMU-testable,
|
||||||
|
because QEMU can't re-present the boot-controller device); **the boot volume** is the volume
|
||||||
that resolves `/system/configuration` on its own media, findable on any port or
|
that resolves `/system/configuration` on its own media, findable on any port or
|
||||||
partition; and **duplicate identity is a known S4 gap** — two cloned sticks
|
partition; and **duplicate identity is a known S4 gap** — two cloned sticks
|
||||||
share one content id, so today they collide on `/volumes/<id>` (the kernel
|
share one content id, so today they collide on `/volumes/<id>` (the kernel
|
||||||
|
|||||||
@@ -7,12 +7,25 @@
|
|||||||
//! `runtime.dma.alloc`), so whole sectors move without crossing the IPC size
|
//! `runtime.dma.alloc`), so whole sectors move without crossing the IPC size
|
||||||
//! limit — the same handoff usb-storage uses toward the controller.
|
//! limit — the same handoff usb-storage uses toward the controller.
|
||||||
|
|
||||||
|
const std = @import("std");
|
||||||
const envelope = @import("envelope");
|
const envelope = @import("envelope");
|
||||||
const ipc = @import("ipc");
|
const ipc = @import("ipc");
|
||||||
const block_protocol = @import("block-protocol");
|
const block_protocol = @import("block-protocol");
|
||||||
|
|
||||||
const Protocol = block_protocol.Protocol;
|
const Protocol = block_protocol.Protocol;
|
||||||
|
|
||||||
|
/// The medium_changed event payload, re-exported so a consumer decodes it without
|
||||||
|
/// reaching into the wire-format module.
|
||||||
|
pub const MediumChanged = block_protocol.MediumChanged;
|
||||||
|
|
||||||
|
/// Decode a medium_changed event from a buffered-message payload a subscriber
|
||||||
|
/// received (a `Received.isMessage` wake). Null if the bytes are too short to be
|
||||||
|
/// one — a caller ignores anything that is not a well-formed event.
|
||||||
|
pub fn decodeMediumChanged(payload: []const u8) ?MediumChanged {
|
||||||
|
if (payload.len < envelope.prefix_size + @sizeOf(MediumChanged)) return null;
|
||||||
|
return std.mem.bytesToValue(MediumChanged, payload[envelope.prefix_size..][0..@sizeOf(MediumChanged)]);
|
||||||
|
}
|
||||||
|
|
||||||
pub const Geometry = struct { block_size: u32, block_count: u64 };
|
pub const Geometry = struct { block_size: u32, block_count: u64 };
|
||||||
|
|
||||||
pub const Device = struct {
|
pub const Device = struct {
|
||||||
@@ -88,6 +101,30 @@ pub const Device = struct {
|
|||||||
if (status.status != 0) return null;
|
if (status.status != 0) return null;
|
||||||
return reply[0..answer.len];
|
return reply[0..answer.len];
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Subscribe `subscriber` (an endpoint) to this device's medium_changed
|
||||||
|
/// events: the reserved `subscribe` verb carries the subscriber's endpoint as
|
||||||
|
/// the capability, and the driver then ipc.sends each medium transition to it.
|
||||||
|
pub fn subscribeMedium(self: Device, subscriber: ipc.Handle) bool {
|
||||||
|
var packet: [block_protocol.message_maximum]u8 = undefined;
|
||||||
|
const framed = envelope.encodeSubscribe(0, &packet) orelse return false; // interest 0: every event (block has one)
|
||||||
|
var reply: [block_protocol.message_maximum]u8 = undefined;
|
||||||
|
const answer = ipc.callCap(self.endpoint, framed, &reply, subscriber) catch return false;
|
||||||
|
const status = envelope.statusOf(reply[0..answer.len]) orelse return false;
|
||||||
|
return status.status == 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Unsubscribe from this device's medium_changed events. Call before closing
|
||||||
|
/// the channel so the driver's bounded subscriber table frees the slot rather
|
||||||
|
/// than holding a dead endpoint until an exit sweep notices.
|
||||||
|
pub fn unsubscribeMedium(self: Device) bool {
|
||||||
|
var packet: [block_protocol.message_maximum]u8 = undefined;
|
||||||
|
const framed = envelope.encodeUnsubscribe(&packet) orelse return false;
|
||||||
|
var reply: [block_protocol.message_maximum]u8 = undefined;
|
||||||
|
const answer = ipc.callCap(self.endpoint, framed, &reply, null) catch return false;
|
||||||
|
const status = envelope.statusOf(reply[0..answer.len]) orelse return false;
|
||||||
|
return status.status == 0;
|
||||||
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
// There is deliberately no open-by-name here: `block` is not a registry name.
|
// There is deliberately no open-by-name here: `block` is not a registry name.
|
||||||
|
|||||||
@@ -67,6 +67,15 @@ pub const Callbacks = struct {
|
|||||||
/// A notification that is not a signal — a subscribed exit event, a bound
|
/// A notification that is not a signal — a subscribed exit event, a bound
|
||||||
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
|
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
|
||||||
on_notification: ?*const fn (badge: u64) void = null,
|
on_notification: ?*const fn (badge: u64) void = null,
|
||||||
|
/// A buffered async message (`Received.isMessage`): a pushed event from a
|
||||||
|
/// provider this service subscribed to, its payload in the receive buffer.
|
||||||
|
/// Unlike `on_message`, it never goes through the protocol dispatch — so an
|
||||||
|
/// event whose reserved op number collides with one of this service's own
|
||||||
|
/// verbs (a `block` `medium_changed` reaching the volume manager, whose own
|
||||||
|
/// protocol numbers `hello` the same) is decoded by hand here, not
|
||||||
|
/// mis-dispatched. Default null: the badge alone still reaches
|
||||||
|
/// `on_notification`, exactly as before this callback existed.
|
||||||
|
on_buffered_message: ?*const fn (message: []const u8) void = null,
|
||||||
/// The reload signal. Default: ignored.
|
/// The reload signal. Default: ignored.
|
||||||
on_reload: ?*const fn () void = null,
|
on_reload: ?*const fn () void = null,
|
||||||
/// The terminate signal, called before the loop returns. The clean exit is
|
/// The terminate signal, called before the loop returns. The clean exit is
|
||||||
@@ -372,6 +381,13 @@ pub fn run(comptime maximum_message: usize, callbacks: Callbacks) void {
|
|||||||
if (got.isChildExit()) {
|
if (got.isChildExit()) {
|
||||||
if (callbacks.subscribers) |subscribers| subscribers.forget(got.childProcessId());
|
if (callbacks.subscribers) |subscribers| subscribers.forget(got.childProcessId());
|
||||||
}
|
}
|
||||||
|
// A buffered async message (a pushed event) carries a payload; hand it
|
||||||
|
// to the service that asked for it. The badge still reaches
|
||||||
|
// on_notification below, so a coalesced timer/exit riding the same wake
|
||||||
|
// is not lost — and a service without this callback is unchanged.
|
||||||
|
if (got.isMessage()) {
|
||||||
|
if (callbacks.on_buffered_message) |onBuffered| onBuffered(receive[0..got.len]);
|
||||||
|
}
|
||||||
if (callbacks.on_notification) |onNotification| onNotification(got.badge);
|
if (callbacks.on_notification) |onNotification| onNotification(got.badge);
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -79,6 +79,7 @@
|
|||||||
# 'kernel' as the supervisor. Nothing else changes: the binary must still match.
|
# 'kernel' as the supervisor. Nothing else changes: the binary must still match.
|
||||||
/system/services/input, kernel, bind, input
|
/system/services/input, kernel, bind, input
|
||||||
/system/services/device-manager, kernel, bind, device-manager
|
/system/services/device-manager, kernel, bind, device-manager
|
||||||
|
/system/services/volume-manager, kernel, bind, volume-manager
|
||||||
/system/services/fat, kernel, bind, vfs
|
/system/services/fat, kernel, bind, vfs
|
||||||
/system/services/display, kernel, bind, display
|
/system/services/display, kernel, bind, display
|
||||||
/system/services/discovery, kernel, bind, power
|
/system/services/discovery, kernel, bind, power
|
||||||
@@ -107,6 +108,9 @@
|
|||||||
# The volume manager reaches the device manager to be routed to each storage
|
# The volume manager reaches the device manager to be routed to each storage
|
||||||
# provider's block channel, then confines a filesystem to each volume.
|
# provider's block channel, then confines a filesystem to each volume.
|
||||||
/system/services/volume-manager, /system/services/init, open, device-manager
|
/system/services/volume-manager, /system/services/init, open, device-manager
|
||||||
|
# ...and again under the kernel supervisor for the manual-tree drills (S5's
|
||||||
|
# volume-driver-restart spawns the volume manager directly, not via init).
|
||||||
|
/system/services/volume-manager, kernel, open, device-manager
|
||||||
/system/services/display, /system/services/init, open, scanout
|
/system/services/display, /system/services/init, open, scanout
|
||||||
/system/services/display, /system/services/init, open, display
|
/system/services/display, /system/services/init, open, display
|
||||||
/system/services/display, /system/services/init, open, input
|
/system/services/display, /system/services/init, open, input
|
||||||
|
|||||||
|
Can't render this file because it contains an unexpected character in line 12 and column 15.
|
@@ -229,6 +229,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
|||||||
fatMountTest(boot_information);
|
fatMountTest(boot_information);
|
||||||
} else if (eql(case, "exfat-volume")) {
|
} else if (eql(case, "exfat-volume")) {
|
||||||
exfatVolumeTest(boot_information);
|
exfatVolumeTest(boot_information);
|
||||||
|
} else if (eql(case, "volume-driver-restart")) {
|
||||||
|
volumeDriverRestartTest(boot_information);
|
||||||
} else if (eql(case, "device-list")) {
|
} else if (eql(case, "device-list")) {
|
||||||
deviceListTest(boot_information);
|
deviceListTest(boot_information);
|
||||||
} else if (eql(case, "pci-scan")) {
|
} else if (eql(case, "pci-scan")) {
|
||||||
@@ -3685,6 +3687,44 @@ fn displayReattachTest(boot_information: *const BootInformation) void {
|
|||||||
while (true) scheduler.yield();
|
while (true) scheduler.yield();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Storage-driver-crash rebuild (S5): the device manager runs in
|
||||||
|
/// "test-storage-restart" mode and kills the usb-storage driver once, a moment
|
||||||
|
/// after its volume has mounted. The driver's device stays in the tree, so the
|
||||||
|
/// volume manager's presence poll alone would miss the death and leave fat wedged
|
||||||
|
/// on a dead channel; its channel-liveness probe must notice, reap the volume, and
|
||||||
|
/// rebuild on the restarted driver's fresh channel — a SECOND mount of the same
|
||||||
|
/// id-path is the proof. (A pre-S5 manager, checking only device presence, never
|
||||||
|
/// reaps, so the second mount never appears.)
|
||||||
|
fn volumeDriverRestartTest(boot_information: *const BootInformation) void {
|
||||||
|
log("DANOS-TEST-BEGIN: volume-driver-restart\n", .{});
|
||||||
|
if (boot_information.initial_ramdisk_len == 0) {
|
||||||
|
check("bootloader handed over an initial_ramdisk", false);
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
|
||||||
|
const rd = initial_ramdisk.Reader.init(image) orelse {
|
||||||
|
check("initial_ramdisk image is valid", false);
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
process.setInitialRamdisk(image);
|
||||||
|
_ = spawnRegistry(rd);
|
||||||
|
var manager: u32 = 0;
|
||||||
|
var i: u32 = 0;
|
||||||
|
while (i < rd.count) : (i += 1) {
|
||||||
|
const item = rd.entry(i) orelse continue;
|
||||||
|
if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
|
||||||
|
manager = process.spawnProcessSupervised(item.blob, 4, &.{ item.name, "test-storage-restart" }, scheduler.currentId(), null) catch 0;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
check("device-manager spawned (test-storage-restart mode)", manager != 0);
|
||||||
|
check("volume-manager spawned", spawnNamed(rd, "volume-manager"));
|
||||||
|
check("fat-test client spawned", spawnNamed(rd, "fat-test"));
|
||||||
|
scheduler.setPriority(1); // below the tree, so it runs
|
||||||
|
while (true) scheduler.yield();
|
||||||
|
}
|
||||||
|
|
||||||
/// Process arguments, end to end: spawn args-echo bare (its argv[0] is the
|
/// Process arguments, end to end: spawn args-echo bare (its argv[0] is the
|
||||||
/// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through
|
/// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through
|
||||||
/// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument
|
/// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument
|
||||||
|
|||||||
@@ -170,6 +170,8 @@ var test_usb_killed = false;
|
|||||||
var test_pci_restart_mode = false;
|
var test_pci_restart_mode = false;
|
||||||
var test_scanout_restart_mode = false;
|
var test_scanout_restart_mode = false;
|
||||||
var test_scanout_killed = false;
|
var test_scanout_killed = false;
|
||||||
|
var test_storage_restart_mode = false;
|
||||||
|
var test_storage_killed = false;
|
||||||
var test_kill_pid: u32 = 0;
|
var test_kill_pid: u32 = 0;
|
||||||
var test_kill_due_ns: u64 = 0;
|
var test_kill_due_ns: u64 = 0;
|
||||||
|
|
||||||
@@ -600,6 +602,16 @@ fn onHello(_: void, invocation: Invocation(device_manager_protocol.Hello), _: An
|
|||||||
test_kill_due_ns = time.clock() + 1_500_000_000;
|
test_kill_due_ns = time.clock() + 1_500_000_000;
|
||||||
_ = time.timerOnce(manager_endpoint, 1600);
|
_ = time.timerOnce(manager_endpoint, 1600);
|
||||||
}
|
}
|
||||||
|
// Storage-driver-crash drill (S5): once, a moment after usb-storage hellos —
|
||||||
|
// long enough that its volume has mounted — kill it. The manager re-delegates
|
||||||
|
// the still-present device to a restarted driver on a fresh channel; the volume
|
||||||
|
// manager's channel-liveness probe must notice the dead channel and rebuild.
|
||||||
|
if (test_storage_restart_mode and !test_storage_killed and std.mem.eql(u8, driver.name(), "/system/drivers/usb-storage")) {
|
||||||
|
test_storage_killed = true;
|
||||||
|
test_kill_pid = invocation.sender;
|
||||||
|
test_kill_due_ns = time.clock() + 2_000_000_000; // after the ~0.6s mount
|
||||||
|
_ = time.timerOnce(manager_endpoint, 2100);
|
||||||
|
}
|
||||||
return 0;
|
return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -763,6 +775,7 @@ pub fn main(init: process.Init) void {
|
|||||||
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
|
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
|
||||||
test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart");
|
test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart");
|
||||||
test_scanout_restart_mode = std.mem.eql(u8, mode, "test-scanout-restart");
|
test_scanout_restart_mode = std.mem.eql(u8, mode, "test-scanout-restart");
|
||||||
|
test_storage_restart_mode = std.mem.eql(u8, mode, "test-storage-restart");
|
||||||
}
|
}
|
||||||
service.run(device_manager_protocol.message_maximum, .{
|
service.run(device_manager_protocol.message_maximum, .{
|
||||||
.service = "device-manager",
|
.service = "device-manager",
|
||||||
|
|||||||
@@ -306,6 +306,12 @@ fn bringUpVolume() bool {
|
|||||||
return false;
|
return false;
|
||||||
};
|
};
|
||||||
dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device };
|
dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device };
|
||||||
|
// Consume this device's medium_changed events (the second of the removal
|
||||||
|
// lifecycle's two triggers: the device stays in the tree while its medium
|
||||||
|
// leaves — a card reader, an eject). Best effort: a provider that never
|
||||||
|
// publishes the event simply never wakes us, and device-pull is still caught
|
||||||
|
// by the presence poll.
|
||||||
|
_ = opened.device.subscribeMedium(service_endpoint);
|
||||||
const device = opened.device;
|
const device = opened.device;
|
||||||
// Attach the read buffer to THIS device (a no-op without an enforcing IOMMU).
|
// Attach the read buffer to THIS device (a no-op without an enforcing IOMMU).
|
||||||
// The handle is kept, not closed, so it can be re-attached after a replug. A
|
// The handle is kept, not closed, so it can be re-attached after a replug. A
|
||||||
@@ -370,6 +376,10 @@ fn bringUpVolume() bool {
|
|||||||
/// Close a device's channel and free its slot. No volumes are touched (the caller
|
/// Close a device's channel and free its slot. No volumes are touched (the caller
|
||||||
/// ensures none remain, or there never were any).
|
/// ensures none remain, or there never were any).
|
||||||
fn dropDevice(dev: *StorageDevice) void {
|
fn dropDevice(dev: *StorageDevice) void {
|
||||||
|
// Free the driver's subscriber slot before the channel closes. On a still-live
|
||||||
|
// channel (a medium eject) this frees the slot; on a dead one (a device pull)
|
||||||
|
// the call fails fast and the exit sweep frees it anyway.
|
||||||
|
_ = dev.channel.unsubscribeMedium();
|
||||||
_ = ipc.close(dev.channel.endpoint);
|
_ = ipc.close(dev.channel.endpoint);
|
||||||
dev.* = .{};
|
dev.* = .{};
|
||||||
}
|
}
|
||||||
@@ -395,13 +405,25 @@ fn removeDevice(dev: *StorageDevice) void {
|
|||||||
dropDevice(dev);
|
dropDevice(dev);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Whether a device's block channel still answers — a geometry() probe. A storage
|
||||||
|
/// driver that DIED while its device stays in the tree (it crashed; the device
|
||||||
|
/// manager will re-delegate the device to a restarted driver on a FRESH channel)
|
||||||
|
/// leaves a dead channel here, even though isDevicePresent still reports the device
|
||||||
|
/// present. geometry() on the dead endpoint fails fast, so this catches the crash
|
||||||
|
/// that presence-polling alone cannot — the V4 review's open edge.
|
||||||
|
fn channelAlive(dev: *StorageDevice) bool {
|
||||||
|
return dev.channel.geometry() != null;
|
||||||
|
}
|
||||||
|
|
||||||
/// One poll tick. Device removal is reconciled FIRST and supersedes a pending
|
/// One poll tick. Device removal is reconciled FIRST and supersedes a pending
|
||||||
/// restart: a volume whose device left is retired before its restart could fire,
|
/// restart: a volume whose device left (a pull) OR whose driver died on a channel
|
||||||
/// so nothing respawns against a dead channel. Then due restarts fire for present
|
/// that no longer answers is retired before its restart could fire, so nothing
|
||||||
/// volumes; then, if no device is adopted, a present device is brought up.
|
/// respawns against a dead channel. Dropping the device frees its slot, so the
|
||||||
|
/// adopt loop below re-adopts the still-present device on the restarted driver's
|
||||||
|
/// fresh channel — the rebuild. Then due restarts fire for present volumes.
|
||||||
fn pollTick() void {
|
fn pollTick() void {
|
||||||
for (&devices) |*dev| {
|
for (&devices) |*dev| {
|
||||||
if (dev.used and !isDevicePresent(dev.device_id)) removeDevice(dev);
|
if (dev.used and (!isDevicePresent(dev.device_id) or !channelAlive(dev))) removeDevice(dev);
|
||||||
}
|
}
|
||||||
for (&volumes) |*v| {
|
for (&volumes) |*v| {
|
||||||
if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) {
|
if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) {
|
||||||
@@ -533,6 +555,43 @@ fn onNotification(badge: u64) void {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn anyVolumeOn(device_id: u64) bool {
|
||||||
|
for (&volumes) |*v| if (v.used and v.device_id == device_id) return true;
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A storage device published `medium_changed` — the second removal trigger: the
|
||||||
|
/// device stays in the tree while its medium leaves or returns (a card reader, an
|
||||||
|
/// eject). This arrives as a buffered async message, NOT a protocol request, so it
|
||||||
|
/// never reaches `Serve.dispatch` (its event op number collides with the manager's
|
||||||
|
/// own `hello`); it is decoded here by hand. Single-volume scope: the event names
|
||||||
|
/// no device, so `absent` retires every adopted device (its volumes unmount and
|
||||||
|
/// the poll re-adopts the still-present device with its now-empty medium), and
|
||||||
|
/// `present` frees any empty adopted device so the poll re-probes and remounts it.
|
||||||
|
///
|
||||||
|
/// We act on every edge and do NOT dedup on `change_count`. The driver publishes
|
||||||
|
/// exactly once per transition, each with a unique monotonic count, so a count is
|
||||||
|
/// never legitimately re-sent within one subscription — an equality dedup could
|
||||||
|
/// only ever fire spuriously, and it did: `change_count` restarts at 0 in each
|
||||||
|
/// driver instance (usb-storage.zig), so a global "last count" carried across a
|
||||||
|
/// driver restart (S5's own crash-rebuild) mistook the fresh instance's first
|
||||||
|
/// edge for a re-delivery and dropped a real eject, wedging a mount over absent
|
||||||
|
/// media. Both branches are idempotent (a freed device stops matching `dev.used`)
|
||||||
|
/// and the poll reconciles, so reacting to each genuine edge is safe.
|
||||||
|
fn onMediumEvent(payload: []const u8) void {
|
||||||
|
const event = block.decodeMediumChanged(payload) orelse return;
|
||||||
|
if (event.present == 0) {
|
||||||
|
std.log.info("medium left a storage device; unmounting its volume(s)", .{});
|
||||||
|
for (&devices) |*dev| {
|
||||||
|
if (dev.used) removeDevice(dev);
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
for (&devices) |*dev| {
|
||||||
|
if (dev.used and !anyVolumeOn(dev.device_id)) removeDevice(dev);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
pub fn main(init: process.Init) void {
|
pub fn main(init: process.Init) void {
|
||||||
_ = init;
|
_ = init;
|
||||||
service.run(volume_manager_protocol.message_maximum, .{
|
service.run(volume_manager_protocol.message_maximum, .{
|
||||||
@@ -540,5 +599,6 @@ pub fn main(init: process.Init) void {
|
|||||||
.init = initialise,
|
.init = initialise,
|
||||||
.on_message = onMessage,
|
.on_message = onMessage,
|
||||||
.on_notification = onNotification,
|
.on_notification = onNotification,
|
||||||
|
.on_buffered_message = onMediumEvent,
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -796,6 +796,42 @@ CASES = [
|
|||||||
"expect": r"(?s)fat: mounted /volumes/fat-12345678"
|
"expect": r"(?s)fat: mounted /volumes/fat-12345678"
|
||||||
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting",
|
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting",
|
||||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
|
# S5 medium_changed: the SECOND removal trigger. QMP-eject the MEDIUM (the
|
||||||
|
# block backend, not the device) — the usb-storage device stays in the tree,
|
||||||
|
# but its TEST UNIT READY poll reports not-ready and publishes medium_changed
|
||||||
|
# (absent). The volume manager, now a subscriber, runs the same unmount path as
|
||||||
|
# a device pull. Discrimination: before S5 the manager never subscribed, so the
|
||||||
|
# event reached no one and the mount persisted (device-presence polling cannot
|
||||||
|
# see a medium leave while the device stays). One lifecycle, two triggers.
|
||||||
|
{"name": "volume-medium-change",
|
||||||
|
"build_case": "fat-mount",
|
||||||
|
"smp": 4,
|
||||||
|
"timeout": 150,
|
||||||
|
"qmp_sequence": [
|
||||||
|
{"delay": 8, "command": "eject", "arguments": {"device": "bootusb", "force": True}},
|
||||||
|
],
|
||||||
|
"expect": r"(?s)fat: mounted /volumes/fat-12345678"
|
||||||
|
r"[\s\S]*usb-storage: medium absent"
|
||||||
|
r"[\s\S]*volume-manager: medium left a storage device; unmounting"
|
||||||
|
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting",
|
||||||
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
|
# S5 storage-driver-crash rebuild. The device manager (test-storage-restart
|
||||||
|
# mode) kills usb-storage once, ~2s in — after its volume mounted. The device
|
||||||
|
# stays in the tree, so device-presence polling alone would leave fat wedged on
|
||||||
|
# the dead channel; the volume manager's channel-liveness probe (a geometry()
|
||||||
|
# that fails on the dead endpoint) must notice, reap the volume, and rebuild on
|
||||||
|
# the restarted driver's fresh channel — a SECOND mount of the same id-path.
|
||||||
|
# Discrimination: a pre-S5 manager checks only isDevicePresent (still true), so
|
||||||
|
# it never reaps and the second mount never appears (it would restart fat on
|
||||||
|
# the stale channel and crash-loop).
|
||||||
|
{"name": "volume-driver-restart",
|
||||||
|
"build_case": "volume-driver-restart",
|
||||||
|
"smp": 4,
|
||||||
|
"timeout": 150,
|
||||||
|
"expect": r"(?s)fat: mounted /volumes/fat-12345678"
|
||||||
|
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting"
|
||||||
|
r"[\s\S]*fat: mounted /volumes/fat-12345678",
|
||||||
|
"fail": r"failing repeatedly; giving up|\[FAIL\]|DANOS-TEST-RESULT: FAIL"},
|
||||||
# Volume-manager discovery + probe (V3a, docs/volume-manager-plan.md). Reuses
|
# Volume-manager discovery + probe (V3a, docs/volume-manager-plan.md). Reuses
|
||||||
# the fat-mount kernel build (the default boot now spawns the volume manager
|
# the fat-mount kernel build (the default boot now spawns the volume manager
|
||||||
# from init.csv). It acquires the mass-storage block channel through the
|
# from init.csv). It acquires the mass-storage block channel through the
|
||||||
|
|||||||
Reference in New Issue
Block a user