docs: the removal lifecycle closes — three triggers, one path (S5)
Storage removal is now robust to all three ways a volume can leave, and the docs
say so. storage-architecture.md and storage-design-rationale.md move medium_changed
from "planned to be consumed" to consumed, and record the third trigger:
- the DEVICE leaving the tree (a pulled stick) — presence polling
- the MEDIUM leaving while its device stays (a reader) — the volume manager
now consumes the pushed medium_changed event
- the storage DRIVER crashing while its device stays — a channel-liveness
geometry() probe reaps the volume and rebuilds it on the restarted driver's
fresh channel; presence polling alone cannot see this (the V4 open edge)
The re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a
physical unplug/replug is bench-verified (QEMU cannot re-present a usb-storage
device_add). The transport-native eject signal (SCSI UNIT ATTENTION, AHCI
PxSSTS, NVMe namespace-change AER) in place of the TEST UNIT READY poll stays
the documented future refinement.
This commit is contained in:
@@ -22,14 +22,17 @@
|
|||||||
> (`system/services/exfat`): full read + write, directories, rename, and on-disk
|
> (`system/services/exfat`): full read + write, directories, rename, and on-disk
|
||||||
> up-case folding, reusing `library/kernel/file-system-harness` wholesale — the
|
> up-case folding, reusing `library/kernel/file-system-harness` wholesale — the
|
||||||
> reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an
|
> reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an
|
||||||
> `exfat-<serial>` id-path. **Still pending**: the `filesystem UUID` rung (ext-
|
> `exfat-<serial>` id-path. Removal is robust to all three triggers now: a
|
||||||
> family superblocks, which need such an engine); the volume manager *consuming*
|
> pulled device (presence polling), a medium that leaves while its device stays
|
||||||
> `medium_changed`
|
> (the volume manager CONSUMES `medium_changed`), and a storage driver that
|
||||||
> (removal is detected by device-presence polling; the event is published but only
|
> crashes while its device stays present (a channel-liveness `geometry()` probe
|
||||||
> a card-reader medium change needs the subscription); the remount-on-replug
|
> reaps the volume and rebuilds it on the restarted driver's fresh channel). The
|
||||||
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller
|
> re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a physical
|
||||||
> device, so it is bench-verified); and arbitration when two volumes both resolve
|
> unplug/replug exercises the same path and is bench-verified (QEMU cannot
|
||||||
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
|
> re-present a usb-storage `device_add`). **Still pending**: the `filesystem UUID`
|
||||||
|
> rung (ext-family superblocks, which need such an engine); and arbitration when
|
||||||
|
> two volumes both resolve the boot markers (S3 mounts both and logs each claim;
|
||||||
|
> picking one is deferred). A few
|
||||||
> markers below are left where a duty is still pending.
|
> markers below are left where a duty is still pending.
|
||||||
|
|
||||||
## The model
|
## The model
|
||||||
@@ -187,25 +190,26 @@ surprise-removal path — kill the filesystem process, retire its mounts,
|
|||||||
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
||||||
error forever (Plan 9's dead-server wart).
|
error forever (Plan 9's dead-server wart).
|
||||||
|
|
||||||
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
The path folds **three triggers into one lifecycle**: the *device* leaving (a
|
||||||
storage driver dies — channel death, the table below), and the *medium*
|
pulled stick — presence polling); the *medium* leaving while the device stays
|
||||||
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
(an SD card pulled from its reader, an ATAPI tray opened, a USB card reader);
|
||||||
tray opened — including USB card readers today). The second trigger is the
|
and a storage *driver crashing* while its device stays in the tree. The second
|
||||||
pushed `medium_changed` event on the block protocol — published today from a
|
trigger is the pushed `medium_changed` event on the block protocol, published
|
||||||
TEST UNIT READY poll; still *planned* is the volume manager *consuming* it
|
from a TEST UNIT READY poll — the volume manager now **consumes** it (subscribed
|
||||||
(today removal is driven only by device-presence polling) and translating the
|
per device), running the same kill-retire path and re-probing on medium return,
|
||||||
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe
|
so a swapped card is never served with the previous card's filesystem state. The
|
||||||
namespace-change AER) in place of the poll. On the event the volume manager
|
third is caught by a channel-liveness `geometry()` probe: presence polling alone
|
||||||
runs the same kill-retire path, then re-probes on medium return exactly as on
|
sees the device still present, but the channel is dead, so the manager reaps the
|
||||||
device return. Without it, a swapped card would be served with the previous
|
volume and rebuilds it on the restarted driver's fresh channel. Still *planned*
|
||||||
card's filesystem state.
|
is translating the transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS,
|
||||||
|
NVMe namespace-change AER) in place of the presence poll.
|
||||||
|
|
||||||
| Layer | Observes | Must do | Guarantees |
|
| Layer | Observes | Must do | Guarantees |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
||||||
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||||||
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||||||
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
| Volume manager *(built)* | a device leaving the tree (poll), a `medium_changed` event, or a dead channel under a still-present device (a crashed driver — `geometry()` liveness probe) | kill that volume's filesystem service (its mounts retire), then re-adopt + remount on return or on the restarted driver's fresh channel | one removal path for all three triggers; mounts never dangle; the manager never serves from behind a dead channel |
|
||||||
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
|
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
|
||||||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
||||||
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||||||
|
|||||||
@@ -209,18 +209,23 @@ matrix-proven shape; genuinely open.
|
|||||||
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
||||||
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
||||||
filesystem) once danos outgrows FAT; per-process namespaces.
|
filesystem) once danos outgrows FAT; per-process namespaces.
|
||||||
7. **The media-presence event** (settled in principle; lands with the volume
|
7. **The media-presence event** (the consuming half is BUILT; the
|
||||||
manager): the block protocol gains a pushed event — `medium_changed`, with
|
transport-native signal stays future): the block protocol carries a pushed
|
||||||
present/absent and a change counter — produced by the storage driver from
|
event — `medium_changed`, with present/absent and a change counter —
|
||||||
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
|
produced today by the storage driver from a TEST UNIT READY poll (the
|
||||||
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
|
transport's native signal — SCSI UNIT ATTENTION, PxSSTS for AHCI,
|
||||||
consumed by the volume manager, which runs the SAME kill-retire-remount
|
namespace-change AER for NVMe — is the future refinement in place of the
|
||||||
path it runs on channel death — one lifecycle, two triggers. The driver
|
poll) and now **consumed** by the volume manager, which subscribes per
|
||||||
reports presence, never content; a pushed event carries no capability,
|
device and runs the SAME kill-retire-remount path it runs on channel death.
|
||||||
which the kernel already guarantees. The device staying while its medium
|
The driver reports presence, never content; a pushed event carries no
|
||||||
leaves is the one removable-media case the channel-death trigger cannot
|
capability, which the kernel already guarantees. The device staying while
|
||||||
see; without this event a swapped SD card would be served with the old
|
its medium leaves is the one removable-media case the channel-death trigger
|
||||||
card's filesystem state.
|
cannot see; without this event a swapped SD card would be served with the
|
||||||
|
old card's filesystem state. A THIRD trigger closes the last gap — a
|
||||||
|
storage driver that *crashes* while its device stays present: channel death
|
||||||
|
there is invisible to presence polling, so the volume manager probes channel
|
||||||
|
liveness (`geometry()`) each tick and reaps-then-rebuilds the volume on the
|
||||||
|
restarted driver's fresh channel. One lifecycle, three triggers.
|
||||||
8. **Volume identity, and the mount map as danos's fstab** (settled). The
|
8. **Volume identity, and the mount map as danos's fstab** (settled). The
|
||||||
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
|
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
|
||||||
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
||||||
|
|||||||
Reference in New Issue
Block a user