usb/fat: transfer events matched by slot+endpoint; storage failures heal

B2 — the 1-in-3 boot-time READ CAPACITY failure, root-caused: the xHCI
library's awaitTransfer claimed ANY unclaimed transfer event as its own
completion. An interrupt-endpoint event whose TRB pointer no longer
matched the armed subscription (an error or stale completion from the
keyboard/mouse polling concurrently with storage bring-up) fell through
and was misread as the bulk transfer's completion — desynchronizing the
mass-storage bulk protocol in controller state that SURVIVED driver
restarts, so every retry failed too. Awaited transfers now match the
event's slot id and endpoint DCI; foreign events are dropped and named.
Twelve consecutive runs of the previously-flaky cases pass; the full
suite is green with none of its old intermittents.

B1 — and when storage does fail transiently, the system now heals
instead of giving up forever: a nonzero exit maps to ExitReason.aborted
(a deliberate FAILURE exit — supervisors restart those with backoff,
unlike a clean .exited), usb-storage exits nonzero when a PRESENT
device fails bring-up, and the fat service no longer blocks its harness
polling for a block device and then dies — it serves immediately
(requests fail politely), retries on a 500 ms timer, and mounts
whenever storage appears, including after a driver restart.
This commit is contained in:
Daniel Samson
2026-07-21 20:27:55 +01:00
parent 7082699f5f
commit 1638845a4b
6 changed files with 92 additions and 29 deletions
+5 -2
View File
@@ -192,9 +192,12 @@ fn system_call(state: *architecture.CpuState) void {
.exit => {
exit_code = architecture.systemCallArg(state, 0);
// A scheduled process tears down fully (terminateCurrent); a borrowed
// test thread unwinds back to the kernel that entered it.
// test thread unwinds back to the kernel that entered it. A NONZERO
// code is a deliberate failure exit (`.aborted`): "the work exists
// but I could not do it" — supervisors restart those, unlike a clean
// `.exited` ("nothing for me here"), which they let lie.
if (scheduler.currentIsUserProcess()) {
scheduler.current().exit_reason = .exited;
scheduler.current().exit_reason = if (exit_code == 0) .exited else .aborted;
terminateCurrent();
} else architecture.userExit();
},