kernel: the device rides system_spawn, so a driver never runs without it

Delegation moves out of onHello and into the spawn itself. The manager holds
the hardware and names it in the call that creates the driver; the kernel
checks the device is the caller's to give, then hands it over as part of
making the child.

The reason is the window. A transfer after spawning always leaves an
interval in which the child is running and does not yet hold its device. It
would close on QEMU every time and open occasionally on a machine with
different core counts and timing — the exact failure shape this track exists
to delete, and not one worth introducing while removing the others. Fused
into the spawn there is no interval: the child does not exist until it holds
the device.

Ownership is checked BEFORE the child is created, so a refusal leaves
nothing running rather than a driver without the hardware it was spawned
for. The IOMMU confinement moves with the device, as it does on the transfer
path. systemCall6 is added for the sixth argument; r9 was free, and abi
gains a no_device sentinel matching the protocol's.

No driver had to change to receive a device, which is what makes this
better than requiring every driver to hello: ps2-bus keeps its legacy
status, and discovery — which has no assignment at all, since it is what
produces the device tree — is unaffected.

The attacker fixture now tries the spawn as a back door: name someone else's
device, and both the spawn and any child must be refused. Verifying that
assertion exposed a bug in the fixture itself. The kernel case's pass marker
was "device-authority: ok", which matches the FIRST per-assertion line, so
its wait loop exited before any failure was printed — the case would have
passed with failures in it, and had been able to since D2. The verdict lines
now carry a distinct VERDICT prefix, and with the ownership check removed
the case genuinely fails. A green test that cannot go red is worse than no
test.

Suite 118/118.
This commit is contained in:
Daniel Samson
2026-08-08 19:39:32 +01:00
parent 637bf2e0b1
commit f23f073624
8 changed files with 95 additions and 24 deletions
+26
View File
@@ -1017,6 +1017,13 @@ fn systemSpawn(state: *architecture.CpuState) void {
const arguments_ptr = architecture.systemCallArg(state, 2);
const arguments_len = architecture.systemCallArg(state, 3);
const exit_handle = architecture.systemCallArg(state, 4);
// A device the caller holds and gives to the child. Fused into the spawn rather
// than transferred after it, because a separate transfer leaves a window in which
// the child is running and does not yet hold its device — a race that would close
// on one machine and open on another, which is the failure shape this whole track
// exists to remove (docs/bounds-track-plan.md, "the grant rides system_spawn").
// Here the child cannot observe the gap: it does not exist until it holds it.
const device_to_give = architecture.systemCallArg(state, 5);
const t = scheduler.current();
if (len == 0 or len > scheduler.maximum_task_name or ptr >= user_half_end or ptr + len > user_half_end) return fail(state);
if (arguments_len > maximum_argument_bytes) return fail(state);
@@ -1061,7 +1068,26 @@ fn systemSpawn(state: *architecture.CpuState) void {
}
}
// Refuse before creating anything if the device is not the caller's to give — a
// spawn that half-succeeds would leave a child running without the hardware it was
// spawned for, which is worse than not spawning it.
if (device_to_give != abi.no_device and devices_broker.ownerOf(device_to_give) != t.id)
return failErr(state, ipc.EPERM);
const child = spawnProcessSupervised(item.blob, 4, argv[0..argc], t.id, exit_endpoint) catch return fail(state);
if (device_to_give != abi.no_device) {
devices_broker.transfer(device_to_give, t.id, child) catch {
// Cannot happen — ownership was checked above and the lock has not been
// dropped — but a spawned child holding nothing is not something to guess
// about, so say so rather than leave it silent.
log.print("/system/kernel: WARNING spawn gave device {d} to task {d} and the transfer failed\n", .{ device_to_give, child });
};
if (devices_broker.pciAddressOf(device_to_give)) |_| {
iommu.reassign(device_to_give, child);
dmaBindOwnerRegionsInto(child, device_to_give);
}
}
architecture.setSystemCallResult(state, child);
}
+6 -2
View File
@@ -4257,8 +4257,12 @@ fn deviceAuthorityTest(boot_information: *const BootInformation) void {
process.setInitialRamdisk(image);
check("device-authority-test spawned", spawnNamedWithArg(rd, "device-authority-test", "run"));
const pass_marker = "device-authority: ok";
const fail_marker = "device-authority: FAIL";
// The VERDICT prefix matters: the fixture prints one "device-authority: ok <name>"
// line per assertion, so a marker of "device-authority: ok" matches the FIRST
// passing assertion and this loop exits before any later failure is printed — the
// case then passes with failures in it, which it did until this was caught.
const pass_marker = "device-authority: VERDICT ok";
const fail_marker = "device-authority: VERDICT FAILED";
scheduler.setPriority(1);
const deadline = architecture.millis() + 20000;
var saw_pass = false;