Release a dead process's device claims (M17.1)
Every path out of a process (exit, fault, kill) now releases its device claims alongside its IRQ and MSI bindings, so a restarted driver can claim its hardware again — the cleanup half of process-lifecycle.md's iron rule 1. MSI vectors were already swept by irq.releaseOwner; claims were the gap. The claim-release test proves kill -> release -> re-claim, plus the broker release in isolation.
This commit is contained in:
@@ -29,7 +29,9 @@ only when its definition of green holds.
|
|||||||
- [x] **Phase 0** — baseline: docs committed, feat/usb merged to main, pushed;
|
- [x] **Phase 0** — baseline: docs committed, feat/usb merged to main, pushed;
|
||||||
`usb-xhci-libary.zig` renamed to `usb-xhci-library.zig`; existing QEMU
|
`usb-xhci-libary.zig` renamed to `usb-xhci-library.zig`; existing QEMU
|
||||||
suite green from the worktree (48/48, 2026-07-12).
|
suite green from the worktree (48/48, 2026-07-12).
|
||||||
- [ ] **M17.1** — kernel releases claims/MSI on death (branch `feat/process-lifecycle`)
|
- [x] **M17.1** — kernel releases claims/MSI on death (claims: `releaseAllOwnedBy`
|
||||||
|
in the reap; MSI was already swept by `irq.releaseOwner`; `claim-release`
|
||||||
|
test; suite 49/49)
|
||||||
- [ ] **M17.2** — exit reasons
|
- [ ] **M17.2** — exit reasons
|
||||||
- [ ] **M17.3** — published exit events + VFS subscriber
|
- [ ] **M17.3** — published exit events + VFS subscriber
|
||||||
- [ ] **M17.4** — signals, timer notifications, `runtime.process`, the service harness
|
- [ ] **M17.4** — signals, timer notifications, `runtime.process`, the service harness
|
||||||
|
|||||||
@@ -95,8 +95,11 @@ the architecture layer calls up into `tick`.
|
|||||||
|
|
||||||
## Known gaps (bring-up honesty)
|
## Known gaps (bring-up honesty)
|
||||||
|
|
||||||
- Device **claims** are not released on death (pre-existing: the fault path has
|
- ~~Device claims are not released on death~~ Closed (M17.1): every path out of a
|
||||||
the same gap) — a killed driver's device stays claimed until reboot.
|
process releases its device claims alongside its IRQ and MSI bindings
|
||||||
|
(`releaseTaskResourcesLocked`), so a restarted driver can claim its hardware
|
||||||
|
again — the cleanup half of [process-lifecycle.md](process-lifecycle.md)'s iron
|
||||||
|
rule 1. The `claim-release` test proves the kill → release → re-claim cycle.
|
||||||
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
|
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
|
||||||
- There is no exit *status* in the notification, only the id; a supervisor that
|
- There is no exit *status* in the notification, only the id; a supervisor that
|
||||||
needs the code can grow a wait-style call later.
|
needs the code can grow a wait-style call later.
|
||||||
@@ -109,4 +112,5 @@ the architecture layer calls up into `tick`.
|
|||||||
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
|
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
|
||||||
notifications), `supervision` (the whole user-side surface via the process-test
|
notifications), `supervision` (the whole user-side surface via the process-test
|
||||||
service: spawn supervised → enumerate → kill blocked and spinning children →
|
service: spawn supervised → enumerate → kill blocked and spinning children →
|
||||||
notifications → gone). See test/qemu_test.py.
|
notifications → gone), `claim-release` (a killed claim-holder's device is
|
||||||
|
claimable again). See test/qemu_test.py.
|
||||||
|
|||||||
@@ -104,6 +104,19 @@ pub fn ownerOf(id: u64) ?u32 {
|
|||||||
return claimed[@intCast(id)];
|
return claimed[@intCast(id)];
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Release every claim held by `owner` — called by the process layer on every
|
||||||
|
/// path out of a process (exit, fault, kill), so a restarted driver can claim its
|
||||||
|
/// hardware again (docs/process-lifecycle.md iron rule 1: cleanup is the kernel's
|
||||||
|
/// job). The devices stay in the table — they describe hardware, which did not go
|
||||||
|
/// away — only their ownership clears.
|
||||||
|
pub fn releaseAllOwnedBy(owner: u32) void {
|
||||||
|
for (claimed[0..count]) |*slot| {
|
||||||
|
if (slot.*) |o| {
|
||||||
|
if (o == owner) slot.* = null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Resource `index` of device `id`, or null if out of range.
|
/// Resource `index` of device `id`, or null if out of range.
|
||||||
pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
|
pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
|
||||||
if (id >= count) return null;
|
if (id >= count) return null;
|
||||||
|
|||||||
@@ -552,7 +552,11 @@ pub var fault_kill_count: u64 = 0;
|
|||||||
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
|
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
|
||||||
/// ISR call notifyFromIsr on freed memory the next time the device fired.
|
/// ISR call notifyFromIsr on freed memory the next time the device fired.
|
||||||
/// `releaseOwner` also leaves the line masked, so a dead driver's device goes
|
/// `releaseOwner` also leaves the line masked, so a dead driver's device goes
|
||||||
/// quiet rather than storming.
|
/// quiet rather than storming. (It drops MSI vectors by the same owner sweep.)
|
||||||
|
/// - Device claims are released with the IRQ bindings, so a restarted driver can
|
||||||
|
/// claim the same hardware again — the cleanup half of process-lifecycle.md's
|
||||||
|
/// iron rule 1. Claims hold no pointers, so ordering is free; they go here so
|
||||||
|
/// the exit notification (below, last) observes a fully-released child.
|
||||||
/// - A client this task still owes a reply to (it died between receive and reply)
|
/// - A client this task still owes a reply to (it died between receive and reply)
|
||||||
/// is failed with -EPEER rather than left blocked forever — a dead server must
|
/// is failed with -EPEER rather than left blocked forever — a dead server must
|
||||||
/// not hang its callers.
|
/// not hang its callers.
|
||||||
@@ -566,6 +570,7 @@ pub var fault_kill_count: u64 = 0;
|
|||||||
/// Precondition: the big kernel lock is held.
|
/// Precondition: the big kernel lock is held.
|
||||||
fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
|
fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
|
||||||
irq.releaseOwner(t.id);
|
irq.releaseOwner(t.id);
|
||||||
|
devices_broker.releaseAllOwnedBy(t.id);
|
||||||
if (t.ipc_client) |client| {
|
if (t.ipc_client) |client| {
|
||||||
t.ipc_client = null;
|
t.ipc_client = null;
|
||||||
client.ipc_status = -ipc.EPEER;
|
client.ipc_status = -ipc.EPEER;
|
||||||
|
|||||||
@@ -132,6 +132,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
|||||||
processKillTest(boot_information);
|
processKillTest(boot_information);
|
||||||
} else if (eql(case, "supervision")) {
|
} else if (eql(case, "supervision")) {
|
||||||
supervisionTest(boot_information);
|
supervisionTest(boot_information);
|
||||||
|
} else if (eql(case, "claim-release")) {
|
||||||
|
claimReleaseTest(boot_information);
|
||||||
} else if (eql(case, "initial-ramdisk")) {
|
} else if (eql(case, "initial-ramdisk")) {
|
||||||
initialRamdiskTest(boot_information);
|
initialRamdiskTest(boot_information);
|
||||||
} else if (eql(case, "vfs")) {
|
} else if (eql(case, "vfs")) {
|
||||||
@@ -1454,6 +1456,61 @@ fn processKillTest(boot_information: *const BootInformation) void {
|
|||||||
result();
|
result();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// M17.1: a dead process's device claims are released by the reap, so a restarted
|
||||||
|
/// driver can claim its hardware again (docs/process-lifecycle.md iron rule 1).
|
||||||
|
/// First the broker release in isolation — two owners, one released, the other's
|
||||||
|
/// claim must survive. Then the death-path wiring with a real child: the claim is
|
||||||
|
/// made on the child's behalf (the broker is kernel-callable), the child is
|
||||||
|
/// killed, and once the exit notification arrives — posted last, after release —
|
||||||
|
/// the device must be unclaimed and claimable again.
|
||||||
|
fn claimReleaseTest(boot_information: *const BootInformation) void {
|
||||||
|
log("DANOS-TEST-BEGIN: claim-release\n", .{});
|
||||||
|
|
||||||
|
var buffer: [2]device_abi.DeviceDescriptor = undefined;
|
||||||
|
const total = devices_broker.enumerate(&buffer);
|
||||||
|
check("the device tree is seeded (>= 2 devices)", total >= 2);
|
||||||
|
if (total < 2) {
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
// The broker release in isolation.
|
||||||
|
check("device 0 claimed by owner 111", devices_broker.claim(0, 111));
|
||||||
|
check("device 1 claimed by owner 222", devices_broker.claim(1, 222));
|
||||||
|
devices_broker.releaseAllOwnedBy(111);
|
||||||
|
check("owner 111's claim is released", devices_broker.ownerOf(0) == null);
|
||||||
|
check("owner 222's claim survives", (devices_broker.ownerOf(1) orelse 0) == 222);
|
||||||
|
devices_broker.releaseAllOwnedBy(222);
|
||||||
|
check("cleanup released owner 222", devices_broker.ownerOf(1) == null);
|
||||||
|
|
||||||
|
// The death-path wiring: a real process dies holding a claim.
|
||||||
|
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
|
||||||
|
if (boot_information.init_len == 0) {
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
|
||||||
|
const me = scheduler.currentId();
|
||||||
|
const endpoint = ipcsync.createIpcEndpoint() orelse {
|
||||||
|
check("exit endpoint allocated", false);
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
const child = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
|
||||||
|
check("supervised child spawned", child != 0);
|
||||||
|
check("device 0 claimed on the child's behalf", devices_broker.claim(0, child));
|
||||||
|
|
||||||
|
check("the kill is accepted", process.killProcess(me, child) == 0);
|
||||||
|
var badge: u64 = 0;
|
||||||
|
var received_cap: u64 = 0;
|
||||||
|
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
|
||||||
|
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | child);
|
||||||
|
check("death released the child's claim", devices_broker.ownerOf(0) == null);
|
||||||
|
check("the device is claimable again", devices_broker.claim(0, me));
|
||||||
|
devices_broker.releaseAllOwnedBy(me);
|
||||||
|
result();
|
||||||
|
}
|
||||||
|
|
||||||
/// The whole user-side surface at once: spawn process-test's supervisor role,
|
/// The whole user-side surface at once: spawn process-test's supervisor role,
|
||||||
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two
|
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two
|
||||||
/// children supervised, sees them in process_enumerate, kills them (one blocked,
|
/// children supervised, sees them in process_enumerate, kills them (one blocked,
|
||||||
|
|||||||
@@ -240,6 +240,12 @@ CASES = [
|
|||||||
"smp": 4,
|
"smp": 4,
|
||||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
|
# M17.1: a dead process's device claims are released by the reap — kill a child
|
||||||
|
# holding a claim, the device must be claimable again (process-lifecycle.md).
|
||||||
|
{"name": "claim-release",
|
||||||
|
"smp": 4,
|
||||||
|
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||||
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
|
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
|
||||||
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
|
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
|
||||||
{"name": "initial-ramdisk",
|
{"name": "initial-ramdisk",
|
||||||
|
|||||||
Reference in New Issue
Block a user