kernel+tests: M4 shared-fate — nine group-death test cases, per-space arena cursors
Eight new QEMU cases (thread-fault-group, kill-threaded-group, kill-via-worker-tid, racing-triggers, exit-group, leader-thread-exit, thread-exit-solo, shm-mapping-ref) driving seven new thread-test modes; a shared checkGroupDead asserts the contract everywhere: one notification, badged with the leader, reason on the leader's record, no member listed, claims released first. The shm-mapping-ref case flushed out the per-task DMA/shared-memory arena cursor bug directly (a sibling's regions mapped over the worker's), so both cursors moved to the AddressSpaceRef like the mmap/MMIO cursors before them (threading M7 pattern). Runtime gains Thread.tryExitCurrent for the leader -EPERM refusal path. Docs updated: threading.md's shared-fate gap is closed, process-management.md and process-lifecycle.md describe the leader re-key, plan status = implemented. Full suite: 100/100.
This commit is contained in:
@@ -55,6 +55,9 @@ pattern reused. Signals are the same pattern reused a third time.
|
||||
- **`process_signal(id, signal)`** — posts the signal as an asynchronous
|
||||
notification to the target's bound endpoint: badge = `notify_badge_bit |
|
||||
notify_signal_bit | pending signals`. Non-blocking for the sender, always.
|
||||
Signals address the *process*: `id` may name any member of a threaded process
|
||||
and resolves to its leader — whose endpoint the harness binds — with authority
|
||||
mirroring `process_kill` ([shared-fate-plan.md](shared-fate-plan.md)).
|
||||
- **Pending signals coalesce** in a per-process bitmask while the target has no
|
||||
signal endpoint bound, and the whole mask arrives as one notification at bind —
|
||||
POSIX's own semantics for non-realtime signals (two pending SIGTERMs are one
|
||||
|
||||
@@ -57,9 +57,13 @@ dangle even if the supervisor dies first.
|
||||
|
||||
### `process_kill(id) -> 0 / -ESRCH / -EPERM`
|
||||
|
||||
Only the supervisor may kill; kernel tasks are not killable processes. Like a
|
||||
signal, delivery is prompt but asynchronous — 0 means the kill is accepted and
|
||||
irrevocable; the exit notification confirms completion.
|
||||
Only the supervisor may kill; kernel tasks are not killable processes. The kill
|
||||
is a **whole-process** kill ([shared-fate-plan.md](shared-fate-plan.md)): `id`
|
||||
may name any member of a threaded process — it resolves to the group's leader,
|
||||
authorization is checked against the *leader's* supervisor, and every thread
|
||||
dies. Like a signal, delivery is prompt but asynchronous — 0 means the kill is
|
||||
accepted and irrevocable; the exit notification (badged with the leader, posted
|
||||
once the last member is gone) confirms completion.
|
||||
|
||||
## How a kill lands (the kernel mechanics)
|
||||
|
||||
|
||||
@@ -1,7 +1,11 @@
|
||||
# Shared fate: whole-process death (plan)
|
||||
|
||||
**Status: approved 2026-07-22; leader `thread_exit` → `-EPERM`. Implementation in
|
||||
progress (M1–M4 below).**
|
||||
**Status: implemented 2026-07-22 (branch shared-fate), M1–M4 all landed; leader
|
||||
`thread_exit` → `-EPERM` as decided. One scope addition forced by M4: the
|
||||
per-task DMA/shared-memory arena cursors moved to the per-space object (the
|
||||
`shm-mapping-ref` test could not distinguish corruption-by-remap from
|
||||
corruption-by-free while sibling threads overlapped the arena) — the same move
|
||||
the mmap/MMIO cursors made in threading M7.**
|
||||
|
||||
[threading.md](threading.md) promises that a process dies *whole* — a fault in any
|
||||
thread, or a kill, takes down every thread. The kernel doesn't do that yet: every
|
||||
@@ -244,10 +248,10 @@ refcount, and no group-kill special case is needed at all.
|
||||
`.ready` with `in_system_call = true`, and the tick's reap loop will reap it —
|
||||
an existing hazard the fan-out inherits but must not add new instances of.
|
||||
Filed to investigate separately.
|
||||
- **Per-task DMA/shm cursors** (`dma_map_next`, `shared_memory_map_next`): two
|
||||
sibling threads allocating overlap the same arena — pre-existing thread bug,
|
||||
adjacent to but not part of this plan (the mmap/MMIO cursors already moved
|
||||
per-space for exactly this reason).
|
||||
- **Per-task DMA/shm cursors** — *fixed during M4 after all*: the
|
||||
`shm-mapping-ref` test tripped the overlap (the sibling's churn regions mapped
|
||||
over the worker's region), so both cursors moved to the `AddressSpaceRef`
|
||||
like the mmap/MMIO cursors before them.
|
||||
|
||||
## Milestones
|
||||
|
||||
|
||||
+9
-11
@@ -252,20 +252,18 @@ stays single-threaded and lean.
|
||||
halts" property intact under lock contention — no busy-wait.
|
||||
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): the contract is that
|
||||
killing a process kills *all* its threads and only then drops the last address-space
|
||||
ref. **The kernel does not implement that fan-out yet**: `process_kill` reaps only
|
||||
the one task it resolves, and no death path loops over the tasks sharing an address
|
||||
space — the refcount keeps the space (and the sibling threads) alive and running.
|
||||
The gap is hit in practice: of the only threaded binaries (the `display` service and
|
||||
the `thread-test` harness), `display` is a boot service that handles no `.terminate`
|
||||
signal, so init's stop sequence escalates to `process_kill` on every orderly
|
||||
shutdown — benign only because poweroff follows. Whether to implement the fan-out or
|
||||
amend the contract is a decision still to be made.
|
||||
ref — and the kernel now implements exactly that
|
||||
([shared-fate-plan.md](shared-fate-plan.md)): every death path (`exit` from any
|
||||
thread, a fault, `process_kill` aimed at any member id) fans out through the whole
|
||||
group via a `dying` latch on the address space; the supervisor's one exit
|
||||
notification — badged with the leader — fires only when the last member is gone.
|
||||
A worker's voluntary `thread_exit` stays per-thread; the leader's is refused
|
||||
(`-EPERM`).
|
||||
- **Resilience** ([resilience.md](resilience.md)): by the same contract, a faulting
|
||||
thread kills its whole process (shared fate); the supervisor restarts the
|
||||
**process**, which respawns its threads from a known-good state — restart
|
||||
granularity stays the process. Today a CPU fault kills only the faulting task
|
||||
(`killCurrentProcess` tears down a single task), so sibling threads keep running —
|
||||
the same implementation gap as above.
|
||||
granularity stays the process. The leader's recorded exit reason carries the fault
|
||||
class even when a worker faulted, so restart policy is unchanged.
|
||||
- **IPC — two consequences threads forced ([ipc.md](ipc.md)):**
|
||||
- *Handles do not cross threads.* The handle table lives on the `Task`
|
||||
([scheduler.zig](../system/kernel/scheduler.zig)), so a handle number is meaningful
|
||||
|
||||
Reference in New Issue
Block a user