docs+code: spell out aspace/vaddr/paddr per coding standards
Expand the abbreviations flagged in docs/coding-standards.md (names spelled
out in full unless an acronym) across the kernel, runtime, ABI, tests, and
docs:
aspace -> address_space (AspaceRef -> AddressSpaceRef, retainAspace ->
retainAddressSpace, loaded_aspace -> loaded_address_space, the
liveAspaceCount/aspaceDestroyCount test hooks, etc.)
vaddr -> virtual_address
paddr -> physical_address
The kernel test case and its serial markers are renamed to match:
aspace-refcount -> address-space-refcount (kernel dispatch string and
test/qemu_test.py case name kept in sync). Prose in docs uses the natural
"address space"/"virtual address"; backticked field/identifier references
use the code spelling.
Also expand the bare "AS" abbreviation in three ABI comments and reframe the
set_thread_pointer ABI/handler docs to lead with the arch-neutral concept
(user-space TLS thread pointer; x86_64 IA32_FS_BASE, aarch64 TPIDR_EL0)
rather than x86 FS-first, matching scheduler.zig's existing framing.
Foreign ABI names preserved: the ELF p_vaddr field and mmap/mmio remain.
Verified: zig build, zig build test, and the full 25-case QEMU guardrail
suite all green.
This commit is contained in:
+1
-1
@@ -107,7 +107,7 @@ Start with the north star:
|
|||||||
- **[threading.md](threading.md) — threads, the std-shaped way.** **Built** (M1–M6):
|
- **[threading.md](threading.md) — threads, the std-shaped way.** **Built** (M1–M6):
|
||||||
`runtime.Thread` mirrors `std.Thread`'s API (spawn/join/detach, Mutex/Condition/
|
`runtime.Thread` mirrors `std.Thread`'s API (spawn/join/detach, Mutex/Condition/
|
||||||
Semaphore) over a **private** thread ABI — several tasks sharing one address space via
|
Semaphore) over a **private** thread ABI — several tasks sharing one address space via
|
||||||
a `thread_spawn` syscall, futex-backed blocking, aspace refcounting. Why it's the
|
a `thread_spawn` syscall, futex-backed blocking, address-space refcounting. Why it's the
|
||||||
native type and not literal `std.Thread` (the [private ABI](syscall.md)), and why
|
native type and not literal `std.Thread` (the [private ABI](syscall.md)), and why
|
||||||
threads stay a narrow opt-in against the [resilience](resilience.md) default. Build
|
threads stay a narrow opt-in against the [resilience](resilience.md) default. Build
|
||||||
plan + gates: [threading-plan.md](threading-plan.md).
|
plan + gates: [threading-plan.md](threading-plan.md).
|
||||||
|
|||||||
@@ -62,7 +62,7 @@ is the only backend), and `zig build test` stays green.
|
|||||||
- [x] [abi.zig](../system/abi.zig): `shm_create` (34) / `shm_map` (35) syscalls + a
|
- [x] [abi.zig](../system/abi.zig): `shm_create` (34) / `shm_map` (35) syscalls + a
|
||||||
`shm_test` service id. Handlers in process.zig: `shm_create(len)` allocates contiguous,
|
`shm_test` service id. Handlers in process.zig: `shm_create(len)` allocates contiguous,
|
||||||
zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability
|
zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability
|
||||||
handle, maps them into the caller's shm arena → returns vaddr + handle; `shm_map(cap)`
|
handle, maps them into the caller's shm arena → returns virtual_address + handle; `shm_map(cap)`
|
||||||
maps the same physical pages into the receiver. Reclaimed on death (see below).
|
maps the same physical pages into the receiver. Reclaimed on death (see below).
|
||||||
- [x] The capability core (ipc-synchronous.zig) is now **kind-tagged**: `scheduler.Task`'s
|
- [x] The capability core (ipc-synchronous.zig) is now **kind-tagged**: `scheduler.Task`'s
|
||||||
handle table holds `HandleObject{kind, ptr}`; `closeHandles` and `shareCapability`
|
handle table holds `HandleObject{kind, ptr}`; `closeHandles` and `shareCapability`
|
||||||
|
|||||||
+2
-2
@@ -77,9 +77,9 @@ deferred (docs/display.md, "What v1 does not do"). v2 builds it: the natural gen
|
|||||||
of M13 capability-passing from *endpoints* to *memory objects* —
|
of M13 capability-passing from *endpoints* to *memory objects* —
|
||||||
|
|
||||||
```
|
```
|
||||||
shm_create(len) -> {handle, vaddr} // a shareable, page-aligned RAM region
|
shm_create(len) -> {handle, virtual_address} // a shareable, page-aligned RAM region
|
||||||
… pass `handle` as the send_cap on an ipc_call …
|
… pass `handle` as the send_cap on an ipc_call …
|
||||||
shm_map(cap) -> vaddr // the receiver maps the same physical pages
|
shm_map(cap) -> virtual_address // the receiver maps the same physical pages
|
||||||
```
|
```
|
||||||
|
|
||||||
The payoff is leverage: the **same** primitive unlocks **both** native GPU drivers *and*
|
The payoff is leverage: the **same** primitive unlocks **both** native GPU drivers *and*
|
||||||
|
|||||||
+2
-2
@@ -220,8 +220,8 @@ both are clean additions behind the interfaces v1 establishes.
|
|||||||
to render into its *own* buffer and hand the compositor a *reference*, not a stream of
|
to render into its *own* buffer and hand the compositor a *reference*, not a stream of
|
||||||
commands. That needs the missing cross-process shared-memory primitive — best built as
|
commands. That needs the missing cross-process shared-memory primitive — best built as
|
||||||
the natural generalization of the existing M13 [capability passing](driver-model.md)
|
the natural generalization of the existing M13 [capability passing](driver-model.md)
|
||||||
from *endpoints* to *memory objects* (`shm_create(len) → {cap, vaddr}`, pass `cap` on
|
from *endpoints* to *memory objects* (`shm_create(len) → {cap, virtual_address}`, pass `cap` on
|
||||||
an `ipc_call`, receiver `shm_map(cap) → vaddr`). v1 avoids it because server-owned
|
an `ipc_call`, receiver `shm_map(cap) → virtual_address`). v1 avoids it because server-owned
|
||||||
surfaces already prove the whole pipeline.
|
surfaces already prove the whole pipeline.
|
||||||
|
|
||||||
- **Runtime mode-setting (a native backend).** Detecting the EDID mode list and changing
|
- **Runtime mode-setting (a native backend).** Detecting the EDID mode list and changing
|
||||||
|
|||||||
@@ -240,8 +240,8 @@ once per page, maps writeback-cached, and never reveals a physical address.
|
|||||||
**The fix.**
|
**The fix.**
|
||||||
|
|
||||||
```
|
```
|
||||||
dma_alloc(len, flags) -> vaddr (rax), paddr (rdx)
|
dma_alloc(len, flags) -> virtual_address (rax), physical_address (rdx)
|
||||||
dma_free(vaddr, len) -> 0
|
dma_free(virtual_address, len) -> 0
|
||||||
|
|
||||||
flags: dma_coherent (1) uncacheable; the default and the only one that's portable
|
flags: dma_coherent (1) uncacheable; the default and the only one that's portable
|
||||||
dma_wc (2) write-combining — needs PAT programmed; for framebuffers
|
dma_wc (2) write-combining — needs PAT programmed; for framebuffers
|
||||||
|
|||||||
+1
-1
@@ -74,7 +74,7 @@ The driver syscall numbers (`system/abi.zig`) with the device types they carry
|
|||||||
|---|------|---------|
|
|---|------|---------|
|
||||||
| 11 | `device_enumerate(buf, max) -> total` | Snapshot the device table |
|
| 11 | `device_enumerate(buf, max) -> total` | Snapshot the device table |
|
||||||
| 12 | `device_claim(id) -> ok` | Take **exclusive** ownership |
|
| 12 | `device_claim(id) -> ok` | Take **exclusive** ownership |
|
||||||
| 13 | `mmio_map(id, res_idx) -> vaddr` | Map a claimed device's register window |
|
| 13 | `mmio_map(id, res_idx) -> virtual_address` | Map a claimed device's register window |
|
||||||
| 14 | `irq_bind(id, res_idx, endpoint)` | Deliver that device's IRQ as a notification |
|
| 14 | `irq_bind(id, res_idx, endpoint)` | Deliver that device's IRQ as a notification |
|
||||||
| 15 | `irq_ack(id, res_idx)` | Re-arm the IRQ after servicing the device |
|
| 15 | `irq_ack(id, res_idx)` | Re-arm the IRQ after servicing the device |
|
||||||
| 16 | `device_register(parent_id, desc) -> id` | Publish a child of a device you claimed |
|
| 16 | `device_register(parent_id, desc) -> id` | Publish a child of a device you claimed |
|
||||||
|
|||||||
+21
-21
@@ -105,28 +105,28 @@ first unchecked box.
|
|||||||
## M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅
|
## M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅
|
||||||
|
|
||||||
The one invariant change threads require, landed and proven **before** anything shares
|
The one invariant change threads require, landed and proven **before** anything shares
|
||||||
an address space. Today aspace is 1:1 with a task and teardown destroys it on any user
|
an address space. Today address space is 1:1 with a task and teardown destroys it on any user
|
||||||
task's exit; make destruction happen on the **last** exit.
|
task's exit; make destruction happen on the **last** exit.
|
||||||
|
|
||||||
- [x] A refcount keyed by the address-space root, held in `scheduler.zig`
|
- [x] A refcount keyed by the address-space root, held in `scheduler.zig`
|
||||||
(`aspace_refs`): `retainAspace` takes a reference in `spawnUserLocked` (on the
|
(`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the
|
||||||
success path, after the slot + stack are secured), all under the big kernel lock.
|
success path, after the slot + stack are secured), all under the big kernel lock.
|
||||||
- [x] Both task-teardown paths ([scheduler.zig](../system/kernel/scheduler.zig):
|
- [x] Both task-teardown paths ([scheduler.zig](../system/kernel/scheduler.zig):
|
||||||
`exitUserLocked` and `destroyTaskLocked`) call `releaseAspace`, which decrements
|
`exitUserLocked` and `destroyTaskLocked`) call `releaseAspace`, which decrements
|
||||||
and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test
|
and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test
|
||||||
spaces) is destroyed directly, preserving prior behaviour.
|
spaces) is destroyed directly, preserving prior behaviour.
|
||||||
- [x] `-Dtest-case=aspace-refcount`: spawn and reap several ring-3 processes in sequence
|
- [x] `-Dtest-case=address-space-refcount`: spawn and reap several ring-3 processes in sequence
|
||||||
and assert (via test-observable `liveAspaceCount`/`aspaceDestroyCount`) that the
|
and assert (via test-observable `liveAddressSpaceCount`/`addressSpaceDestroyCount`) that the
|
||||||
live-space count returns to **baseline** and destructions advance by exactly that
|
live-space count returns to **baseline** and destructions advance by exactly that
|
||||||
many — each space destroyed exactly once, no leak, no double-free. (Refcount
|
many — each space destroyed exactly once, no leak, no double-free. (Refcount
|
||||||
observables, not raw frame counts, since kernel stacks are still leaked on exit.)
|
observables, not raw frame counts, since kernel stacks are still leaked on exit.)
|
||||||
|
|
||||||
**Gate (met):** `python3 test/qemu_test.py aspace-refcount` passes
|
**Gate (met):** `python3 test/qemu_test.py address-space-refcount` passes
|
||||||
(`aspace-refcount: spaces released to baseline ok` → `DANOS-TEST-RESULT: PASS`), and the
|
(`address-space-refcount: spaces released to baseline ok` → `DANOS-TEST-RESULT: PASS`), and the
|
||||||
full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `smp`,
|
full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `smp`,
|
||||||
`affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`,
|
`affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`,
|
||||||
`vfs-client-death`, `ipc`, `ipc-cap`, `display-service`); default `zig build` clean,
|
`vfs-client-death`, `ipc`, `ipc-cap`, `display-service`); default `zig build` clean,
|
||||||
`zig build test` green. The reframing is invisible until an aspace is actually shared.
|
`zig build test` green. The reframing is invisible until an address space is actually shared.
|
||||||
|
|
||||||
## M2 — `thread_spawn` + `thread_exit`: a thread runs in the shared address space ✅
|
## M2 — `thread_spawn` + `thread_exit`: a thread runs in the shared address space ✅
|
||||||
|
|
||||||
@@ -135,7 +135,7 @@ space and exits cleanly.
|
|||||||
|
|
||||||
- [x] [abi.zig](../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in
|
- [x] [abi.zig](../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in
|
||||||
process.zig; `thread_spawn` calls `scheduler.spawnThread` (shares the caller's
|
process.zig; `thread_spawn` calls `scheduler.spawnThread` (shares the caller's
|
||||||
aspace, `retainAspace`); `thread_exit` ends the task like a process `exit(0)`
|
address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)`
|
||||||
(`terminateCurrent` → `releaseAspace`). The closure pointer is delivered in the new
|
(`terminateCurrent` → `releaseAspace`). The closure pointer is delivered in the new
|
||||||
thread's **rdi** via a new `jump_to_user_arg` asm path (`t.user_arg`, 0 for a
|
thread's **rdi** via a new `jump_to_user_arg` asm path (`t.user_arg`, 0 for a
|
||||||
process) — no naked runtime asm.
|
process) — no naked runtime asm.
|
||||||
@@ -152,14 +152,14 @@ space and exits cleanly.
|
|||||||
address space.
|
address space.
|
||||||
|
|
||||||
**Gate (met):** `python3 test/qemu_test.py thread-spawn` passes
|
**Gate (met):** `python3 test/qemu_test.py thread-spawn` passes
|
||||||
(`thread-test: child ran in shared aspace ok` → `DANOS-TEST-RESULT: PASS`); guardrail set
|
(`thread-test: child ran in shared address space ok` → `DANOS-TEST-RESULT: PASS`); guardrail set
|
||||||
16/16 green (incl. `args`/`init`/`process`, which exercise the new `jump_to_user_arg`
|
16/16 green (incl. `args`/`init`/`process`, which exercise the new `jump_to_user_arg`
|
||||||
process path with arg 0) plus `aspace-refcount`; `zig build` clean, `zig build test`
|
process path with arg 0) plus `address-space-refcount`; `zig build` clean, `zig build test`
|
||||||
green.
|
green.
|
||||||
|
|
||||||
> **Note (deferred to M3+):** the mmap arena is per-*task* (`heap_next`), so two threads
|
> **Note (deferred to M3+):** the mmap arena is per-*task* (`heap_next`), so two threads
|
||||||
> in one aspace that both `mmap` would collide. Fine for M2 (only the parent maps, for the
|
> in one address space that both `mmap` would collide. Fine for M2 (only the parent maps, for the
|
||||||
> child's stack); make the arena per-aspace and the runtime heap thread-safe alongside the
|
> child's stack); make the arena per-address-space and the runtime heap thread-safe alongside the
|
||||||
> `Mutex` work (M5).
|
> `Mutex` work (M5).
|
||||||
|
|
||||||
## M3 — `join` + `detach` + real parallelism ✅
|
## M3 — `join` + `detach` + real parallelism ✅
|
||||||
@@ -184,7 +184,7 @@ green.
|
|||||||
**Gate (met):** `python3 test/qemu_test.py thread-join` passes (`thread-test: join ok` →
|
**Gate (met):** `python3 test/qemu_test.py thread-join` passes (`thread-test: join ok` →
|
||||||
`DANOS-TEST-RESULT: PASS`), robust across 4 runs; guardrail 17/17 green (incl. `smp`,
|
`DANOS-TEST-RESULT: PASS`), robust across 4 runs; guardrail 17/17 green (incl. `smp`,
|
||||||
`affinity`, `process-kill`, and `args`/`init`/`process` on the exit-endpoint spawn path)
|
`affinity`, `process-kill`, and `args`/`init`/`process` on the exit-endpoint spawn path)
|
||||||
plus `aspace-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green.
|
plus `address-space-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green.
|
||||||
|
|
||||||
> **Note (deferred):** a detached thread's stack is freed only at process exit (not by the
|
> **Note (deferred):** a detached thread's stack is freed only at process exit (not by the
|
||||||
> reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And
|
> reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And
|
||||||
@@ -212,7 +212,7 @@ plus `aspace-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green
|
|||||||
**Gate (met):** `python3 test/qemu_test.py thread-futex` passes, robust across 3 runs —
|
**Gate (met):** `python3 test/qemu_test.py thread-futex` passes, robust across 3 runs —
|
||||||
the case's **ordered** regex asserts `waiting → waking → woke → PASS` on the serial
|
the case's **ordered** regex asserts `waiting → waking → woke → PASS` on the serial
|
||||||
stream (the handoff proof), and `thread-futex: timeout ok` confirms the timeout.
|
stream (the handoff proof), and `thread-futex: timeout ok` confirms the timeout.
|
||||||
Guardrail 18/18 green (incl. `sleep`/`event`/`ipc` blocking paths) + `aspace-refcount`,
|
Guardrail 18/18 green (incl. `sleep`/`event`/`ipc` blocking paths) + `address-space-refcount`,
|
||||||
`thread-spawn`, `thread-join`; `zig build` clean, `zig build test` green.
|
`thread-spawn`, `thread-join`; `zig build` clean, `zig build test` green.
|
||||||
|
|
||||||
> **Note:** the kernel test checks only the freshest verdict marker via `bufferHas` (the
|
> **Note:** the kernel test checks only the freshest verdict marker via `bufferHas` (the
|
||||||
@@ -293,9 +293,9 @@ The organising principle, so Phase 2 reinforces danos's goals rather than erodin
|
|||||||
|
|
||||||
- **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks,
|
- **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks,
|
||||||
and futex words live in the process's **address space**, and the kernel's per-process
|
and futex words live in the process's **address space**, and the kernel's per-process
|
||||||
state is keyed by the aspace root — so the M1 refcount + `destroyAddressSpace` already
|
state is keyed by the address-space root — so the M1 refcount + `destroyAddressSpace` already
|
||||||
free all of it when the last thread exits. A crashed or killed threaded process leaves
|
free all of it when the last thread exits. A crashed or killed threaded process leaves
|
||||||
**nothing** behind. Phase 2 closes the one thing that is *not* aspace-owned — the
|
**nothing** behind. Phase 2 closes the one thing that is *not* address-space-owned — the
|
||||||
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
|
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
|
||||||
[resilience](resilience.md) restart guarantee, extended to threads.
|
[resilience](resilience.md) restart guarantee, extended to threads.
|
||||||
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
|
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
|
||||||
@@ -313,9 +313,9 @@ threads in one process that both allocate corrupt each other. The thread *machin
|
|||||||
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
|
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
|
||||||
multi-threaded code would hit it. Closed it:
|
multi-threaded code would hit it. Closed it:
|
||||||
|
|
||||||
- [x] **Kernel — per-address-space mmap arena.** Grew M1's `aspace_refs` entry into the
|
- [x] **Kernel — per-address-space mmap arena.** Grew M1's `address_space_refs` entry into the
|
||||||
per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`);
|
per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`);
|
||||||
`scheduler.aspaceMmapNextPtr`/`aspaceDeviceMapNextPtr` expose them. `systemMmap`
|
`scheduler.addressSpaceMmapNextPtr`/`addressSpaceDeviceMapNextPtr` expose them. `systemMmap`
|
||||||
reserves a disjoint range under a *brief* lock, then maps **per page** under a
|
reserves a disjoint range under a *brief* lock, then maps **per page** under a
|
||||||
short-held lock — not the whole grant — because the big lock is held with interrupts
|
short-held lock — not the whole grant — because the big lock is held with interrupts
|
||||||
disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed
|
disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed
|
||||||
@@ -359,7 +359,7 @@ death lost one, so a crash loop bled kernel memory. The reaper fixes it and serv
|
|||||||
stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered.
|
stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered.
|
||||||
|
|
||||||
**Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`,
|
**Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`,
|
||||||
`supervision`, `process-kill`, `aspace-refcount`, `smp`, `affinity` all still green (24/24
|
`supervision`, `process-kill`, `address-space-refcount`, `smp`, `affinity` all still green (24/24
|
||||||
full guardrail); `zig build`/`zig build test` clean.
|
full guardrail); `zig build`/`zig build test` clean.
|
||||||
|
|
||||||
> **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap
|
> **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap
|
||||||
@@ -397,7 +397,7 @@ guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-re
|
|||||||
|
|
||||||
> **Deferred:** detached-thread **user-stack** reclaim (still freed at process exit, as in
|
> **Deferred:** detached-thread **user-stack** reclaim (still freed at process exit, as in
|
||||||
> M3). Doing it in the reaper needs the saved address space + stack range and a
|
> M3). Doing it in the reaper needs the saved address space + stack range and a
|
||||||
> translate/unmap in a not-currently-loaded aspace — real complexity for a bounded leak.
|
> translate/unmap in a not-currently-loaded address space — real complexity for a bounded leak.
|
||||||
> A follow-up when a consumer needs it.
|
> A follow-up when a consumer needs it.
|
||||||
|
|
||||||
### M10 — Per-thread TLS: the `fs.base` mechanism ✅
|
### M10 — Per-thread TLS: the `fs.base` mechanism ✅
|
||||||
@@ -456,7 +456,7 @@ clean.
|
|||||||
|
|
||||||
## Deferred (explicitly not in this plan)
|
## Deferred (explicitly not in this plan)
|
||||||
|
|
||||||
- **Cross-process shared-memory futex** — the `(aspace, vaddr)` key can become a
|
- **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a
|
||||||
physical-address key so two processes share a futex through an [shm](display-v2.md)
|
physical-address key so two processes share a futex through an [shm](display-v2.md)
|
||||||
region. Not needed for intra-process threads.
|
region. Not needed for intra-process threads.
|
||||||
- **Per-thread priorities / affinity distinct from the process** — threads inherit the
|
- **Per-thread priorities / affinity distinct from the process** — threads inherit the
|
||||||
|
|||||||
+16
-16
@@ -148,10 +148,10 @@ Plus one invariant change with no new syscall: **address-space reference countin
|
|||||||
|
|
||||||
### Address-space reference counting
|
### Address-space reference counting
|
||||||
|
|
||||||
Today an address space is 1:1 with a task: `spawnUserLocked` records `aspace` on the
|
Today an address space is 1:1 with a task: `spawnUserLocked` records `address_space` on the
|
||||||
Task, and teardown does `destroyAddressSpace(t.aspace)` when **any** user task exits
|
Task, and teardown does `destroyAddressSpace(t.address_space)` when **any** user task exits
|
||||||
([scheduler.zig](../system/kernel/scheduler.zig)). With threads, several tasks share
|
([scheduler.zig](../system/kernel/scheduler.zig)). With threads, several tasks share
|
||||||
one `aspace`, so the first to exit would rip the address space out from under its
|
one `address_space`, so the first to exit would rip the address space out from under its
|
||||||
siblings.
|
siblings.
|
||||||
|
|
||||||
Fix: a small refcount keyed by the address-space root (`createAddressSpace` in
|
Fix: a small refcount keyed by the address-space root (`createAddressSpace` in
|
||||||
@@ -162,7 +162,7 @@ that must land and be proven before anything shares an address space.
|
|||||||
|
|
||||||
### `thread_spawn` and the trampoline
|
### `thread_spawn` and the trampoline
|
||||||
|
|
||||||
The scheduler already accepts an arbitrary `aspace` and does **not** smuggle values
|
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle values
|
||||||
through registers — `startUserTask` reads the entry/stack from the Task and
|
through registers — `startUserTask` reads the entry/stack from the Task and
|
||||||
`jumpToUser`s ([scheduler.zig](../system/kernel/scheduler.zig)). That makes the thread
|
`jumpToUser`s ([scheduler.zig](../system/kernel/scheduler.zig)). That makes the thread
|
||||||
path clean:
|
path clean:
|
||||||
@@ -171,7 +171,7 @@ path clean:
|
|||||||
`{ fn_ptr, args_tuple, completion }`, the std "Instance" pattern — and writes the
|
`{ fn_ptr, args_tuple, completion }`, the std "Instance" pattern — and writes the
|
||||||
closure pointer to the **top word of the new stack**.
|
closure pointer to the **top word of the new stack**.
|
||||||
2. It calls `thread_spawn(entry = &threadTrampoline, stack_top, arg = closure_ptr)`.
|
2. It calls `thread_spawn(entry = &threadTrampoline, stack_top, arg = closure_ptr)`.
|
||||||
The kernel calls the same `spawnUserLocked` path with the **caller's aspace**
|
The kernel calls the same `spawnUserLocked` path with the **caller's address space**
|
||||||
(refcount++), `entry`, and `user_sp = stack_top`.
|
(refcount++), `entry`, and `user_sp = stack_top`.
|
||||||
3. `threadTrampoline` (a small runtime shim) reads the closure off its stack, calls
|
3. `threadTrampoline` (a small runtime shim) reads the closure off its stack, calls
|
||||||
the user function, then calls `thread_exit`. No new register ABI — the closure
|
the user function, then calls `thread_exit`. No new register ABI — the closure
|
||||||
@@ -185,7 +185,7 @@ Unlike a process start, there is **no** System V argc/argv/auxv block
|
|||||||
|
|
||||||
- **`thread_exit`** marks the task dead and hands the kernel the thread's user-stack
|
- **`thread_exit`** marks the task dead and hands the kernel the thread's user-stack
|
||||||
range. The kernel reaps the task on the scheduler (already running on a *kernel*
|
range. The kernel reaps the task on the scheduler (already running on a *kernel*
|
||||||
stack, so it can safely unmap the user stack), decrements the aspace refcount, and
|
stack, so it can safely unmap the user stack), decrements the address-space refcount, and
|
||||||
frees the task slot.
|
frees the task slot.
|
||||||
- **`join` — Stage 1** reuses the existing exit-notification machinery
|
- **`join` — Stage 1** reuses the existing exit-notification machinery
|
||||||
([process-lifecycle.md](process-lifecycle.md)): `spawn` passes a per-thread
|
([process-lifecycle.md](process-lifecycle.md)): `spawn` passes a per-thread
|
||||||
@@ -206,12 +206,12 @@ Unlike a process start, there is **no** System V argc/argv/auxv block
|
|||||||
call the futex wrappers on the slow path — the same construction `std.Thread` uses,
|
call the futex wrappers on the slow path — the same construction `std.Thread` uses,
|
||||||
so the algorithms port directly.
|
so the algorithms port directly.
|
||||||
|
|
||||||
Keying: threads share an address space, so a **virtual address within that aspace**
|
Keying: threads share an address space, so a **virtual address within that address space**
|
||||||
identifies a futex uniquely; the kernel keys its wait queue by `(aspace_root, vaddr)`.
|
identifies a futex uniquely; the kernel keys its wait queue by `(address_space_root, virtual_address)`.
|
||||||
Keying by the **physical** address instead (translate `vaddr -> paddr` on entry) is a
|
Keying by the **physical** address instead (translate `virtual_address -> physical_address` on entry) is a
|
||||||
deliberate forward door: it lets two *processes* share a futex through an
|
deliberate forward door: it lets two *processes* share a futex through an
|
||||||
[shm](display-v2.md) region later, without changing the API. We start with the
|
[shm](display-v2.md) region later, without changing the API. We start with the
|
||||||
private-per-aspace key and note the physical-key upgrade.
|
private-per-address-space key and note the physical-key upgrade.
|
||||||
|
|
||||||
No spinning: a contended lock parks the task in the kernel and the core is free to run
|
No spinning: a contended lock parks the task in the kernel and the core is free to run
|
||||||
other work or `hlt` ([halting.md](halting.md)). This is why futex is a locked
|
other work or `hlt` ([halting.md](halting.md)). This is why futex is a locked
|
||||||
@@ -238,14 +238,14 @@ it may call `runtime.Thread.spawn`. Everyone else stays single-threaded and lean
|
|||||||
## Interaction with the rest of the kernel
|
## Interaction with the rest of the kernel
|
||||||
|
|
||||||
- **Scheduler / SMP** ([scheduling.md](scheduling.md), [smp.md](smp.md)): a thread is
|
- **Scheduler / SMP** ([scheduling.md](scheduling.md), [smp.md](smp.md)): a thread is
|
||||||
just another `Task` with an `aspace` shared with its siblings; the existing
|
just another `Task` with an `address_space` shared with its siblings; the existing
|
||||||
per-core ready queues, priorities, and affinity apply unchanged. Threads of one
|
per-core ready queues, priorities, and affinity apply unchanged. Threads of one
|
||||||
process can run on different cores simultaneously — that is the point.
|
process can run on different cores simultaneously — that is the point.
|
||||||
- **Halting** ([halting.md](halting.md)): futex-parked waiters keep the "idle core
|
- **Halting** ([halting.md](halting.md)): futex-parked waiters keep the "idle core
|
||||||
halts" property intact under lock contention — no busy-wait.
|
halts" property intact under lock contention — no busy-wait.
|
||||||
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): killing a process
|
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): killing a process
|
||||||
must kill *all* its threads and only then drop the last aspace ref. The kill path
|
must kill *all* its threads and only then drop the last address-space ref. The kill path
|
||||||
already targets a process; it fans out to every task on that aspace.
|
already targets a process; it fans out to every task on that address space.
|
||||||
- **Resilience** ([resilience.md](resilience.md)): a faulting thread kills its whole
|
- **Resilience** ([resilience.md](resilience.md)): a faulting thread kills its whole
|
||||||
process (shared fate). The supervisor restarts the **process**, which respawns its
|
process (shared fate). The supervisor restarts the **process**, which respawns its
|
||||||
threads from a known-good state — restart granularity stays the process.
|
threads from a known-good state — restart granularity stays the process.
|
||||||
@@ -258,8 +258,8 @@ The ordered, `/loop`-runnable milestones live in
|
|||||||
a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
|
a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
|
||||||
`zig build test` for host unit tests). The stages below are the shape it expands.
|
`zig build test` for host unit tests). The stages below are the shape it expands.
|
||||||
|
|
||||||
- **Stage 0 — address-space refcount.** Refcount on the aspace root; teardown destroys
|
- **Stage 0 — address-space refcount.** Refcount on the address-space root; teardown destroys
|
||||||
at zero. No API yet; nothing shares an aspace, so refcount is 1 everywhere.
|
at zero. No API yet; nothing shares an address space, so refcount is 1 everywhere.
|
||||||
*Gate:* the full QEMU suite stays green (no regression) — proves the reframing is
|
*Gate:* the full QEMU suite stays green (no regression) — proves the reframing is
|
||||||
invisible until used.
|
invisible until used.
|
||||||
- **Stage 1 — spawn / join / detach.** `thread_spawn` + `thread_exit`, the trampoline,
|
- **Stage 1 — spawn / join / detach.** `thread_spawn` + `thread_exit`, the trampoline,
|
||||||
@@ -293,7 +293,7 @@ are — user code never names a syscall.
|
|||||||
- **No thread priorities distinct from the process.** Threads inherit the process
|
- **No thread priorities distinct from the process.** Threads inherit the process
|
||||||
priority; per-thread priority is a later question if it ever earns its keep.
|
priority; per-thread priority is a later question if it ever earns its keep.
|
||||||
- **No cross-process shared-memory futex yet** — the physical-address key leaves the
|
- **No cross-process shared-memory futex yet** — the physical-address key leaves the
|
||||||
door open, but the first cut is private-per-aspace.
|
door open, but the first cut is private-per-address-space.
|
||||||
- **No `pthread`/POSIX surface.** The API is `std.Thread`-shaped Zig, nothing more.
|
- **No `pthread`/POSIX surface.** The API is `std.Thread`-shaped Zig, nothing more.
|
||||||
|
|
||||||
## The self-hosting endgame
|
## The self-hosting endgame
|
||||||
|
|||||||
+1
-1
@@ -119,7 +119,7 @@ One table entry per kernel call, C ABI (System V AMD64), names prefixed
|
|||||||
returns are `u64`, errors return as negative values exactly as today.
|
returns are `u64`, errors return as negative values exactly as today.
|
||||||
|
|
||||||
The calls that return two values in `rax:rdx` today — `dma_alloc`
|
The calls that return two values in `rax:rdx` today — `dma_alloc`
|
||||||
(vaddr + paddr), `msi_bind` (address + data), `shm_create` (vaddr + handle) —
|
(virtual_address + physical_address), `msi_bind` (address + data), `shm_create` (virtual_address + handle) —
|
||||||
become functions returning a two-`u64` struct. The System V ABI returns a
|
become functions returning a two-`u64` struct. The System V ABI returns a
|
||||||
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
|
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
|
||||||
C-ABI spelling of the existing convention, at zero cost.
|
C-ABI spelling of the existing convention, at zero cost.
|
||||||
|
|||||||
@@ -22,7 +22,7 @@ pub const Region = struct {
|
|||||||
};
|
};
|
||||||
|
|
||||||
/// Grant `len` bytes (rounded up to whole pages) of shareable, zeroed, cacheable RAM.
|
/// Grant `len` bytes (rounded up to whole pages) of shareable, zeroed, cacheable RAM.
|
||||||
/// Returns the region or null on failure. Two return values — vaddr in rax, handle in rdx —
|
/// Returns the region or null on failure. Two return values — virtual_address in rax, handle in rdx —
|
||||||
/// so this is a hand-written stub like `dma.alloc`.
|
/// so this is a hand-written stub like `dma.alloc`.
|
||||||
pub fn create(len: usize) ?Region {
|
pub fn create(len: usize) ?Region {
|
||||||
var rax: usize = undefined;
|
var rax: usize = undefined;
|
||||||
|
|||||||
+7
-7
@@ -39,13 +39,13 @@ pub const SystemCall = enum(u64) {
|
|||||||
ipc_reply_wait = 10, // ipc_reply_wait(h, reply, len, receive, cap) -> receive_len (+badge in rdx)
|
ipc_reply_wait = 10, // ipc_reply_wait(h, reply, len, receive, cap) -> receive_len (+badge in rdx)
|
||||||
device_enumerate = 11, // device_enumerate(buffer, maximum) -> count: snapshot the device table
|
device_enumerate = 11, // device_enumerate(buffer, maximum) -> count: snapshot the device table
|
||||||
device_claim = 12, // device_claim(id) -> ok: take exclusive ownership of a device
|
device_claim = 12, // device_claim(id) -> ok: take exclusive ownership of a device
|
||||||
mmio_map = 13, // mmio_map(id, resource_index) -> vaddr: map a claimed device's MMIO into this AS
|
mmio_map = 13, // mmio_map(id, resource_index) -> virtual_address: map a claimed device's MMIO into this address space
|
||||||
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
|
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
|
||||||
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
|
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
|
||||||
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
|
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
|
||||||
system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint) -> child process id: start a named initial-ramdisk binary as a new ring-3 process
|
system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint) -> child process id: start a named initial-ramdisk binary as a new ring-3 process
|
||||||
dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory
|
dma_alloc = 18, // dma_alloc(len, flags) -> virtual_address (rax), physical_address (rdx): contiguous, pinned, uncacheable DMA memory
|
||||||
dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc
|
dma_free = 19, // dma_free(virtual_address, len) -> 0: release a prior dma_alloc
|
||||||
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
|
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
|
||||||
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
|
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
|
||||||
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
|
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
|
||||||
@@ -60,9 +60,9 @@ pub const SystemCall = enum(u64) {
|
|||||||
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
|
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
|
||||||
klog_read = 32, // klog_read(offset, ptr, len) -> bytes copied: copy the kernel RAM log buffer out to a user buffer (for persisting the boot log to disk)
|
klog_read = 32, // klog_read(offset, ptr, len) -> bytes copied: copy the kernel RAM log buffer out to a user buffer (for persisting the boot log to disk)
|
||||||
wall_clock = 33, // wall_clock() -> Unix epoch seconds (UTC): the RTC wall-clock time, for filesystem timestamps (mtime). Monotonic time is `clock`.
|
wall_clock = 33, // wall_clock() -> Unix epoch seconds (UTC): the RTC wall-clock time, for filesystem timestamps (mtime). Monotonic time is `clock`.
|
||||||
shm_create = 34, // shm_create(len) -> vaddr (rax), handle (rdx): a shareable, zeroed, cacheable RAM region mapped into this AS; the handle is a capability passed to another process as an ipc_call send_cap (docs/display-v2.md)
|
shm_create = 34, // shm_create(len) -> virtual_address (rax), handle (rdx): a shareable, zeroed, cacheable RAM region mapped into this AS; the handle is a capability passed to another process as an ipc_call send_cap (docs/display-v2.md)
|
||||||
shm_map = 35, // shm_map(cap) -> vaddr: map the shared region named by a received capability into this AS (the same physical pages the creator sees)
|
shm_map = 35, // shm_map(cap) -> virtual_address: map the shared region named by a received capability into this address space (the same physical pages the creator sees)
|
||||||
shm_physical = 36, // shm_physical(cap) -> paddr: the guest-physical base of a shared region held by capability, so a driver can program it into a device (e.g. virtio-gpu attach_backing); the pages are contiguous (docs/display-v2.md)
|
shm_physical = 36, // shm_physical(cap) -> physical_address: the guest-physical base of a shared region held by capability, so a driver can program it into a device (e.g. virtio-gpu attach_backing); the pages are contiguous (docs/display-v2.md)
|
||||||
thread_spawn = 37, // thread_spawn(entry, stack_top, arg, exit_endpoint) -> tid: start a task sharing the caller's address space at `entry` on `stack_top`, `arg` in rdi; exit_endpoint (a handle, or no_cap) is notified when it ends — how join waits (docs/threading.md)
|
thread_spawn = 37, // thread_spawn(entry, stack_top, arg, exit_endpoint) -> tid: start a task sharing the caller's address space at `entry` on `stack_top`, `arg` in rdi; exit_endpoint (a handle, or no_cap) is notified when it ends — how join waits (docs/threading.md)
|
||||||
thread_exit = 38, // thread_exit(): end the calling thread, dropping one reference to its address space (destroyed on the last)
|
thread_exit = 38, // thread_exit(): end the calling thread, dropping one reference to its address space (destroyed on the last)
|
||||||
current_core = 39, // current_core() -> index: the dense 0-based index of the core the caller is running on (for parallelism/affinity introspection)
|
current_core = 39, // current_core() -> index: the dense 0-based index of the core the caller is running on (for parallelism/affinity introspection)
|
||||||
@@ -70,7 +70,7 @@ pub const SystemCall = enum(u64) {
|
|||||||
futex_wake = 41, // futex_wake(addr, count) -> woken: wake up to `count` tasks blocked in futex_wait on `addr` in this address space
|
futex_wake = 41, // futex_wake(addr, count) -> woken: wake up to `count` tasks blocked in futex_wait on `addr` in this address space
|
||||||
thread_self = 42, // thread_self() -> tid: the calling thread's kernel task id (runtime.Thread.getCurrentId)
|
thread_self = 42, // thread_self() -> tid: the calling thread's kernel task id (runtime.Thread.getCurrentId)
|
||||||
thread_join = 43, // thread_join(tid) -> 0: block until the thread with id `tid` has exited (runtime.Thread.join; no per-thread IPC endpoint) (docs/threading.md)
|
thread_join = 43, // thread_join(tid) -> 0: block until the thread with id `tid` has exited (runtime.Thread.join; no per-thread IPC endpoint) (docs/threading.md)
|
||||||
set_thread_pointer = 44, // set_thread_pointer(addr) -> 0: set the caller's FS base (x86_64 user TLS thread pointer); restored per task across context switches (docs/threading-plan.md M10)
|
set_thread_pointer = 44, // set_thread_pointer(addr) -> 0: set the caller's thread pointer (user-space TLS base; x86_64 IA32_FS_BASE, aarch64 TPIDR_EL0); restored per task across context switches (docs/threading-plan.md M10)
|
||||||
_,
|
_,
|
||||||
};
|
};
|
||||||
|
|
||||||
|
|||||||
@@ -509,7 +509,7 @@ pub fn unmapInto(pml4: u64, virtual: u64) void {
|
|||||||
/// any address space, not just the live one). Returns null if `virtual` is not
|
/// any address space, not just the live one). Returns null if `virtual` is not
|
||||||
/// mapped at any level. Stops at a 2 MiB huge-page leaf (the physmap uses them),
|
/// mapped at any level. Stops at a 2 MiB huge-page leaf (the physmap uses them),
|
||||||
/// resolving the offset within it. The foundation for cross-address-space copies
|
/// resolving the offset within it. The foundation for cross-address-space copies
|
||||||
/// and for munmap (which needs the frame behind a user vaddr to free it).
|
/// and for munmap (which needs the frame behind a user virtual_address to free it).
|
||||||
pub fn translateIn(pml4: u64, virtual: u64) ?u64 {
|
pub fn translateIn(pml4: u64, virtual: u64) ?u64 {
|
||||||
const pml4e = tableAt(pml4)[(virtual >> 39) & 0x1FF];
|
const pml4e = tableAt(pml4)[(virtual >> 39) & 0x1FF];
|
||||||
if (pml4e & present == 0) return null;
|
if (pml4e & present == 0) return null;
|
||||||
|
|||||||
@@ -355,7 +355,7 @@ pub fn replyWait(endpoint: *Endpoint, reply_ptr: u64, reply_len: u64, receive_pt
|
|||||||
me.ipc_client = null;
|
me.ipc_client = null;
|
||||||
const n = @min(reply_len, client.ipc_reply_cap);
|
const n = @min(reply_len, client.ipc_reply_cap);
|
||||||
client.ipc_received_cap = abi.no_cap;
|
client.ipc_received_cap = abi.no_cap;
|
||||||
if (!copyAcross(me.aspace, reply_ptr, client.aspace, client.ipc_reply_ptr, n)) {
|
if (!copyAcross(me.address_space, reply_ptr, client.address_space, client.ipc_reply_ptr, n)) {
|
||||||
client.ipc_status = -EFAULT;
|
client.ipc_status = -EFAULT;
|
||||||
} else if (send_cap != abi.no_cap) {
|
} else if (send_cap != abi.no_cap) {
|
||||||
// Transfer the reply's capability into the client. A failure fails the
|
// Transfer the reply's capability into the client. A failure fails the
|
||||||
@@ -383,8 +383,8 @@ pub fn replyWait(endpoint: *Endpoint, reply_ptr: u64, reply_len: u64, receive_pt
|
|||||||
}
|
}
|
||||||
if (popPost(endpoint)) |slot| {
|
if (popPost(endpoint)) |slot| {
|
||||||
const n = @min(@as(usize, slot.length), receive_cap);
|
const n = @min(@as(usize, slot.length), receive_cap);
|
||||||
// Copy from the kernel-resident ring slot (source aspace 0) into the receiver.
|
// Copy from the kernel-resident ring slot (source address_space 0) into the receiver.
|
||||||
if (!copyAcross(0, @intFromPtr(&slot.bytes), me.aspace, receive_ptr, n)) {
|
if (!copyAcross(0, @intFromPtr(&slot.bytes), me.address_space, receive_ptr, n)) {
|
||||||
continue; // bad receive buffer: drop this message, keep serving
|
continue; // bad receive buffer: drop this message, keep serving
|
||||||
}
|
}
|
||||||
out_badge.* = slot.sender_id | notify_badge_bit | notify_message_bit;
|
out_badge.* = slot.sender_id | notify_badge_bit | notify_message_bit;
|
||||||
@@ -392,7 +392,7 @@ pub fn replyWait(endpoint: *Endpoint, reply_ptr: u64, reply_len: u64, receive_pt
|
|||||||
}
|
}
|
||||||
if (dequeueSender(endpoint)) |caller| {
|
if (dequeueSender(endpoint)) |caller| {
|
||||||
const n = @min(caller.ipc_send_len, receive_cap);
|
const n = @min(caller.ipc_send_len, receive_cap);
|
||||||
if (!copyAcross(caller.aspace, caller.ipc_send_ptr, me.aspace, receive_ptr, n)) {
|
if (!copyAcross(caller.address_space, caller.ipc_send_ptr, me.address_space, receive_ptr, n)) {
|
||||||
caller.ipc_status = -EFAULT; // bad sender buffer: fail it, keep serving
|
caller.ipc_status = -EFAULT; // bad sender buffer: fail it, keep serving
|
||||||
scheduler.readyLocked(caller);
|
scheduler.readyLocked(caller);
|
||||||
continue;
|
continue;
|
||||||
|
|||||||
+78
-77
@@ -59,7 +59,7 @@ pub const stack_top_virtual: u64 = stack_base_virtual + parameters.user_stack_pa
|
|||||||
/// The mmap grant arena: where `mmap` hands out fresh user pages, above the image
|
/// The mmap grant arena: where `mmap` hands out fresh user pages, above the image
|
||||||
/// and stack but still inside PML4[224] (so no kernel mapping is widened). Each
|
/// and stack but still inside PML4[224] (so no kernel mapping is widened). Each
|
||||||
/// process bump-allocates from `heap_arena_base` upward via a per-address-space cursor
|
/// process bump-allocates from `heap_arena_base` upward via a per-address-space cursor
|
||||||
/// (`scheduler.aspaceMmapNextPtr`, shared by its threads); a 1 GiB window is far more
|
/// (`scheduler.addressSpaceMmapNextPtr`, shared by its threads); a 1 GiB window is far more
|
||||||
/// than any user heap needs today.
|
/// than any user heap needs today.
|
||||||
pub const heap_arena_base: u64 = 0x0000_7000_1000_0000;
|
pub const heap_arena_base: u64 = 0x0000_7000_1000_0000;
|
||||||
pub const heap_arena_end: u64 = heap_arena_base + (1 << 30);
|
pub const heap_arena_end: u64 = heap_arena_base + (1 << 30);
|
||||||
@@ -71,7 +71,7 @@ pub const user_half_end: u64 = 0x0000_8000_0000_0000;
|
|||||||
/// The MMIO-grant arena: where `mmio_map` places device windows, in PML4[226] —
|
/// The MMIO-grant arena: where `mmio_map` places device windows, in PML4[226] —
|
||||||
/// a user-exclusive region distinct from code/stack/heap (PML4[224]), so mapping
|
/// a user-exclusive region distinct from code/stack/heap (PML4[224]), so mapping
|
||||||
/// device pages user-accessible widens no kernel mapping. Per-address-space cursor
|
/// device pages user-accessible widens no kernel mapping. Per-address-space cursor
|
||||||
/// (`scheduler.aspaceDeviceMapNextPtr`).
|
/// (`scheduler.addressSpaceDeviceMapNextPtr`).
|
||||||
pub const device_arena_base: u64 = 0x0000_7100_0000_0000;
|
pub const device_arena_base: u64 = 0x0000_7100_0000_0000;
|
||||||
pub const device_arena_end: u64 = device_arena_base + (4 << 30);
|
pub const device_arena_end: u64 = device_arena_base + (4 << 30);
|
||||||
|
|
||||||
@@ -164,7 +164,7 @@ fn fail(state: *architecture.CpuState) void {
|
|||||||
|
|
||||||
fn system_call(state: *architecture.CpuState) void {
|
fn system_call(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
const user = t.aspace != 0;
|
const user = t.address_space != 0;
|
||||||
if (user) {
|
if (user) {
|
||||||
// A condemned process (process_kill caught it running) dies at its next
|
// A condemned process (process_kill caught it running) dies at its next
|
||||||
// kernel entry — before it can spawn, claim, or message anything else.
|
// kernel entry — before it can spawn, claim, or message anything else.
|
||||||
@@ -317,7 +317,7 @@ fn systemIpcReplyWait(state: *architecture.CpuState) void {
|
|||||||
fn systemIpcSend(state: *architecture.CpuState) void {
|
fn systemIpcSend(state: *architecture.CpuState) void {
|
||||||
const me = scheduler.current();
|
const me = scheduler.current();
|
||||||
const endpoint = ipc.resolveHandle(me, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
|
const endpoint = ipc.resolveHandle(me, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
|
||||||
const r = ipc.send(endpoint, me.aspace, architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), me.id);
|
const r = ipc.send(endpoint, me.address_space, architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), me.id);
|
||||||
architecture.setSystemCallResult(state, @bitCast(r));
|
architecture.setSystemCallResult(state, @bitCast(r));
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -327,7 +327,7 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
|
|||||||
const buffer_ptr = architecture.systemCallArg(state, 0);
|
const buffer_ptr = architecture.systemCallArg(state, 0);
|
||||||
const maximum = architecture.systemCallArg(state, 1);
|
const maximum = architecture.systemCallArg(state, 1);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0 or buffer_ptr >= user_half_end) return fail(state);
|
if (t.address_space == 0 or buffer_ptr >= user_half_end) return fail(state);
|
||||||
const sz = @sizeOf(device_abi.DeviceDescriptor);
|
const sz = @sizeOf(device_abi.DeviceDescriptor);
|
||||||
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
|
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
|
||||||
const out: [*]device_abi.DeviceDescriptor = @ptrFromInt(buffer_ptr);
|
const out: [*]device_abi.DeviceDescriptor = @ptrFromInt(buffer_ptr);
|
||||||
@@ -351,14 +351,14 @@ fn systemDeviceClaim(state: *architecture.CpuState) void {
|
|||||||
} else fail(state);
|
} else fail(state);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// mmio_map(device_id, resource_index) -> vaddr: map a claimed device's MMIO window into
|
/// mmio_map(device_id, resource_index) -> virtual_address: map a claimed device's MMIO window into
|
||||||
/// this address space (strong-uncacheable) and return the register base address.
|
/// this address space (strong-uncacheable) and return the register base address.
|
||||||
/// The claim is the capability — a process can only map hardware it owns.
|
/// The claim is the capability — a process can only map hardware it owns.
|
||||||
fn systemMmioMap(state: *architecture.CpuState) void {
|
fn systemMmioMap(state: *architecture.CpuState) void {
|
||||||
const device_id = architecture.systemCallArg(state, 0);
|
const device_id = architecture.systemCallArg(state, 0);
|
||||||
const resource_index = architecture.systemCallArg(state, 1);
|
const resource_index = architecture.systemCallArg(state, 1);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
// Read the broker table under the lock: ring-3 device_register (M19) now
|
// Read the broker table under the lock: ring-3 device_register (M19) now
|
||||||
// mutates it concurrently on other cores, so a lock-free read here could
|
// mutates it concurrently on other cores, so a lock-free read here could
|
||||||
// see a torn resource (and a torn length used to panic the arithmetic
|
// see a torn resource (and a torn length used to panic the arithmetic
|
||||||
@@ -387,11 +387,11 @@ fn systemMmioMap(state: *architecture.CpuState) void {
|
|||||||
// same as mmap (docs/threading-plan.md M7).
|
// same as mmap (docs/threading-plan.md M7).
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
const cursor = scheduler.aspaceDeviceMapNextPtr(t.aspace) orelse return fail(state);
|
const cursor = scheduler.addressSpaceDeviceMapNextPtr(t.address_space) orelse return fail(state);
|
||||||
if (cursor.* == 0) cursor.* = device_arena_base; // seed the arena lazily
|
if (cursor.* == 0) cursor.* = device_arena_base; // seed the arena lazily
|
||||||
const base_v = cursor.*;
|
const base_v = cursor.*;
|
||||||
if (base_v + pages * page_size > device_arena_end) return fail(state);
|
if (base_v + pages * page_size > device_arena_end) return fail(state);
|
||||||
architecture.mapUserDeviceInto(t.aspace, base_v, r.start, r.len, write_combining);
|
architecture.mapUserDeviceInto(t.address_space, base_v, r.start, r.len, write_combining);
|
||||||
cursor.* = base_v + pages * page_size;
|
cursor.* = base_v + pages * page_size;
|
||||||
architecture.setSystemCallResult(state, base_v + (r.start & (page_size - 1))); // register base
|
architecture.setSystemCallResult(state, base_v + (r.start & (page_size - 1))); // register base
|
||||||
}
|
}
|
||||||
@@ -421,7 +421,7 @@ pub fn resolveIoPort(t: *scheduler.Task, device_id: u64, resource_index: u64, of
|
|||||||
/// is fine. See docs/drivers.md.
|
/// is fine. See docs/drivers.md.
|
||||||
fn systemIoRead(state: *architecture.CpuState) void {
|
fn systemIoRead(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const width = architecture.systemCallArg(state, 3);
|
const width = architecture.systemCallArg(state, 3);
|
||||||
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
|
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
|
||||||
architecture.setSystemCallResult(state, architecture.pioRead(@intCast(width), port));
|
architecture.setSystemCallResult(state, architecture.pioRead(@intCast(width), port));
|
||||||
@@ -432,14 +432,14 @@ fn systemIoRead(state: *architecture.CpuState) void {
|
|||||||
/// gate as `io_read`.
|
/// gate as `io_read`.
|
||||||
fn systemIoWrite(state: *architecture.CpuState) void {
|
fn systemIoWrite(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const width = architecture.systemCallArg(state, 3);
|
const width = architecture.systemCallArg(state, 3);
|
||||||
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
|
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
|
||||||
architecture.pioWrite(@intCast(width), port, @intCast(architecture.systemCallArg(state, 4)));
|
architecture.pioWrite(@intCast(width), port, @intCast(architecture.systemCallArg(state, 4)));
|
||||||
architecture.setSystemCallResult(state, 0);
|
architecture.setSystemCallResult(state, 0);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): grant `len` bytes (rounded up to
|
/// dma_alloc(len, flags) -> virtual_address (rax), physical_address (rdx): grant `len` bytes (rounded up to
|
||||||
/// whole pages) of DMA-capable memory — physically contiguous, zeroed, pinned, and
|
/// whole pages) of DMA-capable memory — physically contiguous, zeroed, pinned, and
|
||||||
/// strong-uncacheable (coherent) — mapping it into the caller's DMA arena and handing
|
/// strong-uncacheable (coherent) — mapping it into the caller's DMA arena and handing
|
||||||
/// back both the virtual address to touch and the physical address to program into the
|
/// back both the virtual address to touch and the physical address to program into the
|
||||||
@@ -451,7 +451,7 @@ fn systemDmaAlloc(state: *architecture.CpuState) void {
|
|||||||
const len = architecture.systemCallArg(state, 0);
|
const len = architecture.systemCallArg(state, 0);
|
||||||
const flags = architecture.systemCallArg(state, 1);
|
const flags = architecture.systemCallArg(state, 1);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0 or len == 0) return fail(state);
|
if (t.address_space == 0 or len == 0) return fail(state);
|
||||||
|
|
||||||
const pages: usize = @intCast((len + page_size - 1) / page_size);
|
const pages: usize = @intCast((len + page_size - 1) / page_size);
|
||||||
const max_phys: u64 = if (flags & abi.dma_below_4g != 0) (@as(u64, 4) << 30) else ~@as(u64, 0);
|
const max_phys: u64 = if (flags & abi.dma_below_4g != 0) (@as(u64, 4) << 30) else ~@as(u64, 0);
|
||||||
@@ -467,14 +467,14 @@ fn systemDmaAlloc(state: *architecture.CpuState) void {
|
|||||||
// Zero through the physmap (the frames aren't mapped in the caller yet), then map.
|
// Zero through the physmap (the frames aren't mapped in the caller yet), then map.
|
||||||
const kernel_view: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(phys));
|
const kernel_view: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(phys));
|
||||||
@memset(kernel_view[0 .. pages * page_size], 0);
|
@memset(kernel_view[0 .. pages * page_size], 0);
|
||||||
architecture.mapUserDmaInto(t.aspace, base_v, phys, pages * page_size);
|
architecture.mapUserDmaInto(t.address_space, base_v, phys, pages * page_size);
|
||||||
|
|
||||||
t.dma_map_next = base_v + pages * page_size;
|
t.dma_map_next = base_v + pages * page_size;
|
||||||
architecture.setSystemCallResult(state, base_v); // virtual address for the CPU
|
architecture.setSystemCallResult(state, base_v); // virtual address for the CPU
|
||||||
architecture.setSystemCallResult2(state, phys); // physical address for the device
|
architecture.setSystemCallResult2(state, phys); // physical address for the device
|
||||||
}
|
}
|
||||||
|
|
||||||
/// dma_free(vaddr, len) -> 0: release a prior `dma_alloc`. Bounded to the DMA arena so
|
/// dma_free(virtual_address, len) -> 0: release a prior `dma_alloc`. Bounded to the DMA arena so
|
||||||
/// it can never unmap-and-free the caller's stack, heap, or an MMIO grant; only pages
|
/// it can never unmap-and-free the caller's stack, heap, or an MMIO grant; only pages
|
||||||
/// actually mapped are freed (an unmapped hole is skipped). Teardown also reclaims any
|
/// actually mapped are freed (an unmapped hole is skipped). Teardown also reclaims any
|
||||||
/// DMA pages left mapped at exit (they carry no `device_grant`, so `freeSubtree` frees
|
/// DMA pages left mapped at exit (they carry no `device_grant`, so `freeSubtree` frees
|
||||||
@@ -483,21 +483,21 @@ fn systemDmaFree(state: *architecture.CpuState) void {
|
|||||||
const base_v = architecture.systemCallArg(state, 0);
|
const base_v = architecture.systemCallArg(state, 0);
|
||||||
const len = architecture.systemCallArg(state, 1);
|
const len = architecture.systemCallArg(state, 1);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const pages: usize = @intCast((len + page_size - 1) / page_size);
|
const pages: usize = @intCast((len + page_size - 1) / page_size);
|
||||||
if (base_v < dma_arena_base or base_v + pages * page_size > dma_arena_end) return fail(state);
|
if (base_v < dma_arena_base or base_v + pages * page_size > dma_arena_end) return fail(state);
|
||||||
|
|
||||||
for (0..pages) |i| {
|
for (0..pages) |i| {
|
||||||
const va = base_v + i * page_size;
|
const va = base_v + i * page_size;
|
||||||
if (architecture.translate(t.aspace, va)) |phys| {
|
if (architecture.translate(t.address_space, va)) |phys| {
|
||||||
architecture.unmapUserPageInto(t.aspace, va);
|
architecture.unmapUserPageInto(t.address_space, va);
|
||||||
pmm.free(phys);
|
pmm.free(phys);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
architecture.setSystemCallResult(state, 0);
|
architecture.setSystemCallResult(state, 0);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// shm_create(len) -> vaddr (rax), handle (rdx): grant `len` bytes (rounded up to whole
|
/// shm_create(len) -> virtual_address (rax), handle (rdx): grant `len` bytes (rounded up to whole
|
||||||
/// pages) of **shareable, zeroed, cacheable** RAM — contiguous frames mapped into the
|
/// pages) of **shareable, zeroed, cacheable** RAM — contiguous frames mapped into the
|
||||||
/// caller's shm arena — and hand back the virtual address plus a capability handle. Unlike
|
/// caller's shm arena — and hand back the virtual address plus a capability handle. Unlike
|
||||||
/// `dma_alloc` the memory is write-back cacheable (for CPU compositing, not device DMA) and
|
/// `dma_alloc` the memory is write-back cacheable (for CPU compositing, not device DMA) and
|
||||||
@@ -508,7 +508,7 @@ fn systemDmaFree(state: *architecture.CpuState) void {
|
|||||||
fn systemShmCreate(state: *architecture.CpuState) void {
|
fn systemShmCreate(state: *architecture.CpuState) void {
|
||||||
const len = architecture.systemCallArg(state, 0);
|
const len = architecture.systemCallArg(state, 0);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0 or len == 0) return fail(state);
|
if (t.address_space == 0 or len == 0) return fail(state);
|
||||||
|
|
||||||
const pages: usize = @intCast((len + page_size - 1) / page_size);
|
const pages: usize = @intCast((len + page_size - 1) / page_size);
|
||||||
if (pages == 0 or pages > maximum_shm_pages) return fail(state);
|
if (pages == 0 or pages > maximum_shm_pages) return fail(state);
|
||||||
@@ -533,20 +533,20 @@ fn systemShmCreate(state: *architecture.CpuState) void {
|
|||||||
return fail(state);
|
return fail(state);
|
||||||
}
|
}
|
||||||
|
|
||||||
architecture.mapUserSharedInto(t.aspace, base_v, phys, pages * page_size);
|
architecture.mapUserSharedInto(t.address_space, base_v, phys, pages * page_size);
|
||||||
t.shm_map_next = base_v + pages * page_size;
|
t.shm_map_next = base_v + pages * page_size;
|
||||||
architecture.setSystemCallResult(state, base_v); // vaddr for the CPU
|
architecture.setSystemCallResult(state, base_v); // virtual_address for the CPU
|
||||||
architecture.setSystemCallResult2(state, @intCast(handle)); // capability handle to pass on
|
architecture.setSystemCallResult2(state, @intCast(handle)); // capability handle to pass on
|
||||||
}
|
}
|
||||||
|
|
||||||
/// shm_map(cap) -> vaddr: map the shared region named by a capability handle the caller
|
/// shm_map(cap) -> virtual_address: map the shared region named by a capability handle the caller
|
||||||
/// received (via an `ipc_call` send_cap) into its shm arena — the same physical frames the
|
/// received (via an `ipc_call` send_cap) into its shm arena — the same physical frames the
|
||||||
/// creator sees — returning the virtual address. The handle already holds a reference (taken
|
/// creator sees — returning the virtual address. The handle already holds a reference (taken
|
||||||
/// when the capability was shared), so this only adds a mapping; it never bumps the refcount.
|
/// when the capability was shared), so this only adds a mapping; it never bumps the refcount.
|
||||||
fn systemShmMap(state: *architecture.CpuState) void {
|
fn systemShmMap(state: *architecture.CpuState) void {
|
||||||
const cap = architecture.systemCallArg(state, 0);
|
const cap = architecture.systemCallArg(state, 0);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
|
|
||||||
const shm = ipc.resolveShm(t, cap) orelse return fail(state); // not an shm handle we hold
|
const shm = ipc.resolveShm(t, cap) orelse return fail(state); // not an shm handle we hold
|
||||||
if (t.shm_map_next == 0) t.shm_map_next = shm_arena_base;
|
if (t.shm_map_next == 0) t.shm_map_next = shm_arena_base;
|
||||||
@@ -554,12 +554,12 @@ fn systemShmMap(state: *architecture.CpuState) void {
|
|||||||
const size = shm.pages * page_size;
|
const size = shm.pages * page_size;
|
||||||
if (base_v + size > shm_arena_end) return fail(state);
|
if (base_v + size > shm_arena_end) return fail(state);
|
||||||
|
|
||||||
architecture.mapUserSharedInto(t.aspace, base_v, shm.phys, size);
|
architecture.mapUserSharedInto(t.address_space, base_v, shm.phys, size);
|
||||||
t.shm_map_next = base_v + size;
|
t.shm_map_next = base_v + size;
|
||||||
architecture.setSystemCallResult(state, base_v);
|
architecture.setSystemCallResult(state, base_v);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// shm_physical(cap) -> paddr: the guest-physical base of a shared region the caller holds a
|
/// shm_physical(cap) -> physical_address: the guest-physical base of a shared region the caller holds a
|
||||||
/// capability for. The frames are contiguous (allocated by `allocContiguous`), so a single
|
/// capability for. The frames are contiguous (allocated by `allocContiguous`), so a single
|
||||||
/// physical base + length describes the whole region — which is exactly what a driver needs
|
/// physical base + length describes the whole region — which is exactly what a driver needs
|
||||||
/// to hand a shm surface to a device (virtio-gpu `attach_backing`). Only a holder of the
|
/// to hand a shm surface to a device (virtio-gpu `attach_backing`). Only a holder of the
|
||||||
@@ -567,7 +567,7 @@ fn systemShmMap(state: *architecture.CpuState) void {
|
|||||||
fn systemShmPhysical(state: *architecture.CpuState) void {
|
fn systemShmPhysical(state: *architecture.CpuState) void {
|
||||||
const cap = architecture.systemCallArg(state, 0);
|
const cap = architecture.systemCallArg(state, 0);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const shm = ipc.resolveShm(t, cap) orelse return fail(state); // not an shm handle we hold
|
const shm = ipc.resolveShm(t, cap) orelse return fail(state); // not an shm handle we hold
|
||||||
architecture.setSystemCallResult(state, shm.phys);
|
architecture.setSystemCallResult(state, shm.phys);
|
||||||
}
|
}
|
||||||
@@ -588,10 +588,10 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
|
|||||||
const parent_id = architecture.systemCallArg(state, 0);
|
const parent_id = architecture.systemCallArg(state, 0);
|
||||||
const descriptor_ptr = architecture.systemCallArg(state, 1);
|
const descriptor_ptr = architecture.systemCallArg(state, 1);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
|
|
||||||
var descriptor: device_abi.DeviceDescriptor = undefined;
|
var descriptor: device_abi.DeviceDescriptor = undefined;
|
||||||
if (!ipc.copyFromUser(t.aspace, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
|
if (!ipc.copyFromUser(t.address_space, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
|
||||||
|
|
||||||
// Under the big kernel lock: the broker's table is also mutated by the
|
// Under the big kernel lock: the broker's table is also mutated by the
|
||||||
// death sweep (releaseAllOwnedBy) and read by enumerate on other cores —
|
// death sweep (releaseAllOwnedBy) and read by enumerate on other cores —
|
||||||
@@ -675,7 +675,7 @@ fn systemThreadSpawn(state: *architecture.CpuState) void {
|
|||||||
const arg = architecture.systemCallArg(state, 2);
|
const arg = architecture.systemCallArg(state, 2);
|
||||||
const exit_handle = architecture.systemCallArg(state, 3);
|
const exit_handle = architecture.systemCallArg(state, 3);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state); // kernel tasks own no address space to share
|
if (t.address_space == 0) return fail(state); // kernel tasks own no address space to share
|
||||||
if (entry == 0 or entry >= user_half_end) return fail(state);
|
if (entry == 0 or entry >= user_half_end) return fail(state);
|
||||||
if (stack_top == 0 or stack_top > user_half_end) return fail(state);
|
if (stack_top == 0 or stack_top > user_half_end) return fail(state);
|
||||||
// The endpoint the thread notifies on exit (how join waits), or none.
|
// The endpoint the thread notifies on exit (how join waits), or none.
|
||||||
@@ -683,17 +683,17 @@ fn systemThreadSpawn(state: *architecture.CpuState) void {
|
|||||||
null
|
null
|
||||||
else
|
else
|
||||||
ipc.resolveHandle(t, exit_handle) orelse return failErr(state, ipc.EBADF);
|
ipc.resolveHandle(t, exit_handle) orelse return failErr(state, ipc.EBADF);
|
||||||
const tid = spawnThreadSupervised(t.aspace, entry, stack_top, arg, t.priority, t.id, exit_endpoint) orelse return fail(state);
|
const tid = spawnThreadSupervised(t.address_space, entry, stack_top, arg, t.priority, t.id, exit_endpoint) orelse return fail(state);
|
||||||
architecture.setSystemCallResult(state, tid);
|
architecture.setSystemCallResult(state, tid);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Spawn a thread sharing `aspace`, taking the exit-endpoint reference under the **same**
|
/// Spawn a thread sharing `address_space`, taking the exit-endpoint reference under the **same**
|
||||||
/// lock as the spawn (as `spawnProcessSupervised` does), so the thread cannot die before
|
/// lock as the spawn (as `spawnProcessSupervised` does), so the thread cannot die before
|
||||||
/// its reference exists. Returns the new thread id, or null on resource exhaustion.
|
/// its reference exists. Returns the new thread id, or null on resource exhaustion.
|
||||||
fn spawnThreadSupervised(aspace: u64, entry: u64, stack_top: u64, arg: u64, priority: scheduler.Priority, supervisor: u32, exit_endpoint: ?*ipc.Endpoint) ?u32 {
|
fn spawnThreadSupervised(address_space: u64, entry: u64, stack_top: u64, arg: u64, priority: scheduler.Priority, supervisor: u32, exit_endpoint: ?*ipc.Endpoint) ?u32 {
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
const tid = scheduler.spawnUserLocked(aspace, entry, stack_top, arg, priority, "thread", supervisor, if (exit_endpoint) |e| @ptrCast(e) else null) orelse return null;
|
const tid = scheduler.spawnUserLocked(address_space, entry, stack_top, arg, priority, "thread", supervisor, if (exit_endpoint) |e| @ptrCast(e) else null) orelse return null;
|
||||||
if (exit_endpoint) |endpoint| endpoint.refcount += 1; // the thread holds it birth-to-death
|
if (exit_endpoint) |endpoint| endpoint.refcount += 1; // the thread holds it birth-to-death
|
||||||
return tid;
|
return tid;
|
||||||
}
|
}
|
||||||
@@ -708,13 +708,14 @@ fn systemThreadSelf(state: *architecture.CpuState) void {
|
|||||||
architecture.setSystemCallResult(state, scheduler.currentId());
|
architecture.setSystemCallResult(state, scheduler.currentId());
|
||||||
}
|
}
|
||||||
|
|
||||||
/// set_thread_pointer(addr) -> 0: set the caller's FS base (its user-space TLS thread
|
/// set_thread_pointer(addr) -> 0: set the caller's user-space TLS thread pointer. The
|
||||||
/// pointer). The kernel never uses FS; the scheduler restores this per task across context
|
/// arch layer maps it to IA32_FS_BASE on x86_64, `TPIDR_EL0` on aarch64; the kernel
|
||||||
/// switches (docs/threading-plan.md M10). `addr` must be a user-half address.
|
/// never reads it, and the scheduler restores it per task across context switches
|
||||||
|
/// (docs/threading-plan.md M10). `addr` must be a user-half address.
|
||||||
fn systemSetThreadPointer(state: *architecture.CpuState) void {
|
fn systemSetThreadPointer(state: *architecture.CpuState) void {
|
||||||
const addr = architecture.systemCallArg(state, 0);
|
const addr = architecture.systemCallArg(state, 0);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state); // kernel tasks have no user TLS
|
if (t.address_space == 0) return fail(state); // kernel tasks have no user TLS
|
||||||
if (addr >= user_half_end) return fail(state);
|
if (addr >= user_half_end) return fail(state);
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
scheduler.setThreadPointerLocked(addr);
|
scheduler.setThreadPointerLocked(addr);
|
||||||
@@ -728,7 +729,7 @@ fn systemSetThreadPointer(state: *architecture.CpuState) void {
|
|||||||
fn systemThreadJoin(state: *architecture.CpuState) void {
|
fn systemThreadJoin(state: *architecture.CpuState) void {
|
||||||
const tid: u32 = @truncate(architecture.systemCallArg(state, 0));
|
const tid: u32 = @truncate(architecture.systemCallArg(state, 0));
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state); // kernel tasks don't join
|
if (t.address_space == 0) return fail(state); // kernel tasks don't join
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
scheduler.joinThreadLocked(tid);
|
scheduler.joinThreadLocked(tid);
|
||||||
sync.leave(flags);
|
sync.leave(flags);
|
||||||
@@ -745,12 +746,12 @@ fn systemFutexWait(state: *architecture.CpuState) void {
|
|||||||
const expected: u32 = @truncate(architecture.systemCallArg(state, 1));
|
const expected: u32 = @truncate(architecture.systemCallArg(state, 1));
|
||||||
const timeout_ns = architecture.systemCallArg(state, 2);
|
const timeout_ns = architecture.systemCallArg(state, 2);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
if (addr == 0 or (addr & 3) != 0 or addr + 4 > user_half_end) return fail(state);
|
if (addr == 0 or (addr & 3) != 0 or addr + 4 > user_half_end) return fail(state);
|
||||||
|
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
var word_bytes: [4]u8 = undefined;
|
var word_bytes: [4]u8 = undefined;
|
||||||
if (!ipc.copyFromUser(t.aspace, addr, &word_bytes)) {
|
if (!ipc.copyFromUser(t.address_space, addr, &word_bytes)) {
|
||||||
sync.leave(flags);
|
sync.leave(flags);
|
||||||
return fail(state);
|
return fail(state);
|
||||||
}
|
}
|
||||||
@@ -774,10 +775,10 @@ fn systemFutexWake(state: *architecture.CpuState) void {
|
|||||||
const addr = architecture.systemCallArg(state, 0);
|
const addr = architecture.systemCallArg(state, 0);
|
||||||
const count: u32 = @truncate(architecture.systemCallArg(state, 1));
|
const count: u32 = @truncate(architecture.systemCallArg(state, 1));
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
if (addr == 0 or (addr & 3) != 0 or addr + 4 > user_half_end) return fail(state);
|
if (addr == 0 or (addr & 3) != 0 or addr + 4 > user_half_end) return fail(state);
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
const woken = scheduler.futexWakeLocked(t.aspace, addr, count);
|
const woken = scheduler.futexWakeLocked(t.address_space, addr, count);
|
||||||
sync.leave(flags);
|
sync.leave(flags);
|
||||||
architecture.setSystemCallResult(state, woken);
|
architecture.setSystemCallResult(state, woken);
|
||||||
}
|
}
|
||||||
@@ -792,7 +793,7 @@ fn systemProcessEnumerate(state: *architecture.CpuState) void {
|
|||||||
const buffer_ptr = architecture.systemCallArg(state, 0);
|
const buffer_ptr = architecture.systemCallArg(state, 0);
|
||||||
const maximum = architecture.systemCallArg(state, 1);
|
const maximum = architecture.systemCallArg(state, 1);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0 or buffer_ptr >= user_half_end) return fail(state);
|
if (t.address_space == 0 or buffer_ptr >= user_half_end) return fail(state);
|
||||||
const sz = @sizeOf(abi.ProcessDescriptor);
|
const sz = @sizeOf(abi.ProcessDescriptor);
|
||||||
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
|
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
|
||||||
const out: [*]abi.ProcessDescriptor = @ptrFromInt(buffer_ptr);
|
const out: [*]abi.ProcessDescriptor = @ptrFromInt(buffer_ptr);
|
||||||
@@ -805,7 +806,7 @@ fn systemProcessEnumerate(state: *architecture.CpuState) void {
|
|||||||
/// cannot be a weapon (ids are never reused, so a stale one just misses).
|
/// cannot be a weapon (ids are never reused, so a stale one just misses).
|
||||||
fn systemProcessKill(state: *architecture.CpuState) void {
|
fn systemProcessKill(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const id = architecture.systemCallArg(state, 0);
|
const id = architecture.systemCallArg(state, 0);
|
||||||
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
|
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
|
||||||
const r = killProcess(t.id, @intCast(id));
|
const r = killProcess(t.id, @intCast(id));
|
||||||
@@ -938,7 +939,7 @@ pub fn killProcess(caller_id: u32, target_id: u32) i64 {
|
|||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
|
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
|
||||||
if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes
|
if (target.address_space == 0) return -ipc.ESRCH; // kernel tasks are not processes
|
||||||
if (target.supervisor != caller_id) return -ipc.EPERM;
|
if (target.supervisor != caller_id) return -ipc.EPERM;
|
||||||
target.exit_reason = .killed;
|
target.exit_reason = .killed;
|
||||||
if (target.state == .running) {
|
if (target.state == .running) {
|
||||||
@@ -1008,7 +1009,7 @@ var exit_subscribers: [exit_subscriber_capacity]?ExitSubscriber = .{null} ** exi
|
|||||||
/// secret between cooperating processes. -ENOSPC when the table is full.
|
/// secret between cooperating processes. -ENOSPC when the table is full.
|
||||||
fn systemProcessSubscribe(state: *architecture.CpuState) void {
|
fn systemProcessSubscribe(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
|
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
@@ -1028,7 +1029,7 @@ fn systemProcessSubscribe(state: *architecture.CpuState) void {
|
|||||||
/// delivered immediately on bind, coalesced into one notification.
|
/// delivered immediately on bind, coalesced into one notification.
|
||||||
fn systemSignalBind(state: *architecture.CpuState) void {
|
fn systemSignalBind(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
|
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
@@ -1048,7 +1049,7 @@ fn systemSignalBind(state: *architecture.CpuState) void {
|
|||||||
/// targets accumulate the signal in their pending mask.
|
/// targets accumulate the signal in their pending mask.
|
||||||
fn systemProcessSignal(state: *architecture.CpuState) void {
|
fn systemProcessSignal(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const id = architecture.systemCallArg(state, 0);
|
const id = architecture.systemCallArg(state, 0);
|
||||||
const signal = architecture.systemCallArg(state, 1);
|
const signal = architecture.systemCallArg(state, 1);
|
||||||
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
|
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
|
||||||
@@ -1056,7 +1057,7 @@ fn systemProcessSignal(state: *architecture.CpuState) void {
|
|||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
const target = scheduler.taskByIdLocked(@intCast(id)) orelse return failErr(state, ipc.ESRCH);
|
const target = scheduler.taskByIdLocked(@intCast(id)) orelse return failErr(state, ipc.ESRCH);
|
||||||
if (target.aspace == 0) return failErr(state, ipc.ESRCH);
|
if (target.address_space == 0) return failErr(state, ipc.ESRCH);
|
||||||
if (target.supervisor != t.id and target.id != t.id) return failErr(state, ipc.EPERM);
|
if (target.supervisor != t.id and target.id != t.id) return failErr(state, ipc.EPERM);
|
||||||
target.pending_signals |= @as(u32, 1) << @intCast(signal);
|
target.pending_signals |= @as(u32, 1) << @intCast(signal);
|
||||||
if (target.signal_endpoint) |raw| {
|
if (target.signal_endpoint) |raw| {
|
||||||
@@ -1093,7 +1094,7 @@ fn timerSweepLocked() void {
|
|||||||
/// timer_bind(endpoint, ms): arm a one-shot timer. -ENOSPC when the table is full.
|
/// timer_bind(endpoint, ms): arm a one-shot timer. -ENOSPC when the table is full.
|
||||||
fn systemTimerBind(state: *architecture.CpuState) void {
|
fn systemTimerBind(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
|
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
|
||||||
const ms = architecture.systemCallArg(state, 1);
|
const ms = architecture.systemCallArg(state, 1);
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
@@ -1110,7 +1111,7 @@ fn systemTimerBind(state: *architecture.CpuState) void {
|
|||||||
|
|
||||||
fn systemProcessExitReason(state: *architecture.CpuState) void {
|
fn systemProcessExitReason(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const id = architecture.systemCallArg(state, 0);
|
const id = architecture.systemCallArg(state, 0);
|
||||||
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
|
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
|
||||||
const r = exitReasonOf(t.id, @intCast(id));
|
const r = exitReasonOf(t.id, @intCast(id));
|
||||||
@@ -1136,7 +1137,7 @@ fn ownedGsi(t: *scheduler.Task, device_id: u64, resource_index: u64) ?u32 {
|
|||||||
/// IPC_ReplyWait and is woken by the ISR; see system/kernel/irq.zig for the cycle.
|
/// IPC_ReplyWait and is woken by the ISR; see system/kernel/irq.zig for the cycle.
|
||||||
fn systemIrqBind(state: *architecture.CpuState) void {
|
fn systemIrqBind(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const gsi = ownedGsi(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1)) orelse
|
const gsi = ownedGsi(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1)) orelse
|
||||||
return fail(state);
|
return fail(state);
|
||||||
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 2)) orelse return fail(state);
|
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 2)) orelse return fail(state);
|
||||||
@@ -1156,7 +1157,7 @@ fn systemIrqBind(state: *architecture.CpuState) void {
|
|||||||
fn systemMsiBind(state: *architecture.CpuState) void {
|
fn systemMsiBind(state: *architecture.CpuState) void {
|
||||||
const device_id = architecture.systemCallArg(state, 0);
|
const device_id = architecture.systemCallArg(state, 0);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
|
const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
|
||||||
if (owner != t.id) return fail(state); // not claimed by this process
|
if (owner != t.id) return fail(state); // not claimed by this process
|
||||||
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 1)) orelse return failErr(state, ipc.EBADF);
|
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 1)) orelse return failErr(state, ipc.EBADF);
|
||||||
@@ -1175,7 +1176,7 @@ fn systemMsiBind(state: *architecture.CpuState) void {
|
|||||||
/// more arrives until the driver says it has serviced the hardware.
|
/// more arrives until the driver says it has serviced the hardware.
|
||||||
fn systemIrqAck(state: *architecture.CpuState) void {
|
fn systemIrqAck(state: *architecture.CpuState) void {
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state);
|
if (t.address_space == 0) return fail(state);
|
||||||
const gsi = ownedGsi(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1)) orelse
|
const gsi = ownedGsi(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1)) orelse
|
||||||
return fail(state);
|
return fail(state);
|
||||||
|
|
||||||
@@ -1263,7 +1264,7 @@ fn systemKlogRead(state: *architecture.CpuState) void {
|
|||||||
fn systemMmap(state: *architecture.CpuState) void {
|
fn systemMmap(state: *architecture.CpuState) void {
|
||||||
const len = architecture.systemCallArg(state, 0);
|
const len = architecture.systemCallArg(state, 0);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0) return fail(state); // not a user process — nothing to map into
|
if (t.address_space == 0) return fail(state); // not a user process — nothing to map into
|
||||||
const pages = (len + page_size - 1) / page_size;
|
const pages = (len + page_size - 1) / page_size;
|
||||||
if (pages == 0 or pages > maximum_mmap_pages) return fail(state);
|
if (pages == 0 or pages > maximum_mmap_pages) return fail(state);
|
||||||
|
|
||||||
@@ -1275,7 +1276,7 @@ fn systemMmap(state: *architecture.CpuState) void {
|
|||||||
const base = reserve: {
|
const base = reserve: {
|
||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
const cursor = scheduler.aspaceMmapNextPtr(t.aspace) orelse return fail(state);
|
const cursor = scheduler.addressSpaceMmapNextPtr(t.address_space) orelse return fail(state);
|
||||||
if (cursor.* == 0) cursor.* = heap_arena_base; // seed the arena lazily
|
if (cursor.* == 0) cursor.* = heap_arena_base; // seed the arena lazily
|
||||||
const b = cursor.*;
|
const b = cursor.*;
|
||||||
if (b + pages * page_size > heap_arena_end) return fail(state); // arena exhausted
|
if (b + pages * page_size > heap_arena_end) return fail(state); // arena exhausted
|
||||||
@@ -1295,8 +1296,8 @@ fn systemMmap(state: *architecture.CpuState) void {
|
|||||||
var i: usize = 0;
|
var i: usize = 0;
|
||||||
while (i < mapped) : (i += 1) {
|
while (i < mapped) : (i += 1) {
|
||||||
const va = base + i * page_size;
|
const va = base + i * page_size;
|
||||||
if (architecture.translate(t.aspace, va)) |physical| {
|
if (architecture.translate(t.address_space, va)) |physical| {
|
||||||
architecture.unmapUserPageInto(t.aspace, va);
|
architecture.unmapUserPageInto(t.address_space, va);
|
||||||
pmm.free(physical);
|
pmm.free(physical);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1305,7 +1306,7 @@ fn systemMmap(state: *architecture.CpuState) void {
|
|||||||
};
|
};
|
||||||
const destination: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(frame));
|
const destination: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(frame));
|
||||||
@memset(destination[0..page_size], 0); // hand out zeroed memory
|
@memset(destination[0..page_size], 0); // hand out zeroed memory
|
||||||
architecture.mapUserPageInto(t.aspace, base + mapped * page_size, frame, true, false); // RW + NX
|
architecture.mapUserPageInto(t.address_space, base + mapped * page_size, frame, true, false); // RW + NX
|
||||||
sync.leave(flags);
|
sync.leave(flags);
|
||||||
}
|
}
|
||||||
architecture.setSystemCallResult(state, base); // the cursor was already advanced at reserve
|
architecture.setSystemCallResult(state, base); // the cursor was already advanced at reserve
|
||||||
@@ -1320,14 +1321,14 @@ fn systemMunmap(state: *architecture.CpuState) void {
|
|||||||
const base = architecture.systemCallArg(state, 0);
|
const base = architecture.systemCallArg(state, 0);
|
||||||
const len = architecture.systemCallArg(state, 1);
|
const len = architecture.systemCallArg(state, 1);
|
||||||
const t = scheduler.current();
|
const t = scheduler.current();
|
||||||
if (t.aspace == 0 or base % page_size != 0) return fail(state);
|
if (t.address_space == 0 or base % page_size != 0) return fail(state);
|
||||||
const pages = (len + page_size - 1) / page_size;
|
const pages = (len + page_size - 1) / page_size;
|
||||||
if (base < heap_arena_base or base + pages * page_size > heap_arena_end) return fail(state);
|
if (base < heap_arena_base or base + pages * page_size > heap_arena_end) return fail(state);
|
||||||
|
|
||||||
for (0..pages) |i| {
|
for (0..pages) |i| {
|
||||||
const va = base + i * page_size;
|
const va = base + i * page_size;
|
||||||
if (architecture.translate(t.aspace, va)) |physical| {
|
if (architecture.translate(t.address_space, va)) |physical| {
|
||||||
architecture.unmapUserPageInto(t.aspace, va);
|
architecture.unmapUserPageInto(t.address_space, va);
|
||||||
pmm.free(physical);
|
pmm.free(physical);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1391,7 +1392,7 @@ const maximum_segments = 16;
|
|||||||
const maximum_pages = 256; // 1 MiB loader budget; the user region caps at 2 MiB anyway
|
const maximum_pages = 256; // 1 MiB loader budget; the user region caps at 2 MiB anyway
|
||||||
|
|
||||||
const Segment = struct {
|
const Segment = struct {
|
||||||
vaddr: u64,
|
virtual_address: u64,
|
||||||
memsz: u64,
|
memsz: u64,
|
||||||
filesz: u64,
|
filesz: u64,
|
||||||
off: u64,
|
off: u64,
|
||||||
@@ -1439,7 +1440,7 @@ fn parseSegments(image: []const u8, segs: *[maximum_segments]Segment) InitError!
|
|||||||
if (w and x) return error.BadSegment; // W^X, even for init
|
if (w and x) return error.BadSegment; // W^X, even for init
|
||||||
|
|
||||||
const seg = Segment{
|
const seg = Segment{
|
||||||
.vaddr = phdr.p_vaddr,
|
.virtual_address = phdr.p_vaddr,
|
||||||
.memsz = phdr.p_memsz,
|
.memsz = phdr.p_memsz,
|
||||||
.filesz = phdr.p_filesz,
|
.filesz = phdr.p_filesz,
|
||||||
.off = phdr.p_offset,
|
.off = phdr.p_offset,
|
||||||
@@ -1448,9 +1449,9 @@ fn parseSegments(image: []const u8, segs: *[maximum_segments]Segment) InitError!
|
|||||||
};
|
};
|
||||||
// No overlap with any earlier segment (page-granular, since mapping is).
|
// No overlap with any earlier segment (page-granular, since mapping is).
|
||||||
for (segs[0..count]) |other| {
|
for (segs[0..count]) |other| {
|
||||||
const a_end = seg.vaddr + seg.pages() * page_size;
|
const a_end = seg.virtual_address + seg.pages() * page_size;
|
||||||
const b_end = other.vaddr + other.pages() * page_size;
|
const b_end = other.virtual_address + other.pages() * page_size;
|
||||||
if (seg.vaddr < b_end and other.vaddr < a_end) return error.BadSegment;
|
if (seg.virtual_address < b_end and other.virtual_address < a_end) return error.BadSegment;
|
||||||
}
|
}
|
||||||
total_pages += seg.pages();
|
total_pages += seg.pages();
|
||||||
if (total_pages > maximum_pages) return error.ProgramTooBig;
|
if (total_pages > maximum_pages) return error.ProgramTooBig;
|
||||||
@@ -1461,17 +1462,17 @@ fn parseSegments(image: []const u8, segs: *[maximum_segments]Segment) InitError!
|
|||||||
|
|
||||||
// The entry point must land inside an executable segment.
|
// The entry point must land inside an executable segment.
|
||||||
for (segs[0..count]) |seg| {
|
for (segs[0..count]) |seg| {
|
||||||
if (seg.executable and ehdr.e_entry >= seg.vaddr and ehdr.e_entry < seg.vaddr + seg.memsz)
|
if (seg.executable and ehdr.e_entry >= seg.virtual_address and ehdr.e_entry < seg.virtual_address + seg.memsz)
|
||||||
return .{ .count = count, .entry = ehdr.e_entry };
|
return .{ .count = count, .entry = ehdr.e_entry };
|
||||||
}
|
}
|
||||||
return error.BadEntry;
|
return error.BadEntry;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Load one page of a segment into address space `aspace`: a fresh frame, zeroed
|
/// Load one page of a segment into address space `address_space`: a fresh frame, zeroed
|
||||||
/// and filled through the physmap, mapped user-accessible with the segment's W^X.
|
/// and filled through the physmap, mapped user-accessible with the segment's W^X.
|
||||||
/// On a later failure the whole address space is torn down, which frees every
|
/// On a later failure the whole address space is torn down, which frees every
|
||||||
/// frame mapped into it — so no per-page rollback list is needed here.
|
/// frame mapped into it — so no per-page rollback list is needed here.
|
||||||
fn loadPageInto(aspace: u64, image: []const u8, seg: Segment, page_index: u64) InitError!void {
|
fn loadPageInto(address_space: u64, image: []const u8, seg: Segment, page_index: u64) InitError!void {
|
||||||
const frame = pmm.alloc() orelse return error.OutOfMemory;
|
const frame = pmm.alloc() orelse return error.OutOfMemory;
|
||||||
const destination: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(frame));
|
const destination: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(frame));
|
||||||
@memset(destination[0..page_size], 0);
|
@memset(destination[0..page_size], 0);
|
||||||
@@ -1480,7 +1481,7 @@ fn loadPageInto(aspace: u64, image: []const u8, seg: Segment, page_index: u64) I
|
|||||||
const n = @min(page_size, seg.filesz - page_off);
|
const n = @min(page_size, seg.filesz - page_off);
|
||||||
@memcpy(destination[0..n], image[seg.off + page_off ..][0..n]);
|
@memcpy(destination[0..n], image[seg.off + page_off ..][0..n]);
|
||||||
}
|
}
|
||||||
architecture.mapUserPageInto(aspace, seg.vaddr + page_off, frame, seg.writable, seg.executable);
|
architecture.mapUserPageInto(address_space, seg.virtual_address + page_off, frame, seg.writable, seg.executable);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Build the System V AMD64 process-entry block at the top of a process's stack
|
/// Build the System V AMD64 process-entry block at the top of a process's stack
|
||||||
@@ -1567,11 +1568,11 @@ pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []c
|
|||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
|
|
||||||
const aspace = architecture.createAddressSpace() orelse return error.OutOfMemory;
|
const address_space = architecture.createAddressSpace() orelse return error.OutOfMemory;
|
||||||
errdefer architecture.destroyAddressSpace(aspace);
|
errdefer architecture.destroyAddressSpace(address_space);
|
||||||
|
|
||||||
for (segs[0..parsed.count]) |seg| {
|
for (segs[0..parsed.count]) |seg| {
|
||||||
for (0..seg.pages()) |i| try loadPageInto(aspace, image, seg, i);
|
for (0..seg.pages()) |i| try loadPageInto(address_space, image, seg, i);
|
||||||
}
|
}
|
||||||
|
|
||||||
// The stack: `user_stack_pages` zeroed pages below stack_top_virtual, RW + NX.
|
// The stack: `user_stack_pages` zeroed pages below stack_top_virtual, RW + NX.
|
||||||
@@ -1585,10 +1586,10 @@ pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []c
|
|||||||
const page_virtual = stack_base_virtual + i * page_size;
|
const page_virtual = stack_base_virtual + i * page_size;
|
||||||
if (i == parameters.user_stack_pages - 1)
|
if (i == parameters.user_stack_pages - 1)
|
||||||
user_sp = buildEntryStack(stack_page, page_virtual, argv);
|
user_sp = buildEntryStack(stack_page, page_virtual, argv);
|
||||||
architecture.mapUserPageInto(aspace, page_virtual, stack_frame, true, false); // RW + NX
|
architecture.mapUserPageInto(address_space, page_virtual, stack_frame, true, false); // RW + NX
|
||||||
}
|
}
|
||||||
|
|
||||||
const child = scheduler.spawnUserLocked(aspace, parsed.entry, user_sp, 0, priority, argv[0], supervisor, if (exit_endpoint) |endpoint| @ptrCast(endpoint) else null) orelse
|
const child = scheduler.spawnUserLocked(address_space, parsed.entry, user_sp, 0, priority, argv[0], supervisor, if (exit_endpoint) |endpoint| @ptrCast(endpoint) else null) orelse
|
||||||
return error.OutOfMemory;
|
return error.OutOfMemory;
|
||||||
// The child holds a reference to its exit endpoint from birth to death. Taken
|
// The child holds a reference to its exit endpoint from birth to death. Taken
|
||||||
// only now, after nothing can fail; the lock is still held, so the child
|
// only now, after nothing can fail; the lock is still held, so the child
|
||||||
|
|||||||
+48
-48
@@ -78,7 +78,7 @@ pub const Task = struct {
|
|||||||
ipc_wait_endpoint: ?*anyopaque = null,
|
ipc_wait_endpoint: ?*anyopaque = null,
|
||||||
// Physical root of this task's address space, or 0 for a kernel task (which
|
// Physical root of this task's address space, or 0 for a kernel task (which
|
||||||
// runs on the shared kernel page tables). A user task carries its own.
|
// runs on the shared kernel page tables). A user task carries its own.
|
||||||
aspace: u64 = 0,
|
address_space: u64 = 0,
|
||||||
user_ip: u64 = 0, // user-mode entry point (user task only)
|
user_ip: u64 = 0, // user-mode entry point (user task only)
|
||||||
user_sp: u64 = 0, // user-mode stack pointer (user task only)
|
user_sp: u64 = 0, // user-mode stack pointer (user task only)
|
||||||
user_arg: u64 = 0, // value delivered in the user's first argument register at first entry
|
user_arg: u64 = 0, // value delivered in the user's first argument register at first entry
|
||||||
@@ -95,8 +95,8 @@ pub const Task = struct {
|
|||||||
// aarch64. Restored on every context switch to this task (docs/threading-plan.md M10).
|
// aarch64. Restored on every context switch to this task (docs/threading-plan.md M10).
|
||||||
thread_pointer: u64 = 0,
|
thread_pointer: u64 = 0,
|
||||||
// The mmap / MMIO grant-arena cursors moved from Task to the per-address-space object
|
// The mmap / MMIO grant-arena cursors moved from Task to the per-address-space object
|
||||||
// (`AspaceRef`, below) so threads sharing one address space hand out disjoint grants
|
// (`AddressSpaceRef`, below) so threads sharing one address space hand out disjoint grants
|
||||||
// — see aspaceMmapNextPtr / aspaceDeviceMapNextPtr (docs/threading-plan.md M7).
|
// — see addressSpaceMmapNextPtr / addressSpaceDeviceMapNextPtr (docs/threading-plan.md M7).
|
||||||
// --- synchronous IPC (ipc_sync.zig) ---
|
// --- synchronous IPC (ipc_sync.zig) ---
|
||||||
// Per-process handle table: a small-int handle names a kernel capability object.
|
// Per-process handle table: a small-int handle names a kernel capability object.
|
||||||
// Each entry tags its `kind` (an IPC endpoint or a shared-memory object) so the
|
// Each entry tags its `kind` (an IPC endpoint or a shared-memory object) so the
|
||||||
@@ -107,9 +107,9 @@ pub const Task = struct {
|
|||||||
// receive, cleared when it replies). A client, while blocked in Call, records
|
// receive, cleared when it replies). A client, while blocked in Call, records
|
||||||
// its message + reply buffers here and its result lands in `ipc_status`.
|
// its message + reply buffers here and its result lands in `ipc_status`.
|
||||||
ipc_client: ?*Task = null,
|
ipc_client: ?*Task = null,
|
||||||
ipc_send_ptr: u64 = 0, // client: outgoing message (vaddr in this task's AS)
|
ipc_send_ptr: u64 = 0, // client: outgoing message (virtual_address in this task's address space)
|
||||||
ipc_send_len: u64 = 0,
|
ipc_send_len: u64 = 0,
|
||||||
ipc_reply_ptr: u64 = 0, // client: reply buffer (vaddr)
|
ipc_reply_ptr: u64 = 0, // client: reply buffer (virtual_address)
|
||||||
ipc_reply_cap: u64 = 0,
|
ipc_reply_cap: u64 = 0,
|
||||||
ipc_status: i64 = 0, // client: reply length / -errno, written by the replier
|
ipc_status: i64 = 0, // client: reply length / -errno, written by the replier
|
||||||
dma_map_next: u64 = 0, // bump pointer into this task's DMA arena (0 = unseeded)
|
dma_map_next: u64 = 0, // bump pointer into this task's DMA arena (0 = unseeded)
|
||||||
@@ -159,9 +159,9 @@ var tasks = [_]Task{.{}} ** maximum_tasks;
|
|||||||
// One live entry per address space; threads sharing an address space share this entry,
|
// One live entry per address space; threads sharing an address space share this entry,
|
||||||
// so their mmap/mmio grants bump one cursor and never overlap (docs/threading-plan.md M7).
|
// so their mmap/mmio grants bump one cursor and never overlap (docs/threading-plan.md M7).
|
||||||
// `mmap_next`/`device_map_next` are 0 until process.zig seeds them to the arena base.
|
// `mmap_next`/`device_map_next` are 0 until process.zig seeds them to the arena base.
|
||||||
const AspaceRef = struct { root: u64 = 0, count: u32 = 0, mmap_next: u64 = 0, device_map_next: u64 = 0 };
|
const AddressSpaceRef = struct { root: u64 = 0, count: u32 = 0, mmap_next: u64 = 0, device_map_next: u64 = 0 };
|
||||||
var aspace_refs = [_]AspaceRef{.{}} ** maximum_tasks;
|
var address_space_refs = [_]AddressSpaceRef{.{}} ** maximum_tasks;
|
||||||
var aspace_destroy_count: u64 = 0;
|
var address_space_destroy_count: u64 = 0;
|
||||||
|
|
||||||
/// Total bytes of task **kernel** stacks currently allocated from the kernel heap —
|
/// Total bytes of task **kernel** stacks currently allocated from the kernel heap —
|
||||||
/// incremented when a task is created, decremented when the reaper frees a dead task's
|
/// incremented when a task is created, decremented when the reaper frees a dead task's
|
||||||
@@ -202,10 +202,10 @@ fn drainReapListLocked(pc: *PerCpu) void {
|
|||||||
/// Take a reference to address space `root` (0 = a kernel task, which owns none).
|
/// Take a reference to address space `root` (0 = a kernel task, which owns none).
|
||||||
/// Returns false only if the ref table is full — bounded by `maximum_tasks`, so in
|
/// Returns false only if the ref table is full — bounded by `maximum_tasks`, so in
|
||||||
/// practice it never is. Caller holds the kernel lock.
|
/// practice it never is. Caller holds the kernel lock.
|
||||||
fn retainAspace(root: u64) bool {
|
fn retainAddressSpace(root: u64) bool {
|
||||||
if (root == 0) return true;
|
if (root == 0) return true;
|
||||||
var free: ?*AspaceRef = null;
|
var free: ?*AddressSpaceRef = null;
|
||||||
for (&aspace_refs) |*entry| {
|
for (&address_space_refs) |*entry| {
|
||||||
if (entry.count != 0 and entry.root == root) {
|
if (entry.count != 0 and entry.root == root) {
|
||||||
entry.count += 1;
|
entry.count += 1;
|
||||||
return true;
|
return true;
|
||||||
@@ -220,51 +220,51 @@ fn retainAspace(root: u64) bool {
|
|||||||
/// Drop a reference to `root`; destroy the address space when the **last** one drops.
|
/// Drop a reference to `root`; destroy the address space when the **last** one drops.
|
||||||
/// A `root` with no entry — never retained, e.g. a hand-built test space — is
|
/// A `root` with no entry — never retained, e.g. a hand-built test space — is
|
||||||
/// destroyed directly, preserving the pre-refcount behaviour. Caller holds the lock.
|
/// destroyed directly, preserving the pre-refcount behaviour. Caller holds the lock.
|
||||||
fn releaseAspace(root: u64) void {
|
fn releaseAddressSpace(root: u64) void {
|
||||||
if (root == 0) return;
|
if (root == 0) return;
|
||||||
for (&aspace_refs) |*entry| {
|
for (&address_space_refs) |*entry| {
|
||||||
if (entry.count == 0 or entry.root != root) continue;
|
if (entry.count == 0 or entry.root != root) continue;
|
||||||
entry.count -= 1;
|
entry.count -= 1;
|
||||||
if (entry.count == 0) {
|
if (entry.count == 0) {
|
||||||
entry.root = 0;
|
entry.root = 0;
|
||||||
architecture.destroyAddressSpace(root);
|
architecture.destroyAddressSpace(root);
|
||||||
aspace_destroy_count += 1;
|
address_space_destroy_count += 1;
|
||||||
}
|
}
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
architecture.destroyAddressSpace(root);
|
architecture.destroyAddressSpace(root);
|
||||||
aspace_destroy_count += 1;
|
address_space_destroy_count += 1;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Test-observable: how many address spaces are live (entries with a nonzero count).
|
/// Test-observable: how many address spaces are live (entries with a nonzero count).
|
||||||
pub fn liveAspaceCount() u32 {
|
pub fn liveAddressSpaceCount() u32 {
|
||||||
var live: u32 = 0;
|
var live: u32 = 0;
|
||||||
for (&aspace_refs) |*entry| {
|
for (&address_space_refs) |*entry| {
|
||||||
if (entry.count != 0) live += 1;
|
if (entry.count != 0) live += 1;
|
||||||
}
|
}
|
||||||
return live;
|
return live;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Test-observable: total address-space destructions since boot.
|
/// Test-observable: total address-space destructions since boot.
|
||||||
pub fn aspaceDestroyCount() u64 {
|
pub fn addressSpaceDestroyCount() u64 {
|
||||||
return aspace_destroy_count;
|
return address_space_destroy_count;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Pointer to the mmap grant-arena cursor for address space `root`, so the mmap syscall
|
/// Pointer to the mmap grant-arena cursor for address space `root`, so the mmap syscall
|
||||||
/// can read-and-bump it. Per-address-space (not per-task), so sibling threads get
|
/// can read-and-bump it. Per-address-space (not per-task), so sibling threads get
|
||||||
/// disjoint grants. **Caller holds the kernel lock** (the entry is stable while held).
|
/// disjoint grants. **Caller holds the kernel lock** (the entry is stable while held).
|
||||||
/// Null only if `root` was never retained — which can't happen for a live user task.
|
/// Null only if `root` was never retained — which can't happen for a live user task.
|
||||||
pub fn aspaceMmapNextPtr(root: u64) ?*u64 {
|
pub fn addressSpaceMmapNextPtr(root: u64) ?*u64 {
|
||||||
for (&aspace_refs) |*entry| {
|
for (&address_space_refs) |*entry| {
|
||||||
if (entry.count != 0 and entry.root == root) return &entry.mmap_next;
|
if (entry.count != 0 and entry.root == root) return &entry.mmap_next;
|
||||||
}
|
}
|
||||||
return null;
|
return null;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Pointer to the MMIO grant-arena cursor for address space `root` (see
|
/// Pointer to the MMIO grant-arena cursor for address space `root` (see
|
||||||
/// `aspaceMmapNextPtr`). Caller holds the kernel lock.
|
/// `addressSpaceMmapNextPtr`). Caller holds the kernel lock.
|
||||||
pub fn aspaceDeviceMapNextPtr(root: u64) ?*u64 {
|
pub fn addressSpaceDeviceMapNextPtr(root: u64) ?*u64 {
|
||||||
for (&aspace_refs) |*entry| {
|
for (&address_space_refs) |*entry| {
|
||||||
if (entry.count != 0 and entry.root == root) return &entry.device_map_next;
|
if (entry.count != 0 and entry.root == root) return &entry.device_map_next;
|
||||||
}
|
}
|
||||||
return null;
|
return null;
|
||||||
@@ -288,7 +288,7 @@ pub const PerCpu = struct {
|
|||||||
hw_id: u32 = 0, // the core's hardware id (Local APIC id on x86_64)
|
hw_id: u32 = 0, // the core's hardware id (Local APIC id on x86_64)
|
||||||
index: u32 = 0, // dense 0-based core index
|
index: u32 = 0, // dense 0-based core index
|
||||||
online: bool = false, // has this core finished bring-up?
|
online: bool = false, // has this core finished bring-up?
|
||||||
loaded_aspace: u64 = 0, // the address-space root currently loaded on this core
|
loaded_address_space: u64 = 0, // the address-space root currently loaded on this core
|
||||||
loaded_thread_pointer: u64 = 0, // the TLS thread pointer currently loaded on this core (docs/threading-plan.md M10)
|
loaded_thread_pointer: u64 = 0, // the TLS thread pointer currently loaded on this core (docs/threading-plan.md M10)
|
||||||
// Tasks pinned to this core (affinity == index), per priority level + bitmap.
|
// Tasks pinned to this core (affinity == index), per priority level + bitmap.
|
||||||
pinned_head: [number_priorities]?*Task = .{null} ** number_priorities,
|
pinned_head: [number_priorities]?*Task = .{null} ** number_priorities,
|
||||||
@@ -331,7 +331,7 @@ var preemption_enabled = true;
|
|||||||
/// boot, before interrupts are enabled — so no lock is needed here.
|
/// boot, before interrupts are enabled — so no lock is needed here.
|
||||||
pub fn init(boot_priority: Priority) void {
|
pub fn init(boot_priority: Priority) void {
|
||||||
const pc = &cpus[0];
|
const pc = &cpus[0];
|
||||||
pc.* = .{ .index = 0, .online = true, .loaded_aspace = architecture.kernelPageTable() };
|
pc.* = .{ .index = 0, .online = true, .loaded_address_space = architecture.kernelPageTable() };
|
||||||
architecture.setCpuLocal(0, @intFromPtr(pc));
|
architecture.setCpuLocal(0, @intFromPtr(pc));
|
||||||
tasks[0] = .{ .id = 0, .state = .running, .priority = boot_priority };
|
tasks[0] = .{ .id = 0, .state = .running, .priority = boot_priority };
|
||||||
pc.current = &tasks[0];
|
pc.current = &tasks[0];
|
||||||
@@ -370,7 +370,7 @@ pub fn secondaryMain() callconv(.c) noreturn {
|
|||||||
pc.current = t;
|
pc.current = t;
|
||||||
pc.idle = t;
|
pc.idle = t;
|
||||||
pc.online = true;
|
pc.online = true;
|
||||||
pc.loaded_aspace = architecture.kernelPageTable(); // the AP adopted the kernel tables at bring-up
|
pc.loaded_address_space = architecture.kernelPageTable(); // the AP adopted the kernel tables at bring-up
|
||||||
sync.leave(flags);
|
sync.leave(flags);
|
||||||
|
|
||||||
architecture.enableInterrupts(); // the timer now preempts this idle context into work
|
architecture.enableInterrupts(); // the timer now preempts this idle context into work
|
||||||
@@ -456,7 +456,7 @@ pub fn spawnOn(entry: *const fn () void, priority: Priority, cpu: u32) bool {
|
|||||||
return ok;
|
return ok;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Spawn a **user** task: a task with its own address space (`aspace`) that starts
|
/// Spawn a **user** task: a task with its own address space (`address_space`) that starts
|
||||||
/// in user mode at `entry` on `user_sp`, recorded under `name` (its argv[0]).
|
/// in user mode at `entry` on `user_sp`, recorded under `name` (its argv[0]).
|
||||||
/// `supervisor` is the id of the spawning process (0 = the kernel) — the kill
|
/// `supervisor` is the id of the spawning process (0 = the kernel) — the kill
|
||||||
/// authority — and `exit_endpoint` (an *ipc.Endpoint whose reference the caller
|
/// authority — and `exit_endpoint` (an *ipc.Endpoint whose reference the caller
|
||||||
@@ -465,14 +465,14 @@ pub fn spawnOn(entry: *const fn () void, priority: Priority, cpu: u32) bool {
|
|||||||
/// lands in `user_task_trampoline`.
|
/// lands in `user_task_trampoline`.
|
||||||
/// Returns the new process id, or null (creating nothing) if the table is full or
|
/// Returns the new process id, or null (creating nothing) if the table is full or
|
||||||
/// out of memory.
|
/// out of memory.
|
||||||
/// **Caller must hold the kernel lock** (the loader that builds `aspace` holds it
|
/// **Caller must hold the kernel lock** (the loader that builds `address_space` holds it
|
||||||
/// across the whole spawn, so the address space and the task appear atomically).
|
/// across the whole spawn, so the address space and the task appear atomically).
|
||||||
pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, user_arg: u64, priority: Priority, task_name: []const u8, supervisor: u32, exit_endpoint: ?*anyopaque) ?u32 {
|
pub fn spawnUserLocked(address_space: u64, entry: u64, user_sp: u64, user_arg: u64, priority: Priority, task_name: []const u8, supervisor: u32, exit_endpoint: ?*anyopaque) ?u32 {
|
||||||
const t = freeSlot() orelse return null;
|
const t = freeSlot() orelse return null;
|
||||||
const stack = heap.allocator().alloc(u8, stack_size) catch return null;
|
const stack = heap.allocator().alloc(u8, stack_size) catch return null;
|
||||||
// Take this task's reference to the address space before we commit the slot, so a
|
// Take this task's reference to the address space before we commit the slot, so a
|
||||||
// failure here leaves nothing to unwind (the caller still owns the raw `aspace`).
|
// failure here leaves nothing to unwind (the caller still owns the raw `address_space`).
|
||||||
if (!retainAspace(aspace)) {
|
if (!retainAddressSpace(address_space)) {
|
||||||
heap.allocator().free(stack);
|
heap.allocator().free(stack);
|
||||||
return null;
|
return null;
|
||||||
}
|
}
|
||||||
@@ -482,7 +482,7 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, user_arg: u64, pri
|
|||||||
.state = .ready,
|
.state = .ready,
|
||||||
.priority = priority,
|
.priority = priority,
|
||||||
.stack = stack,
|
.stack = stack,
|
||||||
.aspace = aspace,
|
.address_space = address_space,
|
||||||
.user_ip = entry,
|
.user_ip = entry,
|
||||||
.user_sp = user_sp,
|
.user_sp = user_sp,
|
||||||
.user_arg = user_arg,
|
.user_arg = user_arg,
|
||||||
@@ -561,7 +561,7 @@ fn schedule() void {
|
|||||||
/// Make `next` this core's running task: publish its kernel stack (TSS.rsp0, so a
|
/// Make `next` this core's running task: publish its kernel stack (TSS.rsp0, so a
|
||||||
/// user-mode interrupt lands on a good stack) and its address space (only when
|
/// user-mode interrupt lands on a good stack) and its address space (only when
|
||||||
/// it differs from what's loaded — every page-table switch is a full TLB flush),
|
/// it differs from what's loaded — every page-table switch is a full TLB flush),
|
||||||
/// then switch registers/stacks. Kernel tasks (aspace == 0, no kstack_top used
|
/// then switch registers/stacks. Kernel tasks (address_space == 0, no kstack_top used
|
||||||
/// from user mode) resolve to the shared kernel page tables and skip the kernel-
|
/// from user mode) resolve to the shared kernel page tables and skip the kernel-
|
||||||
/// stack write, so this is a no-op beyond the register switch for a pure-kernel
|
/// stack write, so this is a no-op beyond the register switch for a pure-kernel
|
||||||
/// workload. The big kernel lock is held and interrupts are off throughout, so no
|
/// workload. The big kernel lock is held and interrupts are off throughout, so no
|
||||||
@@ -569,10 +569,10 @@ fn schedule() void {
|
|||||||
/// `save_sp` receives the outgoing task's stack pointer.
|
/// `save_sp` receives the outgoing task's stack pointer.
|
||||||
fn switchTo(pc: *PerCpu, save_sp: *usize, next: *Task) void {
|
fn switchTo(pc: *PerCpu, save_sp: *usize, next: *Task) void {
|
||||||
if (next.kstack_top != 0) architecture.setKernelStack(pc.index, next.kstack_top);
|
if (next.kstack_top != 0) architecture.setKernelStack(pc.index, next.kstack_top);
|
||||||
const want = if (next.aspace != 0) next.aspace else architecture.kernelPageTable();
|
const want = if (next.address_space != 0) next.address_space else architecture.kernelPageTable();
|
||||||
if (want != pc.loaded_aspace) {
|
if (want != pc.loaded_address_space) {
|
||||||
architecture.loadPageTable(want);
|
architecture.loadPageTable(want);
|
||||||
pc.loaded_aspace = want;
|
pc.loaded_address_space = want;
|
||||||
}
|
}
|
||||||
// Restore the next task's user TLS thread pointer — only on change, the same
|
// Restore the next task's user TLS thread pointer — only on change, the same
|
||||||
// conditional-load discipline as CR3 above (docs/threading-plan.md M10).
|
// conditional-load discipline as CR3 above (docs/threading-plan.md M10).
|
||||||
@@ -635,12 +635,12 @@ pub fn futexWaitLocked(addr: u64, timeout_ms: u64) FutexResult {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// Wake up to `count` tasks blocked in `futex_wait` on `addr` in address space
|
/// Wake up to `count` tasks blocked in `futex_wait` on `addr` in address space
|
||||||
/// `aspace`. Precondition: the big kernel lock is held. Returns how many woke.
|
/// `address_space`. Precondition: the big kernel lock is held. Returns how many woke.
|
||||||
pub fn futexWakeLocked(aspace: u64, addr: u64, count: u32) u32 {
|
pub fn futexWakeLocked(address_space: u64, addr: u64, count: u32) u32 {
|
||||||
var woken: u32 = 0;
|
var woken: u32 = 0;
|
||||||
for (&tasks) |*t| {
|
for (&tasks) |*t| {
|
||||||
if (woken >= count) break;
|
if (woken >= count) break;
|
||||||
if (t.state == .blocked and t.aspace == aspace and t.futex_addr == addr) {
|
if (t.state == .blocked and t.address_space == address_space and t.futex_addr == addr) {
|
||||||
t.futex_addr = 0; // the "woken, not timed out" signal to futexWaitLocked
|
t.futex_addr = 0; // the "woken, not timed out" signal to futexWaitLocked
|
||||||
t.wake_at = 0;
|
t.wake_at = 0;
|
||||||
t.state = .ready;
|
t.state = .ready;
|
||||||
@@ -907,7 +907,7 @@ fn reapKillPendingLocked() void {
|
|||||||
// task_trampoline, not switchTo's tail), its stack is still queued here. The dying
|
// task_trampoline, not switchTo's tail), its stack is still queued here. The dying
|
||||||
// task switched away before this tick, so it is off its stack — drain now (M8).
|
// task switched away before this tick, so it is off its stack — drain now (M8).
|
||||||
drainReapListLocked(pc);
|
drainReapListLocked(pc);
|
||||||
if (cur.kill_pending and cur.aspace != 0 and !cur.in_system_call) {
|
if (cur.kill_pending and cur.address_space != 0 and !cur.in_system_call) {
|
||||||
if (terminate_current_hook) |hook| hook(); // noreturn
|
if (terminate_current_hook) |hook| hook(); // noreturn
|
||||||
}
|
}
|
||||||
if (reap_task_hook) |hook| {
|
if (reap_task_hook) |hook| {
|
||||||
@@ -977,16 +977,16 @@ pub fn exitUser() noreturn {
|
|||||||
pub fn exitUserLocked() noreturn {
|
pub fn exitUserLocked() noreturn {
|
||||||
const pc = thisCpu();
|
const pc = thisCpu();
|
||||||
const dying = pc.current;
|
const dying = pc.current;
|
||||||
const as = dying.aspace;
|
const as = dying.address_space;
|
||||||
if (as != 0) {
|
if (as != 0) {
|
||||||
const kroot = architecture.kernelPageTable();
|
const kroot = architecture.kernelPageTable();
|
||||||
architecture.loadPageTable(kroot); // off the process tables before freeing them
|
architecture.loadPageTable(kroot); // off the process tables before freeing them
|
||||||
pc.loaded_aspace = kroot;
|
pc.loaded_address_space = kroot;
|
||||||
releaseAspace(as); // destroys only when this was the last task on the space
|
releaseAddressSpace(as); // destroys only when this was the last task on the space
|
||||||
}
|
}
|
||||||
dying.state = .reaping; // dead but its slot stays reserved until the stack is freed
|
dying.state = .reaping; // dead but its slot stays reserved until the stack is freed
|
||||||
wakeJoinersLocked(dying.id); // let any thread_join(dying.id) return (M9)
|
wakeJoinersLocked(dying.id); // let any thread_join(dying.id) return (M9)
|
||||||
dying.aspace = 0;
|
dying.address_space = 0;
|
||||||
dying.kill_pending = false;
|
dying.kill_pending = false;
|
||||||
dying.in_system_call = false;
|
dying.in_system_call = false;
|
||||||
// Queue for reaping: the task we switch to (or the next tick) frees this stack (M8/M9).
|
// Queue for reaping: the task we switch to (or the next tick) frees this stack (M8/M9).
|
||||||
@@ -1007,9 +1007,9 @@ pub fn exitUserLocked() noreturn {
|
|||||||
/// task isn't running). The kernel stack is leaked, as in `exitUser` (no reaper
|
/// task isn't running). The kernel stack is leaked, as in `exitUser` (no reaper
|
||||||
/// yet). Precondition: the big kernel lock is held.
|
/// yet). Precondition: the big kernel lock is held.
|
||||||
pub fn destroyTaskLocked(t: *Task) void {
|
pub fn destroyTaskLocked(t: *Task) void {
|
||||||
if (t.aspace != 0) releaseAspace(t.aspace); // destroys only on the last reference
|
if (t.address_space != 0) releaseAddressSpace(t.address_space); // destroys only on the last reference
|
||||||
reapStackLocked(t); // safe to free now: `t` is not running on any core (M8)
|
reapStackLocked(t); // safe to free now: `t` is not running on any core (M8)
|
||||||
t.aspace = 0;
|
t.address_space = 0;
|
||||||
t.kill_pending = false;
|
t.kill_pending = false;
|
||||||
t.in_system_call = false;
|
t.in_system_call = false;
|
||||||
t.wake_at = 0;
|
t.wake_at = 0;
|
||||||
@@ -1052,7 +1052,7 @@ pub fn enumerate(out: []abi.ProcessDescriptor) u64 {
|
|||||||
|
|
||||||
/// Whether the running task is a user process (has its own address space).
|
/// Whether the running task is a user process (has its own address space).
|
||||||
pub fn currentIsUserProcess() bool {
|
pub fn currentIsUserProcess() bool {
|
||||||
return current().aspace != 0;
|
return current().address_space != 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
pub fn currentId() u32 {
|
pub fn currentId() u32 {
|
||||||
|
|||||||
+42
-42
@@ -139,8 +139,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
|||||||
userPfTest();
|
userPfTest();
|
||||||
} else if (eql(case, "fault-recovery")) {
|
} else if (eql(case, "fault-recovery")) {
|
||||||
faultRecoveryTest(boot_information);
|
faultRecoveryTest(boot_information);
|
||||||
} else if (eql(case, "aspace-refcount")) {
|
} else if (eql(case, "address-space-refcount")) {
|
||||||
aspaceRefcountTest(boot_information);
|
addressSpaceRefcountTest(boot_information);
|
||||||
} else if (eql(case, "thread-spawn")) {
|
} else if (eql(case, "thread-spawn")) {
|
||||||
threadSpawnTest(boot_information);
|
threadSpawnTest(boot_information);
|
||||||
} else if (eql(case, "thread-join")) {
|
} else if (eql(case, "thread-join")) {
|
||||||
@@ -920,12 +920,12 @@ fn userMemTest() void {
|
|||||||
log("DANOS-TEST-BEGIN: usermem\n", .{});
|
log("DANOS-TEST-BEGIN: usermem\n", .{});
|
||||||
const base_free = pmm.stats().free_frames;
|
const base_free = pmm.stats().free_frames;
|
||||||
|
|
||||||
const aspace = architecture.createAddressSpace() orelse {
|
const address_space = architecture.createAddressSpace() orelse {
|
||||||
check("created a fresh address space", false);
|
check("created a fresh address space", false);
|
||||||
result();
|
result();
|
||||||
return;
|
return;
|
||||||
};
|
};
|
||||||
check("created a fresh address space", aspace != 0);
|
check("created a fresh address space", address_space != 0);
|
||||||
|
|
||||||
// Grant three pages into the arena, mapped RW + NX (the mmap contract).
|
// Grant three pages into the arena, mapped RW + NX (the mmap contract).
|
||||||
const npages = 3;
|
const npages = 3;
|
||||||
@@ -934,7 +934,7 @@ fn userMemTest() void {
|
|||||||
var mapped: usize = 0;
|
var mapped: usize = 0;
|
||||||
while (mapped < npages) : (mapped += 1) {
|
while (mapped < npages) : (mapped += 1) {
|
||||||
frames[mapped] = pmm.alloc() orelse break;
|
frames[mapped] = pmm.alloc() orelse break;
|
||||||
architecture.mapUserPageInto(aspace, arena + mapped * abi.page_size, frames[mapped], true, false);
|
architecture.mapUserPageInto(address_space, arena + mapped * abi.page_size, frames[mapped], true, false);
|
||||||
}
|
}
|
||||||
check("granted three user pages", mapped == npages);
|
check("granted three user pages", mapped == npages);
|
||||||
|
|
||||||
@@ -943,7 +943,7 @@ fn userMemTest() void {
|
|||||||
var rw_ok = true;
|
var rw_ok = true;
|
||||||
for (0..npages) |i| {
|
for (0..npages) |i| {
|
||||||
const va = arena + i * abi.page_size;
|
const va = arena + i * abi.page_size;
|
||||||
const physical = architecture.translate(aspace, va) orelse {
|
const physical = architecture.translate(address_space, va) orelse {
|
||||||
translate_ok = false;
|
translate_ok = false;
|
||||||
continue;
|
continue;
|
||||||
};
|
};
|
||||||
@@ -958,13 +958,13 @@ fn userMemTest() void {
|
|||||||
// Release them the way munmap does, then tear down the address space.
|
// Release them the way munmap does, then tear down the address space.
|
||||||
for (0..npages) |i| {
|
for (0..npages) |i| {
|
||||||
const va = arena + i * abi.page_size;
|
const va = arena + i * abi.page_size;
|
||||||
if (architecture.translate(aspace, va)) |physical| {
|
if (architecture.translate(address_space, va)) |physical| {
|
||||||
architecture.unmapUserPageInto(aspace, va);
|
architecture.unmapUserPageInto(address_space, va);
|
||||||
pmm.free(physical);
|
pmm.free(physical);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
check("munmap unmapped every grant", architecture.translate(aspace, arena) == null);
|
check("munmap unmapped every grant", architecture.translate(address_space, arena) == null);
|
||||||
architecture.destroyAddressSpace(aspace);
|
architecture.destroyAddressSpace(address_space);
|
||||||
|
|
||||||
check("no frames leaked (free count restored)", pmm.stats().free_frames == base_free);
|
check("no frames leaked (free count restored)", pmm.stats().free_frames == base_free);
|
||||||
result();
|
result();
|
||||||
@@ -1132,12 +1132,12 @@ fn dmaTest() void {
|
|||||||
|
|
||||||
// Map the run into a fresh address space as coherent DMA and translate each page
|
// Map the run into a fresh address space as coherent DMA and translate each page
|
||||||
// back: the same physical run, in order — proving contiguity and the mapping.
|
// back: the same physical run, in order — proving contiguity and the mapping.
|
||||||
const aspace = architecture.createAddressSpace().?;
|
const address_space = architecture.createAddressSpace().?;
|
||||||
architecture.mapUserDmaInto(aspace, process.dma_arena_base, phys, frames * abi.page_size);
|
architecture.mapUserDmaInto(address_space, process.dma_arena_base, phys, frames * abi.page_size);
|
||||||
var mapped_ok = true;
|
var mapped_ok = true;
|
||||||
for (0..frames) |i| {
|
for (0..frames) |i| {
|
||||||
const va = process.dma_arena_base + i * abi.page_size;
|
const va = process.dma_arena_base + i * abi.page_size;
|
||||||
const got = architecture.translate(aspace, va) orelse {
|
const got = architecture.translate(address_space, va) orelse {
|
||||||
mapped_ok = false;
|
mapped_ok = false;
|
||||||
break;
|
break;
|
||||||
};
|
};
|
||||||
@@ -1147,7 +1147,7 @@ fn dmaTest() void {
|
|||||||
|
|
||||||
// Teardown must reclaim the DMA RAM (the leaves carry no device_grant, so
|
// Teardown must reclaim the DMA RAM (the leaves carry no device_grant, so
|
||||||
// freeSubtree frees them as ordinary frames) — a driver that just dies leaks none.
|
// freeSubtree frees them as ordinary frames) — a driver that just dies leaks none.
|
||||||
architecture.destroyAddressSpace(aspace);
|
architecture.destroyAddressSpace(address_space);
|
||||||
for (0..2) |i| pmm.free(low + i * abi.page_size);
|
for (0..2) |i| pmm.free(low + i * abi.page_size);
|
||||||
check("no frames leaked after DMA teardown", pmm.stats().free_frames == base_free);
|
check("no frames leaked after DMA teardown", pmm.stats().free_frames == base_free);
|
||||||
result();
|
result();
|
||||||
@@ -1379,9 +1379,9 @@ fn spawnFaultingProcess() ?u32 {
|
|||||||
const flags = sync.enter();
|
const flags = sync.enter();
|
||||||
defer sync.leave(flags);
|
defer sync.leave(flags);
|
||||||
|
|
||||||
const aspace = architecture.createAddressSpace() orelse return null;
|
const address_space = architecture.createAddressSpace() orelse return null;
|
||||||
const code_frame = pmm.alloc() orelse {
|
const code_frame = pmm.alloc() orelse {
|
||||||
architecture.destroyAddressSpace(aspace);
|
architecture.destroyAddressSpace(address_space);
|
||||||
return null;
|
return null;
|
||||||
};
|
};
|
||||||
// Fill through the physmap (the user mapping is read-only); pad with int3 so a
|
// Fill through the physmap (the user mapping is read-only); pad with int3 so a
|
||||||
@@ -1389,17 +1389,17 @@ fn spawnFaultingProcess() ?u32 {
|
|||||||
const code: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(code_frame));
|
const code: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(code_frame));
|
||||||
@memset(code[0..abi.page_size], 0xCC);
|
@memset(code[0..abi.page_size], 0xCC);
|
||||||
@memcpy(code[0..blob.len], blob);
|
@memcpy(code[0..blob.len], blob);
|
||||||
architecture.mapUserPageInto(aspace, process.code_virtual, code_frame, false, true); // RO + X
|
architecture.mapUserPageInto(address_space, process.code_virtual, code_frame, false, true); // RO + X
|
||||||
|
|
||||||
const stack_frame = pmm.alloc() orelse {
|
const stack_frame = pmm.alloc() orelse {
|
||||||
architecture.destroyAddressSpace(aspace); // frees code_frame too — it's mapped
|
architecture.destroyAddressSpace(address_space); // frees code_frame too — it's mapped
|
||||||
return null;
|
return null;
|
||||||
};
|
};
|
||||||
architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX
|
architecture.mapUserPageInto(address_space, process.stack_base_virtual, stack_frame, true, false); // RW + NX
|
||||||
|
|
||||||
// Supervised by the calling test task, so exitReasonOf can read the verdict.
|
// Supervised by the calling test task, so exitReasonOf can read the verdict.
|
||||||
const id = scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 0, 4, "fault-probe", scheduler.currentId(), null) orelse {
|
const id = scheduler.spawnUserLocked(address_space, process.code_virtual, process.stack_base_virtual + abi.page_size, 0, 4, "fault-probe", scheduler.currentId(), null) orelse {
|
||||||
architecture.destroyAddressSpace(aspace);
|
architecture.destroyAddressSpace(address_space);
|
||||||
return null;
|
return null;
|
||||||
};
|
};
|
||||||
return id;
|
return id;
|
||||||
@@ -1460,11 +1460,11 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
|
|||||||
/// address spaces returns to baseline while destructions advance by exactly that many.
|
/// address spaces returns to baseline while destructions advance by exactly that many.
|
||||||
/// This is the foundation threads (shared address spaces) build on: the refactor must be
|
/// This is the foundation threads (shared address spaces) build on: the refactor must be
|
||||||
/// invisible while every space still has exactly one task.
|
/// invisible while every space still has exactly one task.
|
||||||
fn aspaceRefcountTest(boot_information: *const BootInformation) void {
|
fn addressSpaceRefcountTest(boot_information: *const BootInformation) void {
|
||||||
_ = boot_information;
|
_ = boot_information;
|
||||||
log("DANOS-TEST-BEGIN: aspace-refcount\n", .{});
|
log("DANOS-TEST-BEGIN: address-space-refcount\n", .{});
|
||||||
const base_live = scheduler.liveAspaceCount();
|
const base_live = scheduler.liveAddressSpaceCount();
|
||||||
const base_destroyed = scheduler.aspaceDestroyCount();
|
const base_destroyed = scheduler.addressSpaceDestroyCount();
|
||||||
const rounds: u32 = 5;
|
const rounds: u32 = 5;
|
||||||
var killed: u32 = 0;
|
var killed: u32 = 0;
|
||||||
var round: u32 = 0;
|
var round: u32 = 0;
|
||||||
@@ -1480,11 +1480,11 @@ fn aspaceRefcountTest(boot_information: *const BootInformation) void {
|
|||||||
if (process.fault_kill_count >= 1) killed += 1;
|
if (process.fault_kill_count >= 1) killed += 1;
|
||||||
}
|
}
|
||||||
check("all probes spawned and were killed", killed == rounds);
|
check("all probes spawned and were killed", killed == rounds);
|
||||||
check("live address-space count returned to baseline", scheduler.liveAspaceCount() == base_live);
|
check("live address-space count returned to baseline", scheduler.liveAddressSpaceCount() == base_live);
|
||||||
check("each address space destroyed exactly once", scheduler.aspaceDestroyCount() == base_destroyed + rounds);
|
check("each address space destroyed exactly once", scheduler.addressSpaceDestroyCount() == base_destroyed + rounds);
|
||||||
if (killed == rounds and scheduler.liveAspaceCount() == base_live and
|
if (killed == rounds and scheduler.liveAddressSpaceCount() == base_live and
|
||||||
scheduler.aspaceDestroyCount() == base_destroyed + rounds)
|
scheduler.addressSpaceDestroyCount() == base_destroyed + rounds)
|
||||||
log("aspace-refcount: spaces released to baseline ok\n", .{});
|
log("address-space-refcount: spaces released to baseline ok\n", .{});
|
||||||
result();
|
result();
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1509,7 +1509,7 @@ fn threadSpawnTest(boot_information: *const BootInformation) void {
|
|||||||
check("thread-test spawned", spawnNamed(rd, "thread-test"));
|
check("thread-test spawned", spawnNamed(rd, "thread-test"));
|
||||||
|
|
||||||
// Wait for the service's verdict marker (it polls shared memory the worker wrote).
|
// Wait for the service's verdict marker (it polls shared memory the worker wrote).
|
||||||
const ok_marker = "thread-test: child ran in shared aspace ok";
|
const ok_marker = "thread-test: child ran in shared address space ok";
|
||||||
const fail_marker = "thread-test: FAIL";
|
const fail_marker = "thread-test: FAIL";
|
||||||
scheduler.setPriority(1);
|
scheduler.setPriority(1);
|
||||||
const deadline = architecture.millis() + 12000;
|
const deadline = architecture.millis() + 12000;
|
||||||
@@ -1740,7 +1740,7 @@ fn threadAllocTest(boot_information: *const BootInformation) void {
|
|||||||
}
|
}
|
||||||
scheduler.setPriority(4);
|
scheduler.setPriority(4);
|
||||||
|
|
||||||
check("concurrent heap allocation stayed corruption-free (shared heap + per-aspace arena)", bufferHas(ok_marker) and !bufferHas(fail_marker));
|
check("concurrent heap allocation stayed corruption-free (shared heap + per-address-space arena)", bufferHas(ok_marker) and !bufferHas(fail_marker));
|
||||||
result();
|
result();
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -3289,20 +3289,20 @@ fn ioPassTest() void {
|
|||||||
log("DANOS-TEST-BEGIN: iopass\n", .{});
|
log("DANOS-TEST-BEGIN: iopass\n", .{});
|
||||||
const base_free = pmm.stats().free_frames;
|
const base_free = pmm.stats().free_frames;
|
||||||
|
|
||||||
const aspace = architecture.createAddressSpace() orelse {
|
const address_space = architecture.createAddressSpace() orelse {
|
||||||
check("created a fresh address space", false);
|
check("created a fresh address space", false);
|
||||||
result();
|
result();
|
||||||
return;
|
return;
|
||||||
};
|
};
|
||||||
const frame = pmm.alloc() orelse {
|
const frame = pmm.alloc() orelse {
|
||||||
architecture.destroyAddressSpace(aspace);
|
architecture.destroyAddressSpace(address_space);
|
||||||
check("allocated a frame to grant", false);
|
check("allocated a frame to grant", false);
|
||||||
result();
|
result();
|
||||||
return;
|
return;
|
||||||
};
|
};
|
||||||
// Map it the way mmio_map does (device grant, strong-uncacheable), then tear the space down.
|
// Map it the way mmio_map does (device grant, strong-uncacheable), then tear the space down.
|
||||||
architecture.mapUserDeviceInto(aspace, process.device_arena_base, frame, abi.page_size, false);
|
architecture.mapUserDeviceInto(address_space, process.device_arena_base, frame, abi.page_size, false);
|
||||||
architecture.destroyAddressSpace(aspace);
|
architecture.destroyAddressSpace(address_space);
|
||||||
|
|
||||||
// The page tables were reclaimed; the device-granted frame must not have been.
|
// The page tables were reclaimed; the device-granted frame must not have been.
|
||||||
check("device-granted frame survived teardown (not reclaimed as RAM)", pmm.stats().free_frames == base_free - 1);
|
check("device-granted frame survived teardown (not reclaimed as RAM)", pmm.stats().free_frames == base_free - 1);
|
||||||
@@ -3354,26 +3354,26 @@ fn displayTest(boot_information: *const BootInformation) void {
|
|||||||
// space, and confirm the leaf's cache type. We never run this space (no CR3 load) —
|
// space, and confirm the leaf's cache type. We never run this space (no CR3 load) —
|
||||||
// we only read back the page-table entries — so aliasing the same physical page at
|
// we only read back the page-table entries — so aliasing the same physical page at
|
||||||
// two cache types below is inert.
|
// two cache types below is inert.
|
||||||
const aspace = architecture.createAddressSpace() orelse {
|
const address_space = architecture.createAddressSpace() orelse {
|
||||||
check("created a fresh address space", false);
|
check("created a fresh address space", false);
|
||||||
result();
|
result();
|
||||||
return;
|
return;
|
||||||
};
|
};
|
||||||
defer architecture.destroyAddressSpace(aspace);
|
defer architecture.destroyAddressSpace(address_space);
|
||||||
|
|
||||||
const page_base = fb.base & ~@as(u64, abi.page_size - 1);
|
const page_base = fb.base & ~@as(u64, abi.page_size - 1);
|
||||||
architecture.mapUserDeviceInto(aspace, process.device_arena_base, page_base, abi.page_size, true);
|
architecture.mapUserDeviceInto(address_space, process.device_arena_base, page_base, abi.page_size, true);
|
||||||
check(
|
check(
|
||||||
"the framebuffer maps write-combining (PAT entry 4: PAT bit set, PCD/PWT clear)",
|
"the framebuffer maps write-combining (PAT entry 4: PAT bit set, PCD/PWT clear)",
|
||||||
architecture.userLeafIsWriteCombining(aspace, process.device_arena_base) == true,
|
architecture.userLeafIsWriteCombining(address_space, process.device_arena_base) == true,
|
||||||
);
|
);
|
||||||
|
|
||||||
// Regression guard: the strong-uncacheable default is still that, so WC is a real
|
// Regression guard: the strong-uncacheable default is still that, so WC is a real
|
||||||
// choice the flag makes, not the only behaviour.
|
// choice the flag makes, not the only behaviour.
|
||||||
architecture.mapUserDeviceInto(aspace, process.device_arena_base + abi.page_size, page_base, abi.page_size, false);
|
architecture.mapUserDeviceInto(address_space, process.device_arena_base + abi.page_size, page_base, abi.page_size, false);
|
||||||
check(
|
check(
|
||||||
"a register window still maps strong-uncacheable",
|
"a register window still maps strong-uncacheable",
|
||||||
architecture.userLeafIsWriteCombining(aspace, process.device_arena_base + abi.page_size) == false,
|
architecture.userLeafIsWriteCombining(address_space, process.device_arena_base + abi.page_size) == false,
|
||||||
);
|
);
|
||||||
|
|
||||||
log("display: mapped {d}x{d} pitch {d} (write-combining)\n", .{ fb.width, fb.height, fb.pitch });
|
log("display: mapped {d}x{d} pitch {d} (write-combining)\n", .{ fb.width, fb.height, fb.pitch });
|
||||||
|
|||||||
@@ -40,7 +40,7 @@ fn runSpawnMode() void {
|
|||||||
runtime.system.yield();
|
runtime.system.yield();
|
||||||
}
|
}
|
||||||
if (spawn_done.load(.acquire) == 1 and shared_value == sentinel) {
|
if (spawn_done.load(.acquire) == 1 and shared_value == sentinel) {
|
||||||
write("thread-test: child ran in shared aspace ok\n");
|
write("thread-test: child ran in shared address space ok\n");
|
||||||
} else {
|
} else {
|
||||||
write("thread-test: FAIL worker did not update shared memory\n");
|
write("thread-test: FAIL worker did not update shared memory\n");
|
||||||
}
|
}
|
||||||
@@ -349,7 +349,7 @@ fn runAllocMode() void {
|
|||||||
for (threads[0..alloc_threads]) |t| t.join();
|
for (threads[0..alloc_threads]) |t| t.join();
|
||||||
|
|
||||||
// Every thread must have completed all rounds with each block intact — proof the
|
// Every thread must have completed all rounds with each block intact — proof the
|
||||||
// shared heap and the per-aspace mmap arena are safe under concurrent allocation.
|
// shared heap and the per-address_space mmap arena are safe under concurrent allocation.
|
||||||
if (allocs_clean.load(.acquire) != alloc_threads) {
|
if (allocs_clean.load(.acquire) != alloc_threads) {
|
||||||
write("thread-alloc: FAIL corruption or OOM under concurrent allocation\n");
|
write("thread-alloc: FAIL corruption or OOM under concurrent allocation\n");
|
||||||
return;
|
return;
|
||||||
|
|||||||
+2
-2
@@ -295,8 +295,8 @@ CASES = [
|
|||||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
|
|
||||||
# docs/threading-plan.md M1: address-space refcount — spaces destroyed exactly
|
# docs/threading-plan.md M1: address-space refcount — spaces destroyed exactly
|
||||||
# once per process, no leak/double-free (the foundation shared-aspace threads need).
|
# once per process, no leak/double-free (the foundation shared-address-space threads need).
|
||||||
{"name": "aspace-refcount",
|
{"name": "address-space-refcount",
|
||||||
"timeout": 60,
|
"timeout": 60,
|
||||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
|
|||||||
Reference in New Issue
Block a user