docs: audit all 'what's next' sections against the code; fix stale comments
Verified every deferred item in the nine docs with a what's-next section and marked what has since landed (reaper, per-process CR3, higher-half kernel, kernel heap, contiguous frame alloc, RSDP capture, cap-passing, driver restart with backoff, PS/2 keyboard) while keeping the genuinely open items. Also corrects interrupts.md's claim that the keyboard skipped the IO-APIC, and updates scheduler.zig/heap.zig comments that predated the reaper and the big kernel lock.
This commit is contained in:
@@ -188,13 +188,15 @@ spinning in unrelated code — is the whole mechanism working end to end.
|
||||
- **Uncacheable MMIO**: device grants are mapped `PCD|PWT` (strong-uncacheable) for
|
||||
user drivers — see [paging.md](paging.md).
|
||||
|
||||
## What's next (not done here)
|
||||
## What's next (partly done since)
|
||||
|
||||
- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and port I/O is
|
||||
now available to ring 3 via the claim-gated `io_read`/`io_write` syscalls
|
||||
([drivers.md](drivers.md)) — so the first *input* device is unblocked; it just needs
|
||||
writing (claim the controller, `irq_bind` GSI 1, read scancodes from `0x60`).
|
||||
- **MSI-X**: `msi_bind` gives one per-device edge-triggered vector (M15); MSI-X's
|
||||
multi-vector table (many queues per device, e.g. NVMe) is the remaining extension.
|
||||
- **The LAPIC's own page** is still mapped writeback-cacheable like the rest of the
|
||||
identity map. QEMU tolerates it; real hardware wants it uncacheable.
|
||||
- **The keyboard** — done, exactly as sketched: the PS/2 bus driver
|
||||
(`system/drivers/ps2-bus/`) claims the port-mapped 8042 controller through the
|
||||
claim-gated `io_read`/`io_write` syscalls ([drivers.md](drivers.md)), binds
|
||||
IRQ 1 (and the aux mouse's IRQ 12), reads scancodes from `0x60`, and decodes
|
||||
them into HID events for the [input service](input.md).
|
||||
- **MSI-X** — still open: `msi_bind` gives one per-device edge-triggered vector
|
||||
(M15); MSI-X's multi-vector table (many queues per device, e.g. NVMe) is the
|
||||
remaining extension.
|
||||
- **The LAPIC's own page** — still mapped writeback-cacheable like the rest of
|
||||
the identity map. QEMU tolerates it; real hardware wants it uncacheable.
|
||||
|
||||
+15
-15
@@ -358,21 +358,21 @@ is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2
|
||||
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on
|
||||
the first DMA driver to protect and test against) and these smaller items:
|
||||
|
||||
- **Releasing a claim.** There is no `dev_release`, and `devices_broker` never drops a claim on
|
||||
exit — only IRQ bindings are released. A dead driver's device stays owned forever,
|
||||
which blocks restart.
|
||||
- **Unregistering children.** `device_register` only appends. A USB device that is
|
||||
unplugged cannot be removed, and a bus driver in a loop can exhaust the 64-entry
|
||||
table.
|
||||
- **Restart.** A supervisor that *spawns* drivers now exists — the device-manager starts
|
||||
them with `system_spawn` — but a supervisor that *restarts* them does not. A driver that
|
||||
dies should release its claim, have its device quiesced, and be respawned; today nothing
|
||||
notices the death. Some pieces (`releaseIrqs`, `device_grant` teardown, the claim table)
|
||||
exist, and `dev_release` (below) is the missing mechanism; the restart policy is the
|
||||
resilience track ([resilience.md](resilience.md)).
|
||||
- **Interrupt priority / threaded IRQ latency.** `notifyFromIsr` enqueues the woken
|
||||
driver but doesn't preempt (`wakeLocked` deliberately leaves that to the caller), so
|
||||
a woken driver waits for the next scheduling point.
|
||||
- **Releasing a claim** — half done. The kernel now drops *all* of a dead driver's
|
||||
claims on every path out of a process (`releaseAllOwnedBy`, called from process
|
||||
teardown), which unblocked restart. A voluntary `dev_release` for a live driver
|
||||
still doesn't exist.
|
||||
- **Unregistering children** — half done. Hot-remove works at the manager layer:
|
||||
the xHCI bus reports `child_removed` on unplug and the device manager prunes its
|
||||
tree. The kernel's own device table is still append-only, so a bus driver in a
|
||||
loop can still exhaust the 64-entry table.
|
||||
- **Restart** — done. The device manager notices a driver's death, reads its exit
|
||||
reason, prunes the children it reported, and respawns it with exponential
|
||||
backoff — with a crash-loop cap that marks a repeat offender `failed` instead
|
||||
of respawning forever ([device-manager.md](device-manager.md)).
|
||||
- **Interrupt priority / threaded IRQ latency** — still open. `notifyFromIsr`
|
||||
enqueues the woken driver but doesn't preempt (`wakeLocked` deliberately leaves
|
||||
that to the caller), so a woken driver waits for the next scheduling point.
|
||||
|
||||
## The driver contract (M17–M18)
|
||||
|
||||
|
||||
@@ -114,11 +114,13 @@ leaves the single region containing it `reserved`, so `init` won't hand it out.
|
||||
later step will move task 0 onto a kernel-owned stack, freeing that last ~1 MiB
|
||||
region too (and giving user mode the clean stack it wants).
|
||||
|
||||
## What's next (not done here)
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Contiguous allocation** — scan for N consecutive free bits — for callers that
|
||||
need physically adjacent frames.
|
||||
- **A kernel stack for task 0**, so the boot stack's region can be freed too (and
|
||||
for the clean stack user mode wants).
|
||||
- **Freeing the `reserved` `loader_data`** (the boot-time map buffers) once the
|
||||
kernel is done reading the memory map.
|
||||
- **Contiguous allocation** — done: `allocContiguous` scans for a run of clear
|
||||
bits, with an optional physical ceiling for DMA (`dma_alloc` is its user), and
|
||||
`allocBelow` serves the SMP trampoline.
|
||||
- **A kernel stack for task 0** — still open: the boot processor's idle task runs
|
||||
on the boot stack to this day, so that region can't be freed.
|
||||
- **Freeing the `reserved` `loader_data`** (the boot-time map buffers) — still
|
||||
open: the bitmap deliberately tracks those frames so they *can* be freed, but
|
||||
nothing frees them yet.
|
||||
|
||||
+9
-7
@@ -65,11 +65,13 @@ the proof that free and the free list actually work, not just alloc; "heap growt
|
||||
forces allocation past the initial page so `grow`/`map` runs; and the `ArrayList`
|
||||
check is the std-integration payoff.
|
||||
|
||||
## What's next (not done here)
|
||||
## What's next (largely still true)
|
||||
|
||||
- **Thread/interrupt safety.** The heap assumes a single caller — no lock yet.
|
||||
It's safe now (nothing allocates from interrupt handlers), but threads or an
|
||||
allocating IRQ handler will need a lock (or `cli` around the critical section).
|
||||
- **Larger alignments** than 16 (for page-aligned buffers, DMA regions).
|
||||
- **`resize`/`remap` in place**, so growing an `ArrayList` needn't always copy.
|
||||
- **Reclaiming empty tail pages** back to the frame allocator when the heap shrinks.
|
||||
- **Thread/interrupt safety** — overtaken by the big kernel lock: SMP arrived
|
||||
with a single kernel lock taken at every kernel entry, which serializes all
|
||||
heap access. The heap still has no lock of its own, and needs none unless the
|
||||
big lock is ever split.
|
||||
- **Larger alignments** than 16 — still unsupported; page-aligned and DMA
|
||||
buffers come straight from the frame allocator instead.
|
||||
- **`resize`/`remap` in place** — still not done; growing an `ArrayList` copies.
|
||||
- **Reclaiming empty tail pages** — still not done; the heap only ever grows.
|
||||
|
||||
+5
-4
@@ -129,10 +129,11 @@ Both items originally deferred here have landed:
|
||||
- **The IO-APIC**: [ioapic.zig](../system/kernel/architecture/x86_64/ioapic.zig)
|
||||
routes external device lines onto vectors — discovered via ACPI's MADT, every
|
||||
input masked at init, lines unmasked one at a time as user-space drivers bind
|
||||
them (see [device-interrupts.md](device-interrupts.md)). The keyboard turned
|
||||
out not to need it: danos's keyboard is USB HID over xHCI, which interrupts via
|
||||
MSI, not an ISA line. The IO-APIC path is still exercised — e.g. by the HPET's
|
||||
GSI routing.
|
||||
them (see [device-interrupts.md](device-interrupts.md)). The keyboard followed
|
||||
exactly as predicted: the PS/2 bus driver (`system/drivers/ps2-bus/`) claims
|
||||
the 8042 controller and binds its IRQ 1 (and the aux mouse's IRQ 12) through
|
||||
this routing. USB HID keyboards arrive over xHCI instead, which interrupts via
|
||||
MSI, and the HPET's GSI routing exercises the same path.
|
||||
- **SSE state**: `isr_common` (and the syscall entry) now `fxsave`/`fxrstor` the
|
||||
full SSE/x87 register file around dispatch. This stopped being optional the
|
||||
moment kernel code touched XMM — a 16-byte struct copy is a `movdqu` — and its
|
||||
|
||||
+13
-8
@@ -87,13 +87,15 @@ elsewhere is not lost.
|
||||
This is what makes a user-space driver possible at all, and it's the subject of
|
||||
[drivers.md](drivers.md).
|
||||
|
||||
## What's next (not done here)
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Priority inheritance** through IPC, so a high-priority client blocked on a
|
||||
low-priority server doesn't suffer unbounded priority inversion.
|
||||
- **Handle transfer.** A server can't hand a client a handle to a third endpoint, so
|
||||
every capability is either well-known (the registry) or inherited — there's no way
|
||||
to delegate one.
|
||||
- **Priority inheritance** through IPC — still open: a high-priority client
|
||||
blocked on a low-priority server suffers unbounded priority inversion.
|
||||
- **Handle transfer.** *Landed as cap-passing (M13)*: `ipc_call` and
|
||||
`ipc_reply_wait` carry an optional capability alongside the bytes (`send_cap`),
|
||||
copying an endpoint or shared-memory handle into the peer's table. First user:
|
||||
[input](input.md) subscribers register by handing over their own endpoint, and
|
||||
class drivers get a private channel to one device.
|
||||
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
|
||||
shape (logging, notifications between servers). *Landed as `ipc_send`* — a
|
||||
non-blocking post to an endpoint's bounded payload queue, delivered through
|
||||
@@ -101,8 +103,11 @@ This is what makes a user-space driver possible at all, and it's the subject of
|
||||
first used by, the [input service](input.md)'s keyboard-event broadcast, where a
|
||||
synchronous push would let one dead subscriber hang the fan-out. A full queue drops
|
||||
the oldest (discrete messages, not a coalescing level like the notification ring).
|
||||
- **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel
|
||||
lock; a bulk transfer wants shared pages, not a copy.
|
||||
- **A bounded reply** — half landed. The copy is still 256 bytes
|
||||
(`MESSAGE_MAXIMUM`) under the big kernel lock, but bulk transfer got its shared
|
||||
pages: `shared_memory_create`/`map`/`physical`, the region handle delegated as
|
||||
a capability (above). virtio-gpu's scanout surface is the first user
|
||||
([display-v2.md](display-v2.md)).
|
||||
|
||||
## Lifecycle conventions over IPC (M17)
|
||||
|
||||
|
||||
+6
-4
@@ -158,11 +158,13 @@ never knows the difference.
|
||||
This page is plumbing plus classification only. The map's first consumer, the
|
||||
**physical frame allocator**, is built directly on the `usable` regions here — which
|
||||
already include the reclaimed boot-services memory the loader folded in (see
|
||||
[frame-allocator.md](frame-allocator.md)). Still to come:
|
||||
[frame-allocator.md](frame-allocator.md)). Of the two items once listed here, one is done:
|
||||
|
||||
- Freeing the `reserved` `loader_data` (these boot-time buffers) once the kernel is
|
||||
done reading the map.
|
||||
- Freeing the `reserved` `loader_data` (these boot-time buffers) once the kernel
|
||||
is done reading the map — still open: the frame allocator's bitmap tracks those
|
||||
frames so they can be freed, but nothing frees them yet.
|
||||
- Capturing the ACPI RSDP from the UEFI configuration table before exit (the same
|
||||
"grab it before ExitBootServices" pattern), for when ACPI parsing arrives.
|
||||
"grab it before ExitBootServices" pattern) — done: the loader stows it in the
|
||||
boot handoff, and ACPI parsing consumes it from there ([acpi.md](acpi.md)).
|
||||
|
||||
See the roadmap in [efi.md](efi.md) for where this sits in the boot flow.
|
||||
|
||||
+13
-9
@@ -124,13 +124,17 @@ Four tests (see [testing.md](testing.md)) pin down the guarantees:
|
||||
> pointer to force a real hardware access. And `invlpg`, like `lgdt`, needs its
|
||||
> operand staged through a register in inline asm.
|
||||
|
||||
## What's next (not done here)
|
||||
## What's next (mostly done since)
|
||||
|
||||
- **A kernel heap** — the first real user of `map`, giving the kernel dynamic
|
||||
allocation. This is the natural next milestone.
|
||||
- **A higher-half kernel**: relink the kernel at a high virtual base so a future
|
||||
user address space can own the low half.
|
||||
- **Per-address-space tables** once there are user processes, and shared/copy-on-
|
||||
write mappings.
|
||||
- **Uncacheable MMIO**: the APIC/framebuffer pages are mapped writeback-cacheable;
|
||||
real hardware wants MMIO marked uncacheable.
|
||||
- **A kernel heap** — done, built on `map` exactly as anticipated
|
||||
([heap.md](heap.md)).
|
||||
- **A higher-half kernel** — done: the kernel is linked at
|
||||
`0xFFFFFFFF80000000` (`linker.ld`), loaded low and running high, and user
|
||||
processes own the low half.
|
||||
- **Per-address-space tables** — done: each user process gets its own root with
|
||||
the kernel half shared, and refcounted shared-memory mappings exist
|
||||
([ipc.md](ipc.md)). Copy-on-write remains unbuilt — nothing has needed it yet.
|
||||
- **Uncacheable MMIO** — half done: user-space device and DMA mappings are
|
||||
strong-uncacheable and the framebuffer is write-combining via the PAT, but the
|
||||
kernel's own `mapMmio` path is still writeback — the LAPIC included (see
|
||||
[device-interrupts.md](device-interrupts.md)).
|
||||
|
||||
+10
-9
@@ -136,13 +136,14 @@ Three tests (see [testing.md](testing.md)) prove the guarantees:
|
||||
- **`event`** blocks a task on a wait queue; waking it (from another task) resumes
|
||||
it, and since it's higher priority it preempts immediately.
|
||||
|
||||
## What's next (not done here)
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Priority inheritance.** Once tasks block on shared resources (locks, IPC),
|
||||
danos will need it to bound priority inversion — a [real-time](vision.md)
|
||||
requirement.
|
||||
- **Task exit / a reaper.** `exit` currently leaks the task's stack; nothing frees
|
||||
finished tasks' memory yet.
|
||||
- **Per-address-space tasks.** Today all tasks share the kernel address space. User
|
||||
processes will each get their own, switching page tables (CR3) on the context
|
||||
switch.
|
||||
- **Priority inheritance** — still open. Tasks now do block on shared resources
|
||||
(IPC rendezvous, the big kernel lock), and nothing yet bounds priority
|
||||
inversion — a [real-time](vision.md) requirement.
|
||||
- **Task exit / a reaper** — done. A dying task goes on its core's reap list in a
|
||||
`.reaping` state; the timer tick drains the list, frees the stack back to the
|
||||
heap, and recycles the task-table slot.
|
||||
- **Per-address-space tasks** — done. User processes each own an address space,
|
||||
and the context switch reloads CR3 when the target's tables differ (see
|
||||
[paging.md](paging.md)).
|
||||
|
||||
@@ -9,8 +9,9 @@
|
||||
//! of free blocks, split on allocation and coalesced with neighbours on free. It
|
||||
//! is exposed as a std.mem.Allocator, so the kernel can use std containers.
|
||||
//!
|
||||
//! Not yet concurrency-safe: it assumes a single caller and no allocation from
|
||||
//! interrupt handlers (ours don't). A lock comes with threads/SMP.
|
||||
//! No lock of its own: every kernel entry takes the big kernel lock (sync.zig),
|
||||
//! which serializes all heap access. A private lock only becomes necessary if
|
||||
//! the big lock is ever split.
|
||||
|
||||
const std = @import("std");
|
||||
const abi = @import("abi");
|
||||
|
||||
@@ -942,8 +942,9 @@ pub fn setPreemption(enabled: bool) void {
|
||||
preemption_enabled = enabled;
|
||||
}
|
||||
|
||||
/// End the current task and switch away for good; never returns. The task's stack
|
||||
/// is leaked for now (no reaper yet). Acquires the kernel lock and hands it off to
|
||||
/// End the current task and switch away for good; never returns. The task goes on
|
||||
/// its core's reap list; the tick-time reaper frees the stack and recycles the
|
||||
/// slot. Acquires the kernel lock and hands it off to
|
||||
/// the task we switch into (which releases it) — this frame never returns to leave.
|
||||
pub fn exit() noreturn {
|
||||
_ = sync.enter();
|
||||
@@ -965,7 +966,7 @@ pub fn exit() noreturn {
|
||||
/// dying task's kernel stack (in the shared kernel half, so it survives the CR3
|
||||
/// switch to the kernel tables that must happen before we free the process's own
|
||||
/// tables — we can't free the page tables we're standing on). The kernel stack
|
||||
/// itself is leaked, as in `exit` (no reaper yet). Never returns.
|
||||
/// itself is reaped later, as in `exit`. Never returns.
|
||||
pub fn exitUser() noreturn {
|
||||
_ = sync.enter();
|
||||
exitUserLocked();
|
||||
|
||||
Reference in New Issue
Block a user