M2 step 5: drop the low half — a true higher-half kernel
paging.init now maps only the physmap, the framebuffer/LAPIC windows, and the kernel's own segments; the entire low canonical half is left to user space. Every higher-half PML4 entry is pre-created so a per-process address space can share the kernel half by copying PML4[256..512), with an assert against late top-half entries and a 4 GiB guard on pre-switch table frames. The AP trampoline's low identity page is now created transiently by arm() and unmapped by disarm(); startAp asserts the page-table root is 32-bit addressable. Docs (paging.md) updated. Suite 27/27; 4-core normal boot reaches /sbin/init. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
f57a73e8a1
commit
bb7597ea0b
+48
-13
@@ -17,20 +17,54 @@ the final 4 KiB page. Each entry holds a physical address plus flag bits —
|
||||
present, writable, and (bit 63) **no-execute**. danos maps everything with 4 KiB
|
||||
pages: precise, and the extra table memory is negligible against available RAM.
|
||||
|
||||
## Higher half: the address-space layout
|
||||
|
||||
danos is a **higher-half kernel**. The kernel is linked to run at
|
||||
`0xFFFF_FFFF_8000_0000` but loaded low (the linker script's `AT()` gives each
|
||||
segment a physical load address at 1 MiB up; the bootloader maps the high link
|
||||
address to the low load address in its bootstrap tables and jumps in). The entire
|
||||
**low canonical half is reserved for user space**; the kernel lives in the top half
|
||||
alongside a **physmap** — a straight window onto all of physical memory at
|
||||
`physmap_base + phys`. Wherever the kernel needs to touch a physical address (a
|
||||
page-table frame, an ACPI table, a device register), it adds that constant:
|
||||
`danos.physToVirt(phys)`. The layout constants live in `src/root.zig`:
|
||||
|
||||
| region | virtual base | PML4 slot |
|
||||
|--------|--------------|-----------|
|
||||
| user image + stack | `0x0000_7000_0000_0000` | 224 (low half) |
|
||||
| kernel heap | `0xFFFF_8000_0000_0000` | 256 |
|
||||
| physmap (all RAM + MMIO windows) | `0xFFFF_8800_0000_0000` + phys | 272 |
|
||||
| kernel image | `0xFFFF_FFFF_8000_0000` | 511 |
|
||||
|
||||
The bootloader builds temporary **bootstrap tables** (identity + a 4 GiB physmap +
|
||||
the high kernel) so it can switch CR3 and jump to the high entry; the kernel then
|
||||
builds its own precise tables below and abandons them. Because both use the same
|
||||
`physmap_base`, any physmap pointer minted before the switch stays valid after it.
|
||||
|
||||
## What gets mapped, and with what permissions
|
||||
|
||||
The address space is built in three passes (`init`):
|
||||
The address space is built in four passes (`init`):
|
||||
|
||||
1. **All RAM, identity-mapped RW + NX.** Every non-MMIO region from the
|
||||
[memory map](memory-map.md) is mapped virtual == physical, read-write and
|
||||
*non-executable*. Identity mapping keeps everything already running valid across
|
||||
the CR3 switch (the frame allocator addresses frames by physical address, page
|
||||
tables are reached the same way, the stack stays put).
|
||||
2. **The framebuffer and the Local APIC**, the device memory we actually touch,
|
||||
also RW + NX. Everything else — unbacked address space, other MMIO — is simply
|
||||
left unmapped, so a stray access faults instead of silently succeeding.
|
||||
3. **The kernel's own segments, overlaid with their true ELF permissions.** This is
|
||||
1. **All RAM in the physmap, RW + NX.** Every non-MMIO region from the
|
||||
[memory map](memory-map.md) is mapped at `physToVirt(phys)`, read-write and
|
||||
*non-executable*. There is **no low/identity mapping** — the low half is user
|
||||
space. (Frames the kernel touches while still building these tables are reached
|
||||
through the loader's bootstrap physmap, which covers the low 4 GiB; both the
|
||||
frame allocator and the table builder scan low-address-up, so those frames stay
|
||||
under that limit.)
|
||||
2. **The framebuffer and the Local APIC**, the device memory the kernel touches
|
||||
directly, as physmap windows (RW + NX). Other MMIO is mapped on demand by
|
||||
`mapMmio`, also into the physmap; everything else is left unmapped, so a stray
|
||||
access faults instead of silently succeeding.
|
||||
3. **The kernel's own segments, overlaid with their true ELF permissions**, at
|
||||
their high link addresses mapped to their low physical load addresses. This is
|
||||
the interesting part.
|
||||
4. **Every higher-half PML4 entry pre-created** (an empty PDPT where none exists
|
||||
yet). The kernel half is then a fixed set of top-level slots, so a per-process
|
||||
address space can share it by copying `PML4[256..512)` once — growth beneath
|
||||
those slots (heap, on-demand MMIO) propagates to every address space because
|
||||
they share the PDPTs. `init` asserts no new higher-half PML4 entry appears
|
||||
afterward.
|
||||
|
||||
### W^X from the ELF program headers
|
||||
|
||||
@@ -55,9 +89,10 @@ reserved bit and fault.
|
||||
|
||||
### The null guard
|
||||
|
||||
Page 0 is deliberately left unmapped. A null (or near-null) pointer dereference now
|
||||
takes a page fault instead of quietly reading or writing real memory — turning a
|
||||
whole class of silent bugs into an immediate, located crash.
|
||||
The whole low half is unmapped except for explicit user mappings, so page 0 (and
|
||||
every near-null address) is unmapped by construction. A null (or near-null) pointer
|
||||
dereference in the kernel takes a page fault instead of quietly reading or writing
|
||||
real memory — turning a whole class of silent bugs into an immediate, located crash.
|
||||
|
||||
## Switching on, and the on-demand API
|
||||
|
||||
|
||||
Reference in New Issue
Block a user