M14b: DMA memory (dma_alloc / dma_free)

An HCD programs a bus-master engine: it needs a descriptor ring that is physically
contiguous, at a physical address it knows, uncacheable, and pinned. mmap gives none
of those. Add dma_alloc(len, flags) -> vaddr (rax), paddr (rdx) and dma_free(vaddr,
len): grant contiguous, zeroed, pinned, strong-uncacheable memory in a per-process DMA
arena (PML4[228]) and hand back both addresses.

Pieces: pmm.allocContiguous(count, max_phys) finds a run of contiguous free frames
below a cap (dma_below_4g for 32-bit engines); mapUserDmaInto maps them uncacheable
(PCD|PWT) but WITHOUT device_grant, so unlike an MMIO grant these frames are real RAM
and freeSubtree returns them on teardown — a driver that dies leaks nothing. dma_free
is bounded to the DMA arena so it can never unmap the caller's stack/heap/MMIO.
dma_write_combining is accepted but falls back to coherent (WC needs PAT programming).

Runtime: runtime.dma.alloc/free (a two-return-value stub, like replyWait). New `dma`
kernel test drives the mechanism directly — contiguity, the below-4G cap, coherent
mapping, and reclaim-on-teardown (no leak). The thin syscall wrappers follow the tested
mmap/mmio_map shape and land their first real use with the first DMA driver. Suite
38/38 plus host tests.
This commit is contained in:
Daniel Samson
2026-07-10 19:43:59 +01:00
parent e7c7e7b94c
commit 125a3b4993
11 changed files with 252 additions and 1 deletions
+28
View File
@@ -173,3 +173,31 @@ pub fn free(address: u64) void {
used_frames -= 1;
if (f < next_hint) next_hint = f;
}
/// Allocate `count` physically **contiguous** frames whose highest byte is below
/// `max_phys` (pass `~0` for no limit; use a real limit for DMA engines with 32-bit
/// addressing). Returns the physical base, or null if no free run of that size fits.
/// A DMA descriptor ring needs contiguity, a known physical address, and pinning —
/// none of which one-frame `alloc` gives. Linear scan for a run of clear bits: fine
/// for the small rings DMA needs; a buddy allocator is a later optimisation. The
/// frames are freed individually with `free`, so there is no bespoke free path.
pub fn allocContiguous(count: usize, max_phys: u64) ?u64 {
if (count == 0) return null;
const limit: usize = @intCast(@min(@as(u64, total_frames), max_phys / page_size));
var start: usize = 1; // frame 0 stays reserved as the "none" address
while (start + count <= limit) {
if (isUsed(start)) {
start += 1;
continue;
}
var run: usize = 0;
while (run < count and !isUsed(start + run)) : (run += 1) {}
if (run == count) {
for (0..count) |i| setUsed(start + i);
used_frames += count;
return @as(u64, start) * page_size;
}
start += run + 1; // the frame at start+run is used; skip past it
}
return null; // no contiguous run of `count` frames below max_phys
}