M14b: DMA memory (dma_alloc / dma_free)

An HCD programs a bus-master engine: it needs a descriptor ring that is physically
contiguous, at a physical address it knows, uncacheable, and pinned. mmap gives none
of those. Add dma_alloc(len, flags) -> vaddr (rax), paddr (rdx) and dma_free(vaddr,
len): grant contiguous, zeroed, pinned, strong-uncacheable memory in a per-process DMA
arena (PML4[228]) and hand back both addresses.

Pieces: pmm.allocContiguous(count, max_phys) finds a run of contiguous free frames
below a cap (dma_below_4g for 32-bit engines); mapUserDmaInto maps them uncacheable
(PCD|PWT) but WITHOUT device_grant, so unlike an MMIO grant these frames are real RAM
and freeSubtree returns them on teardown — a driver that dies leaks nothing. dma_free
is bounded to the DMA arena so it can never unmap the caller's stack/heap/MMIO.
dma_write_combining is accepted but falls back to coherent (WC needs PAT programming).

Runtime: runtime.dma.alloc/free (a two-return-value stub, like replyWait). New `dma`
kernel test drives the mechanism directly — contiguity, the below-4G cap, coherent
mapping, and reclaim-on-teardown (no leak). The thin syscall wrappers follow the tested
mmap/mmio_map shape and land their first real use with the first DMA driver. Suite
38/38 plus host tests.
This commit is contained in:
Daniel Samson
2026-07-10 19:43:59 +01:00
parent e7c7e7b94c
commit 125a3b4993
11 changed files with 252 additions and 1 deletions
+49
View File
@@ -84,6 +84,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
ipcCallTest();
} else if (eql(case, "ipc-cap")) {
capabilityTest();
} else if (eql(case, "dma")) {
dmaTest();
} else if (eql(case, "smp")) {
smpTest();
} else if (eql(case, "affinity")) {
@@ -935,6 +937,53 @@ fn capabilityTest() void {
result();
}
/// DMA memory (M14): the properties a bus-mastering driver needs — physically
/// contiguous, a known physical address, correct cacheability, pinned, and reclaimed
/// on teardown. Exercises the kernel mechanism directly (`pmm.allocContiguous` +
/// `mapUserDmaInto`); the `dma_alloc`/`dma_free` syscalls are thin wrappers over it,
/// following the tested `mmap`/`mmio_map` shape, and land their first real use with the
/// first DMA driver.
fn dmaTest() void {
log("DANOS-TEST-BEGIN: dma\n", .{});
const base_free = pmm.stats().free_frames;
// A contiguous run: aligned, and it consumed exactly that many frames.
const frames = 4;
const phys = pmm.allocContiguous(frames, ~@as(u64, 0)) orelse {
check("allocContiguous(4) succeeded", false);
result();
return;
};
check("contiguous run is page-aligned", phys % abi.page_size == 0);
check("contiguous run consumed 4 frames", pmm.stats().free_frames == base_free - frames);
// The below-4G cap is honoured (legacy 32-bit DMA engines).
const low = pmm.allocContiguous(2, @as(u64, 4) << 30) orelse 0;
check("below-4G run stays under 4 GiB", low != 0 and low + 2 * abi.page_size <= (@as(u64, 4) << 30));
// Map the run into a fresh address space as coherent DMA and translate each page
// back: the same physical run, in order — proving contiguity and the mapping.
const aspace = architecture.createAddressSpace().?;
architecture.mapUserDmaInto(aspace, process.dma_arena_base, phys, frames * abi.page_size);
var mapped_ok = true;
for (0..frames) |i| {
const va = process.dma_arena_base + i * abi.page_size;
const got = architecture.translate(aspace, va) orelse {
mapped_ok = false;
break;
};
if (got != phys + i * abi.page_size) mapped_ok = false;
}
check("DMA pages translate to the contiguous physical run", mapped_ok);
// Teardown must reclaim the DMA RAM (the leaves carry no device_grant, so
// freeSubtree frees them as ordinary frames) — a driver that just dies leaks none.
architecture.destroyAddressSpace(aspace);
for (0..2) |i| pmm.free(low + i * abi.page_size);
check("no frames leaked after DMA teardown", pmm.stats().free_frames == base_free);
result();
}
var proc_worker_run: bool = true;
var proc_worker_ran: bool = false;