kernel: user memory is reached only through a checked copy

A new user-memory module owns every kernel touch of a user buffer:
copyFromUser, the new copyToUser, and the resolve behind both. The walk
accumulates the U/S and writable bits down all four levels with the MMU's
own AND rule — folding a 2 MiB leaf in before it resolves and refusing a
1 GiB leaf outright — so a copy honours what ring 3 itself would be
allowed, closing the presence-only trust model the IPC layer carried since
bring-up. It then confirms the frame is physmap-backed, because that is how
the copy reaches it: an mmio_map'd BAR passes the permission walk and would
otherwise fault ring 0 on an alias the physmap never mapped, on the IPC path
as much as the new one.

The nine stragglers that dereferenced user pointers raw now route through
it, so a bad pointer returns -EFAULT where it used to fault the kernel.
The write direction restructures its callees around kernel bounce buffers:
scheduler and devices-broker enumerate from a slot cursor (a task exiting
between chunks can neither duplicate nor lose an entry), klog_read drains
the ring in chunks, and fs_node stages headers and names contiguously.
fs_resolve copies out before installing the endpoint handle, so a faulting
copy cannot strand a capability; its out-capacity bound no longer adds an
unbounded ring-3 length to the base, which wrapped and trapped the kernel's
own overflow check. debug_write reads the caller's message once.

Suite 107/107 (new user-memory case: seven bad pointers refused, each
paired with a sound call that must still succeed).
This commit is contained in:
Daniel Samson
2026-07-31 20:57:49 +01:00
parent c4f16a5448
commit 8d4a7cf240
16 changed files with 691 additions and 105 deletions
+21 -3
View File
@@ -132,13 +132,31 @@ fn record(node: *platform.Device, parent_id: u64) u64 {
}
/// Copy up to `out.len` device descriptors into `out`; returns the total count
/// available (which may exceed `out.len`).
/// available (which may exceed `out.len`). For kernel callers with a buffer big
/// enough to take the whole table in one go.
pub fn enumerate(out: []device_abi.DeviceDescriptor) usize {
const n = @min(count, out.len);
@memcpy(out[0..n], devices[0..n]);
_ = enumerateFrom(0, out);
return count;
}
/// How many devices the table holds — the total `device_enumerate` reports back
/// however few of them fit in the caller's buffer.
pub fn deviceCount() usize {
return count;
}
/// Copy up to `out.len` descriptors starting at table index `start`, returning how
/// many were filled (0 once `start` reaches the end). The chunked form: the
/// `device_enumerate` system call bounces the table out through a small kernel
/// buffer, one chunk at a time, because a descriptor is far too big to stage a
/// whole user-requested array of them on a 16 KiB kernel stack.
pub fn enumerateFrom(start: usize, out: []device_abi.DeviceDescriptor) usize {
if (start >= count) return 0;
const n = @min(count - start, out.len);
@memcpy(out[0..n], devices[start..][0..n]);
return n;
}
/// Take exclusive ownership of device `id` for task `owner`. Fails if the id is
/// out of range or already claimed.
pub fn claim(id: u64, owner: u32) bool {