kernel: user memory is reached only through a checked copy

A new user-memory module owns every kernel touch of a user buffer:
copyFromUser, the new copyToUser, and the resolve behind both. The walk
accumulates the U/S and writable bits down all four levels with the MMU's
own AND rule — folding a 2 MiB leaf in before it resolves and refusing a
1 GiB leaf outright — so a copy honours what ring 3 itself would be
allowed, closing the presence-only trust model the IPC layer carried since
bring-up. It then confirms the frame is physmap-backed, because that is how
the copy reaches it: an mmio_map'd BAR passes the permission walk and would
otherwise fault ring 0 on an alias the physmap never mapped, on the IPC path
as much as the new one.

The nine stragglers that dereferenced user pointers raw now route through
it, so a bad pointer returns -EFAULT where it used to fault the kernel.
The write direction restructures its callees around kernel bounce buffers:
scheduler and devices-broker enumerate from a slot cursor (a task exiting
between chunks can neither duplicate nor lose an entry), klog_read drains
the ring in chunks, and fs_node stages headers and names contiguously.
fs_resolve copies out before installing the endpoint handle, so a faulting
copy cannot strand a capability; its out-capacity bound no longer adds an
unbounded ring-3 length to the base, which wrapped and trapped the kernel's
own overflow check. debug_write reads the caller's message once.

Suite 107/107 (new user-memory case: seven bad pointers refused, each
paired with a sound call that must still succeed).
This commit is contained in:
Daniel Samson
2026-07-31 20:57:49 +01:00
parent c4f16a5448
commit 8d4a7cf240
16 changed files with 691 additions and 105 deletions
+61 -24
View File
@@ -1203,35 +1203,72 @@ pub fn destroyTaskLocked(t: *Task) void {
/// Snapshot the task table into `out` (up to its length), returning the total
/// number of live tasks — the kernel half of `process_enumerate`, mirroring
/// devices_broker.enumerate. Kernel tasks are included (empty name, supervisor 0):
/// an honest `ps` shows the idle tasks too. `out` may be user memory: the caller's
/// address space is loaded during its system call, and the same bring-up trust
/// applies as for device_enumerate (an unmapped user page faults the kernel).
/// an honest `ps` shows the idle tasks too. `out` is always KERNEL memory: the
/// system call bounces it out to the caller through the checked copy layer
/// (system/kernel/user-memory.zig), so a bad user pointer fails the call instead
/// of faulting ring 0.
pub fn enumerate(out: []abi.ProcessDescriptor) u64 {
var total: u64 = 0;
var cursor: usize = 0;
while (true) {
// Past the buffer, keep walking with an empty chunk: the total is the
// whole live count, however few descriptors the caller had room for.
const room = if (total < out.len) out[@intCast(total)..] else out[out.len..];
const chunk = enumerateFrom(&cursor, room);
total += chunk.live;
if (chunk.done) return total;
}
}
/// What one chunk of the task-table walk found.
pub const TaskChunk = struct {
/// Live tasks passed in this chunk, whether or not they fit in `out` — this
/// is what the running total (and hence `process_enumerate`'s result) counts.
live: usize,
/// How many of those were written into `out` (`@min(live, out.len)`).
filled: usize,
/// The cursor reached the end of the table: this was the last chunk.
done: bool,
};
/// One chunk of the task table: starting at slot `cursor` (advanced past
/// everything scanned), describe up to `out.len` live tasks into `out` — or, with
/// an empty `out`, just count the rest. `cursor == tasks.len` ends the walk.
///
/// The chunked form exists so `process_enumerate` can stage each chunk in a small
/// kernel buffer and copy it out with `user_memory.copyToUser`, rather than
/// handing a user pointer to the kernel's own stores. A *slot* cursor, rather
/// than a "skip the first N live tasks" count, keeps chunks from duplicating or
/// losing an entry when a task exits between them.
pub fn enumerateFrom(cursor: *usize, out: []abi.ProcessDescriptor) TaskChunk {
const flags = sync.enter();
defer sync.leave(flags);
var total: u64 = 0;
for (&tasks) |*t| {
var live: usize = 0;
var filled: usize = 0;
const limit = if (out.len == 0) tasks.len else out.len; // always makes progress
while (cursor.* < tasks.len and live < limit) {
const t = &tasks[cursor.*];
cursor.* += 1;
if (t.state == .free or t.state == .reaping) continue; // reaping = already exited
if (total < out.len) {
const d = &out[total];
d.* = .{
.id = t.id,
.supervisor = t.supervisor,
.leader = t.leader,
.state = @intFromEnum(@as(abi.ProcessState, switch (t.state) {
.ready => .ready,
.running => .running,
.blocked => .blocked,
.free, .reaping => unreachable,
})),
.priority = t.priority,
.name_length = t.name_length,
.name = t.name_buffer,
};
}
total += 1;
live += 1;
if (filled == out.len) continue;
out[filled] = .{
.id = t.id,
.supervisor = t.supervisor,
.leader = t.leader,
.state = @intFromEnum(@as(abi.ProcessState, switch (t.state) {
.ready => .ready,
.running => .running,
.blocked => .blocked,
.free, .reaping => unreachable,
})),
.priority = t.priority,
.name_length = t.name_length,
.name = t.name_buffer,
};
filled += 1;
}
return total;
return .{ .live = live, .filled = filled, .done = cursor.* >= tasks.len };
}
/// Whether the running task is a user process (has its own address space).