isolation M1: ring 3 + a real /sbin/init, end to end

Ring 3 works: user GDT descriptors (sysret-ready layout), TSS.rsp0,
U/S-bit user mappings (W^X preserved), an int 0x80 syscall gate with a
mutable trap frame, and a setjmp-style enter/exit path. /sbin/init is a
real freestanding Zig binary built from sbin/, shipped on the ESP,
loaded by the bootloader (BootInfo.init_base/len), validated and mapped
by an in-kernel user-ELF loader, and run at CPL 3 — syscalls: exit,
ping, write. Tests: user, user-pf (U/S isolation proof, error code
0x5), init. Suite 27/27.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Daniel Samson
2026-07-08 22:15:07 +01:00
co-authored by Claude Fable 5
parent 7501bd1703
commit 546dd44a2a
17 changed files with 838 additions and 25 deletions
+44 -3
View File
@@ -76,6 +76,45 @@ pub fn unmapPage(virt: u64) void {
paging.unmap(virt);
}
/// Map a page accessible from ring 3 (U/S bit at every level). The caller keeps
/// W^X: code read-only + executable, data writable + no-execute.
pub fn mapUserPage(virt: u64, phys: u64, writable: bool, executable: bool) void {
paging.mapUser(virt, phys, writable, executable);
}
// --- ring 3 entry/exit -----------------------------------------------------
/// Drop to ring 3 at `rip` on `rsp` (defined in isr.s). Saves the kernel context,
/// publishes the kernel stack pointer through `rsp0_slot` (this core's TSS.rsp0,
/// so ring-3 interrupts land on a good stack), builds an iretq frame with the
/// user selectors, and iretq's. "Returns" only when the user program triggers
/// the exit path (user_exit_to_kernel).
extern fn enter_user(rip: u64, rsp: u64, rsp0_slot: *align(4) u64) callconv(.c) void;
/// Abandon the in-flight ring-3 trap context and resume the kernel as if
/// `enter_user` had returned (defined in isr.s). Called by the exit syscall.
extern fn user_exit_to_kernel() callconv(.c) noreturn;
/// Run user code at `rip` with stack `rsp` on this core (`cpu` = the caller's CPU
/// index; the arch layer can't ask the scheduler). Returns after the user program
/// exits via syscall. Interrupts are disabled on return (the exit arrives through
/// an interrupt gate) — the caller re-enables.
pub fn enterUser(cpu: usize, rip: u64, rsp: u64) void {
enter_user(rip, rsp, tss.rsp0Ptr(cpu));
}
/// Never returns to the user program: unwind to the kernel context that called
/// `enterUser`. For the exit syscall's handler.
pub fn userExit() noreturn {
user_exit_to_kernel();
}
/// Register the handler for the ring-3 syscall gate (int 0x80, vector 128). The
/// handler may write the trap frame (e.g. rax as the return value).
pub fn setSyscallHandler(handler: *const fn (*idt.CpuState) void) void {
idt.setSyscallHandler(handler);
}
/// CR3 holds the physical address of the active top-level page table.
pub fn readCr3() u64 {
return asm volatile ("mov %%cr3, %[out]"
@@ -84,9 +123,11 @@ pub fn readCr3() u64 {
}
/// IA32_GS_BASE: the hidden base of the GS segment. We repurpose it as the per-CPU
/// data pointer (there's no user mode yet, so no `swapgs` dance — GS base is always
/// the running core's per-CPU block). Set once per core during bring-up, after the
/// GDT is loaded (loading a GS *selector* would otherwise clobber this base).
/// data pointer. Because it's always read back via `rdmsr` (never gs-relative
/// addressing), no `swapgs` dance is needed even with user mode: the MSR is
/// privileged, ring 3 can't touch it, and its value is unaffected by ring
/// transitions. Set once per core during bring-up, after the GDT is loaded
/// (loading a GS *selector* would otherwise clobber this base).
const ia32_gs_base = 0xC000_0101;
/// Publish this core's per-CPU data pointer so `cpuLocal` can retrieve it. Each