diff --git a/docs/danos-file-system-hierarchy-FSH.md b/docs/danos-file-system-hierarchy-FSH.md index 3478036..798e371 100644 --- a/docs/danos-file-system-hierarchy-FSH.md +++ b/docs/danos-file-system-hierarchy-FSH.md @@ -64,14 +64,15 @@ already provide — it claims its device, maps its registers with `mmio_map`, an on `replyWait` for either an interrupt or a client request. `system/drivers/hpet/hpet.zig` is already that program, minus the client half. -The obstacle is not the file type, it is which hardware a ring-3 driver can actually -drive. Port I/O is unavailable to user space — the TSS I/O permission bitmap is absent -and IOPL is never raised — so `in`/`out` from a driver is a #GP. That excludes the -16550 UART at `0x3F8` and PS/2 at `0x60`/`0x64`, which is to say it excludes the -obvious implementations of `/dev/tty`, `/dev/ttyS0` and a keyboard node. Until either -port I/O grants or a memory-mapped UART exist, serial output stays a kernel service -reached through the `write` system call rather than a file. A memory-mapped device such -as the framebuffer has no such problem and is the more likely first real entry here. +The obstacle was never the file type; it is which hardware a ring-3 driver can reach. +Direct `in`/`out` from user space is still a #GP (no TSS I/O bitmap, IOPL never raised), +but a driver no longer needs it: **`io_read`/`io_write`** grant port access the same way +`mmio_map` grants memory — gated by `device_claim` and the device's discovered `io_port` +resource. So the 16550 UART at `0x3F8` and the PS/2 controller at `0x60`/`0x64` (and thus +`/dev/ttyS0` and a keyboard node) are now writable as ordinary ring-3 drivers; the +low-rate legacy hardware that needs port I/O is fine with a syscall per access. A +memory-mapped device such as the framebuffer, needing no port I/O at all, remains the +easiest first entry. ### Block devices @@ -79,20 +80,23 @@ A block device is addressed in fixed-size blocks and, unlike a character device, layer above is free to buffer, reorder, coalesce and retry requests against it. Disks and other persistent storage are the whole population of this class. -**danos cannot host a block driver at all today,** and the reason is worth stating -plainly because it is not a matter of unwritten code. Every storage controller worth -naming is a bus master: it is programmed by handing it the physical address of a -descriptor ring and left to read and write memory on its own. A ring-3 driver cannot -build such a ring, because `mmap` returns writeback-cached, physically discontiguous -pages and never discloses their physical address. Nor should it be allowed to: a device -programmed with an arbitrary physical address writes to arbitrary physical memory, and -page tables do not sit between a device and RAM — an IOMMU does. Granting a DMA-capable -device to a driver process, with no IOMMU programmed, is equivalent to granting ring 0, -which would forfeit the isolation that motivates user-space drivers in the first place. +A block driver is now **writable, but not yet memory-safe.** Every storage controller +worth naming is a bus master: it is programmed by handing it the physical address of a +descriptor ring and left to read and write memory on its own. That ring is exactly what +**`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its +physical address disclosed — and **`/lib/mmio`**'s barriers order the descriptor writes +against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver +can be written today (the M14/M15 work in [driver-model.md](driver-model.md); the earlier +"cannot host a block driver at all" is no longer true). -Block devices therefore wait on DMA-capable memory, memory barriers, and VT-d/DMAR — -the M14–M16 work in [driver-model.md](driver-model.md). A ramdisk over the initial ramdisk is -the one block-shaped thing implementable now, and it needs no driver process. +What is *not* yet true is that it is safe. A device programmed with an arbitrary physical +address writes to arbitrary physical memory, and page tables do not sit between a device +and RAM — an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains +are programmed, so granting a DMA-capable device to a driver process is still equivalent +to granting ring 0. Until per-device domains confine a driver's DMA to the buffers it +`dma_alloc`'d, a block driver works but forfeits the isolation that motivates user-space +drivers — enforcement is the next step, and lands with that first driver. A ramdisk over +the initial ramdisk remains the one block-shaped thing that needs no driver process at all. ### Pseudo-devices diff --git a/docs/device-interrupts.md b/docs/device-interrupts.md index 053f872..fe71a7b 100644 --- a/docs/device-interrupts.md +++ b/docs/device-interrupts.md @@ -156,10 +156,11 @@ spinning in unrelated code — is the whole mechanism working end to end. ## What's next (not done here) -- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and ring 3 - has no port I/O yet, so the first *input* device is blocked on either an I/O - permission bitmap or `io_in`/`io_out` syscalls ([drivers.md](drivers.md)). -- **MSI/MSI-X**: per-device vectors, edge-triggered and unshared, which retire the - I/O APIC's mask/ack cycle and its 24-GSI ceiling. +- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and port I/O is + now available to ring 3 via the claim-gated `io_read`/`io_write` syscalls + ([drivers.md](drivers.md)) — so the first *input* device is unblocked; it just needs + writing (claim the controller, `irq_bind` GSI 1, read scancodes from `0x60`). +- **MSI-X**: `msi_bind` gives one per-device edge-triggered vector (M15); MSI-X's + multi-vector table (many queues per device, e.g. NVMe) is the remaining extension. - **The LAPIC's own page** is still mapped writeback-cacheable like the rest of the identity map. QEMU tolerates it; real hardware wants it uncacheable. diff --git a/docs/driver-model.md b/docs/driver-model.md index 33d5a78..180d980 100644 --- a/docs/driver-model.md +++ b/docs/driver-model.md @@ -152,6 +152,12 @@ If a class driver needs `mmio`, it has become an HCD and should be one. ack cycle. Legacy INTx (`_PRT` parsing + shared lines) is deliberately skipped — MSI is the real answer. QEMU's HPET has no MSI, so delivery is proven with a self-IPI; the first PCI driver is the first real consumer. +- **Port I/O** — `io_read`/`io_write(device_id, resource_index, offset, width[, value])`: + a claimed device's `io_port` resource lets a driver read/write its ports, gated exactly + like `mmio_map` gates memory (direct ring-3 `in`/`out` stays a #GP). This is what makes + a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with + a syscall per access. `io_port` resources were recorded by discovery and ignored — now + they're used. - **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table, maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the platform info). This is detection only — **no translation domains are programmed, so diff --git a/docs/drivers.md b/docs/drivers.md index 2578f94..2c72844 100644 --- a/docs/drivers.md +++ b/docs/drivers.md @@ -266,26 +266,24 @@ controller drivers fit together. Worth knowing before you write the second driver: -- **Ring 3 has no port I/O.** The TSS I/O permission bitmap is absent - (`tss.zig`: `iomap_base = @sizeOf(Tss)`), and IOPL is never raised, so `in`/`out` - from a driver is a #GP. That rules out a user-space 16550 UART (`0x3F8`), PS/2 - (`0x60`/`0x64`), and legacy PCI config (`0xCF8`/`0xCFC`). Everything must be MMIO. - `io_port` resources are recorded by discovery and then ignored. +Several things this list used to warn about are now available (see +[driver-model.md](driver-model.md)): **port I/O** (`io_read`/`io_write`, claim-gated by +the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so a PS/2 or +16550 driver goes through these), **DMA memory** (`dma_alloc`: contiguous, pinned, +uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s +`mb`/`rmb`/`wmb`). What remains: + - **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means granting one grants the other. A `device_register`ed child's *resource* can be narrower than a page, but its *mapping* can't. -- **No DMA memory.** `mmap` gives you writeback-cached, non-contiguous pages and never - tells you their physical address, so you cannot build a descriptor ring. Any driver - for a bus-mastering device is blocked on this. -- **No memory barriers.** There are none in the tree, and `volatile` is not one — it - won't stop the compiler sinking an ordinary store (your DMA descriptor) past a - volatile MMIO store (your doorbell). On x86 you mostly get away with it; on ARM you - will not. See [driver-model.md](driver-model.md#m14). - **DMA is not contained.** A driver that can program a bus-mastering device can make that device write to *any* physical address — page tables don't sit between a device - and RAM; an IOMMU does. Until VT-d/DMAR is programmed, `device_claim` on a DMA-capable - device is effectively equivalent to granting ring 0. This is the largest gap between - the design's promise and what it delivers. + and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains + are programmed, so `device_claim` on a DMA-capable device is still effectively + equivalent to granting ring 0. This is the largest gap between the design's promise and + what it delivers; enforcement lands with the first DMA driver. +- **No `dev_release`.** A claim is never dropped (only IRQ/MSI bindings are, on exit), so + a device stays owned for the life of its driver — which blocks restart. - **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`). @@ -341,14 +339,12 @@ Two companions cover what `hpet` can't, because it never exits: ## What's next (not done here) -The big ones — capability passing (class drivers), DMA + barriers and MSI (host -controller drivers), and the IOMMU — have proposed signatures in -[driver-model.md](driver-model.md). Smaller items: +The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI, +and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as +is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550 +driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on +the first DMA driver to protect and test against) and these smaller items: -- **Port I/O grants**, so a PS/2 or 16550 driver is possible: either a per-device TSS - I/O permission bitmap swapped on context switch, or `io_in`/`io_out` syscalls gated - by the same claim. The legacy devices that need it are all low-rate, so the syscall - is likely fast enough. - **Releasing a claim.** There is no `dev_release`, and `devices_broker` never drops a claim on exit — only IRQ bindings are released. A dead driver's device stays owned forever, which blocks restart. diff --git a/library/runtime/device.zig b/library/runtime/device.zig index 528dcb1..6502314 100644 --- a/library/runtime/device.zig +++ b/library/runtime/device.zig @@ -90,3 +90,21 @@ pub fn msiBind(device_id: u64, endpoint: usize) ?Msi { if (failed(rax)) return null; return .{ .address = rax, .data = @intCast(rdx) }; } + +/// Read `width` bytes (1, 2, or 4) from a port in a claimed device's `io_port` +/// resource, at byte `offset` within it. Ring 3 has no direct `in`/`out`, so a legacy +/// driver (PS/2, 16550 UART) reaches its ports through this claim-gated call — each +/// access is a syscall, which is fine for the low-rate hardware that needs it. Returns +/// null if the capability check fails (device not claimed, wrong resource, out of +/// range). A device that decodes no data returns all-ones, which is a valid value, not +/// a failure. +pub fn ioRead(device_id: u64, resource_index: u64, offset: u64, width: u8) ?u32 { + const r = sc.systemCall4(.io_read, device_id, resource_index, offset, width); + return if (failed(r)) null else @intCast(r); +} + +/// Write `value` (its low `width` bytes, 1/2/4) to a port in a claimed device's +/// `io_port` resource, at byte `offset`. Same capability gate as `ioRead`. +pub fn ioWrite(device_id: u64, resource_index: u64, offset: u64, width: u8, value: u32) bool { + return !failed(sc.systemCall5(.io_write, device_id, resource_index, offset, width, value)); +} diff --git a/system/abi.zig b/system/abi.zig index 29816be..939a981 100644 --- a/system/abi.zig +++ b/system/abi.zig @@ -40,6 +40,8 @@ pub const SystemCall = enum(u64) { dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device + io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource + io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource _, }; diff --git a/system/kernel/process.zig b/system/kernel/process.zig index 5a47cac..222221e 100644 --- a/system/kernel/process.zig +++ b/system/kernel/process.zig @@ -160,6 +160,8 @@ fn system_call(state: *architecture.CpuState) void { .dma_alloc => systemDmaAlloc(state), .dma_free => systemDmaFree(state), .msi_bind => systemMsiBind(state), + .io_read => systemIoRead(state), + .io_write => systemIoWrite(state), _ => fail(state), } } @@ -271,6 +273,49 @@ fn systemMmioMap(state: *architecture.CpuState) void { architecture.setSystemCallResult(state, base_v + (r.start & (page_size - 1))); // register base } +/// Resolve a port-I/O access against the caller's claims. The device must be claimed by +/// `t`, `resource_index` must name one of its `io_port` resources, and the access +/// `[offset, offset+width)` must fall wholly inside it. Returns the absolute 16-bit +/// port, or null if the capability check fails. The claim plus the discovered `io_port` +/// resource are the capability — exactly like `mmio_map` for memory, so a driver can +/// only touch the ports its device actually owns, never a raw `in`/`out` to anywhere. +pub fn resolveIoPort(t: *scheduler.Task, device_id: u64, resource_index: u64, offset: u64, width: u64) ?u16 { + if (width != 1 and width != 2 and width != 4) return null; + const owner = devices_broker.ownerOf(device_id) orelse return null; + if (owner != t.id) return null; // not claimed by this process + const r = devices_broker.resourceOf(device_id, resource_index) orelse return null; + if (r.kind != @intFromEnum(device_abi.ResourceKind.io_port)) return null; + if (offset + width > r.len) return null; // access escapes the claimed port range + const port = r.start + offset; + if (port + width > 0x1_0000) return null; // I/O ports are 16-bit + return @intCast(port); +} + +/// io_read(device_id, resource_index, offset, width) -> value: read `width` bytes (1/2/4) +/// from a port in a claimed device's `io_port` resource. Ring 3 has no direct `in`/`out` +/// (no TSS I/O bitmap, IOPL never raised), so a legacy driver (PS/2, 16550 UART) reaches +/// its ports through this claim-gated call — low-rate hardware, so a syscall per access +/// is fine. See docs/drivers.md. +fn systemIoRead(state: *architecture.CpuState) void { + const t = scheduler.current(); + if (t.aspace == 0) return fail(state); + const width = architecture.systemCallArg(state, 3); + const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state); + architecture.setSystemCallResult(state, architecture.pioRead(@intCast(width), port)); +} + +/// io_write(device_id, resource_index, offset, width, value) -> 0: write `value` (low +/// `width` bytes) to a port in a claimed device's `io_port` resource. Same capability +/// gate as `io_read`. +fn systemIoWrite(state: *architecture.CpuState) void { + const t = scheduler.current(); + if (t.aspace == 0) return fail(state); + const width = architecture.systemCallArg(state, 3); + const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state); + architecture.pioWrite(@intCast(width), port, @intCast(architecture.systemCallArg(state, 4))); + architecture.setSystemCallResult(state, 0); +} + /// dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): grant `len` bytes (rounded up to /// whole pages) of DMA-capable memory — physically contiguous, zeroed, pinned, and /// strong-uncacheable (coherent) — mapping it into the caller's DMA arena and handing diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index cb2a624..1d7f909 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -90,6 +90,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void { msiTest(); } else if (eql(case, "iommu")) { iommuTest(); + } else if (eql(case, "ioport")) { + ioPortTest(); } else if (eql(case, "smp")) { smpTest(); } else if (eql(case, "affinity")) { @@ -1051,6 +1053,53 @@ fn iommuTest() void { result(); } +/// Port I/O grants: ring 3 has no `in`/`out`, so a legacy driver reaches its ports +/// through `io_read`/`io_write`, gated by `device_claim` and the device's `io_port` +/// resource exactly like `mmio_map` gates memory. Target the PS/2 controller's status +/// port (0x64) — discovered on every PC and side-effect-free to read. Proves the +/// capability gate (`resolveIoPort` admits an in-range access, refuses out-of-range, +/// over-wide, and unclaimed) and that the kernel actually performs the `in`. The +/// `io_read`/`io_write` syscalls wrap this with the same ring-3 dispatch every device +/// driver already uses. +fn ioPortTest() void { + log("DANOS-TEST-BEGIN: ioport\n", .{}); + var buffer: [64]device_abi.DeviceDescriptor = undefined; + const n = @min(devices_broker.enumerate(&buffer), buffer.len); + + var found_id: ?u64 = null; + var found_res: u64 = 0; + outer: for (buffer[0..n]) |d| { + for (0..d.resource_count) |ri| { + const r = d.resources[ri]; + if (r.kind == @intFromEnum(device_abi.ResourceKind.io_port) and r.start == 0x64 and r.len >= 1) { + found_id = d.id; + found_res = ri; + break :outer; + } + } + } + const id = found_id orelse { + check("discovered the PS/2 status port (io_port 0x64)", false); + result(); + return; + }; + check("discovered the PS/2 status port (io_port 0x64)", true); + + const me = scheduler.current(); + check("claimed the io_port device", devices_broker.claim(id, me.id)); + check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0, 1) == 0x64); + check("an over-wide access is refused", process.resolveIoPort(me, id, found_res, 0, 2) == null); + check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 1, 1) == null); + check("an unclaimed device id is refused", process.resolveIoPort(me, 0xDEAD_BEEF, found_res, 0, 1) == null); + + // The kernel actually issues the `in`. Reaching this line at all proves it didn't + // fault; a width-1 read must return a single byte. + const status = architecture.pioRead(1, 0x64); + check("reading the PS/2 status port returned a byte", status <= 0xFF); + log("DANOS-IOPORT: PS/2 status = 0x{x}\n", .{status}); + result(); +} + var proc_worker_run: bool = true; var proc_worker_ran: bool = false; diff --git a/test/qemu_test.py b/test/qemu_test.py index 398ce63..2587fb8 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -146,6 +146,11 @@ CASES = [ "qemu_extra": ["-device", "intel-iommu,intremap=off"], "expect": r"DANOS-TEST-RESULT: PASS", "fail": r"DANOS-TEST-RESULT: FAIL"}, + # Port I/O grants: a claimed device's io_port resource lets a driver read/write its + # ports (PS/2 status 0x64), gated by the claim; out-of-range/unclaimed is refused. + {"name": "ioport", + "expect": r"DANOS-TEST-RESULT: PASS", + "fail": r"DANOS-TEST-RESULT: FAIL"}, # Parallelism: needs more than one core, so this case boots with -smp 4. {"name": "smp", "smp": 4,