Port I/O grants (io_read / io_write) + refresh stale driver docs

Ring 3 still has no direct in/out (no TSS I/O bitmap, IOPL never raised — a #GP), but a
driver no longer needs it: io_read(device_id, resource_index, offset, width) and
io_write(..., value) grant port access the same way mmio_map grants memory. The claim
plus the device's discovered io_port resource are the capability — resolveIoPort checks
the device is claimed by the caller, the resource is io_port, and [offset, offset+width)
stays inside it, then issues the in/out via architecture.pioRead/pioWrite. So a PS/2 or
16550 driver is now writable; the low-rate legacy hardware that needs port I/O is fine
with a syscall per access. io_port resources were recorded by discovery and ignored —
now they're used. Runtime: device.ioRead/ioWrite.

New `ioport` test claims QEMU's PS/2 controller (io_port 0x64, discovered via ACPI) and
checks the gate admits an in-range access, refuses over-wide / out-of-range / unclaimed,
and that the kernel actually reads the status port (0x1c). Suite 41/41 plus host tests.

Docs refreshed for the whole M13-M16 + port-I/O reality: drivers.md "Limits" no longer
lists port I/O, DMA memory, or barriers as missing (they exist) and its "what's next"
reflects that; the FSH doc's character-device and block-device sections are corrected
("cannot host a block driver at all" is no longer true — writable now, not yet memory-
safe pending IOMMU enforcement); device-interrupts.md unblocks the keyboard and narrows
the interrupt gap to MSI-X.
This commit is contained in:
Daniel Samson
2026-07-10 20:36:14 +01:00
parent e612d948d2
commit 21657943a8
9 changed files with 174 additions and 48 deletions
+25 -21
View File
@@ -64,14 +64,15 @@ already provide — it claims its device, maps its registers with `mmio_map`, an
on `replyWait` for either an interrupt or a client request. `system/drivers/hpet/hpet.zig` is already on `replyWait` for either an interrupt or a client request. `system/drivers/hpet/hpet.zig` is already
that program, minus the client half. that program, minus the client half.
The obstacle is not the file type, it is which hardware a ring-3 driver can actually The obstacle was never the file type; it is which hardware a ring-3 driver can reach.
drive. Port I/O is unavailable to user space — the TSS I/O permission bitmap is absent Direct `in`/`out` from user space is still a #GP (no TSS I/O bitmap, IOPL never raised),
and IOPL is never raised — so `in`/`out` from a driver is a #GP. That excludes the but a driver no longer needs it: **`io_read`/`io_write`** grant port access the same way
16550 UART at `0x3F8` and PS/2 at `0x60`/`0x64`, which is to say it excludes the `mmio_map` grants memory — gated by `device_claim` and the device's discovered `io_port`
obvious implementations of `/dev/tty`, `/dev/ttyS0` and a keyboard node. Until either resource. So the 16550 UART at `0x3F8` and the PS/2 controller at `0x60`/`0x64` (and thus
port I/O grants or a memory-mapped UART exist, serial output stays a kernel service `/dev/ttyS0` and a keyboard node) are now writable as ordinary ring-3 drivers; the
reached through the `write` system call rather than a file. A memory-mapped device such low-rate legacy hardware that needs port I/O is fine with a syscall per access. A
as the framebuffer has no such problem and is the more likely first real entry here. memory-mapped device such as the framebuffer, needing no port I/O at all, remains the
easiest first entry.
### Block devices ### Block devices
@@ -79,20 +80,23 @@ A block device is addressed in fixed-size blocks and, unlike a character device,
layer above is free to buffer, reorder, coalesce and retry requests against it. Disks layer above is free to buffer, reorder, coalesce and retry requests against it. Disks
and other persistent storage are the whole population of this class. and other persistent storage are the whole population of this class.
**danos cannot host a block driver at all today,** and the reason is worth stating A block driver is now **writable, but not yet memory-safe.** Every storage controller
plainly because it is not a matter of unwritten code. Every storage controller worth worth naming is a bus master: it is programmed by handing it the physical address of a
naming is a bus master: it is programmed by handing it the physical address of a descriptor ring and left to read and write memory on its own. That ring is exactly what
descriptor ring and left to read and write memory on its own. A ring-3 driver cannot **`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its
build such a ring, because `mmap` returns writeback-cached, physically discontiguous physical address disclosed — and **`/lib/mmio`**'s barriers order the descriptor writes
pages and never discloses their physical address. Nor should it be allowed to: a device against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver
programmed with an arbitrary physical address writes to arbitrary physical memory, and can be written today (the M14/M15 work in [driver-model.md](driver-model.md); the earlier
page tables do not sit between a device and RAM — an IOMMU does. Granting a DMA-capable "cannot host a block driver at all" is no longer true).
device to a driver process, with no IOMMU programmed, is equivalent to granting ring 0,
which would forfeit the isolation that motivates user-space drivers in the first place.
Block devices therefore wait on DMA-capable memory, memory barriers, and VT-d/DMAR — What is *not* yet true is that it is safe. A device programmed with an arbitrary physical
the M14–M16 work in [driver-model.md](driver-model.md). A ramdisk over the initial ramdisk is address writes to arbitrary physical memory, and page tables do not sit between a device
the one block-shaped thing implementable now, and it needs no driver process. and RAM — an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
are programmed, so granting a DMA-capable device to a driver process is still equivalent
to granting ring 0. Until per-device domains confine a driver's DMA to the buffers it
`dma_alloc`'d, a block driver works but forfeits the isolation that motivates user-space
drivers — enforcement is the next step, and lands with that first driver. A ramdisk over
the initial ramdisk remains the one block-shaped thing that needs no driver process at all.
### Pseudo-devices ### Pseudo-devices
+6 -5
View File
@@ -156,10 +156,11 @@ spinning in unrelated code — is the whole mechanism working end to end.
## What's next (not done here) ## What's next (not done here)
- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and ring 3 - **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and port I/O is
has no port I/O yet, so the first *input* device is blocked on either an I/O now available to ring 3 via the claim-gated `io_read`/`io_write` syscalls
permission bitmap or `io_in`/`io_out` syscalls ([drivers.md](drivers.md)). ([drivers.md](drivers.md)) — so the first *input* device is unblocked; it just needs
- **MSI/MSI-X**: per-device vectors, edge-triggered and unshared, which retire the writing (claim the controller, `irq_bind` GSI 1, read scancodes from `0x60`).
I/O APIC's mask/ack cycle and its 24-GSI ceiling. - **MSI-X**: `msi_bind` gives one per-device edge-triggered vector (M15); MSI-X's
multi-vector table (many queues per device, e.g. NVMe) is the remaining extension.
- **The LAPIC's own page** is still mapped writeback-cacheable like the rest of the - **The LAPIC's own page** is still mapped writeback-cacheable like the rest of the
identity map. QEMU tolerates it; real hardware wants it uncacheable. identity map. QEMU tolerates it; real hardware wants it uncacheable.
+6
View File
@@ -152,6 +152,12 @@ If a class driver needs `mmio`, it has become an HCD and should be one.
ack cycle. Legacy INTx (`_PRT` parsing + shared lines) is deliberately skipped — MSI ack cycle. Legacy INTx (`_PRT` parsing + shared lines) is deliberately skipped — MSI
is the real answer. QEMU's HPET has no MSI, so delivery is proven with a self-IPI; the is the real answer. QEMU's HPET has no MSI, so delivery is proven with a self-IPI; the
first PCI driver is the first real consumer. first PCI driver is the first real consumer.
- **Port I/O** — `io_read`/`io_write(device_id, resource_index, offset, width[, value])`:
a claimed device's `io_port` resource lets a driver read/write its ports, gated exactly
like `mmio_map` gates memory (direct ring-3 `in`/`out` stays a #GP). This is what makes
a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with
a syscall per access. `io_port` resources were recorded by discovery and ignored — now
they're used.
- **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table, - **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table,
maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the
platform info). This is detection only — **no translation domains are programmed, so platform info). This is detection only — **no translation domains are programmed, so
+18 -22
View File
@@ -266,26 +266,24 @@ controller drivers fit together.
Worth knowing before you write the second driver: Worth knowing before you write the second driver:
- **Ring 3 has no port I/O.** The TSS I/O permission bitmap is absent Several things this list used to warn about are now available (see
(`tss.zig`: `iomap_base = @sizeOf(Tss)`), and IOPL is never raised, so `in`/`out` [driver-model.md](driver-model.md)): **port I/O** (`io_read`/`io_write`, claim-gated by
from a driver is a #GP. That rules out a user-space 16550 UART (`0x3F8`), PS/2 the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so a PS/2 or
(`0x60`/`0x64`), and legacy PCI config (`0xCF8`/`0xCFC`). Everything must be MMIO. 16550 driver goes through these), **DMA memory** (`dma_alloc`: contiguous, pinned,
`io_port` resources are recorded by discovery and then ignored. uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
`mb`/`rmb`/`wmb`). What remains:
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means - **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
granting one grants the other. A `device_register`ed child's *resource* can be narrower granting one grants the other. A `device_register`ed child's *resource* can be narrower
than a page, but its *mapping* can't. than a page, but its *mapping* can't.
- **No DMA memory.** `mmap` gives you writeback-cached, non-contiguous pages and never
tells you their physical address, so you cannot build a descriptor ring. Any driver
for a bus-mastering device is blocked on this.
- **No memory barriers.** There are none in the tree, and `volatile` is not one — it
won't stop the compiler sinking an ordinary store (your DMA descriptor) past a
volatile MMIO store (your doorbell). On x86 you mostly get away with it; on ARM you
will not. See [driver-model.md](driver-model.md#m14).
- **DMA is not contained.** A driver that can program a bus-mastering device can make - **DMA is not contained.** A driver that can program a bus-mastering device can make
that device write to *any* physical address — page tables don't sit between a device that device write to *any* physical address — page tables don't sit between a device
and RAM; an IOMMU does. Until VT-d/DMAR is programmed, `device_claim` on a DMA-capable and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
device is effectively equivalent to granting ring 0. This is the largest gap between are programmed, so `device_claim` on a DMA-capable device is still effectively
the design's promise and what it delivers. equivalent to granting ring 0. This is the largest gap between the design's promise and
what it delivers; enforcement lands with the first DMA driver.
- **No `dev_release`.** A claim is never dropped (only IRQ/MSI bindings are, on exit), so
a device stays owned for the life of its driver — which blocks restart.
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two - **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`). answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
@@ -341,14 +339,12 @@ Two companions cover what `hpet` can't, because it never exits:
## What's next (not done here) ## What's next (not done here)
The big ones — capability passing (class drivers), DMA + barriers and MSI (host The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
controller drivers), and the IOMMU — have proposed signatures in and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as
[driver-model.md](driver-model.md). Smaller items: is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on
the first DMA driver to protect and test against) and these smaller items:
- **Port I/O grants**, so a PS/2 or 16550 driver is possible: either a per-device TSS
I/O permission bitmap swapped on context switch, or `io_in`/`io_out` syscalls gated
by the same claim. The legacy devices that need it are all low-rate, so the syscall
is likely fast enough.
- **Releasing a claim.** There is no `dev_release`, and `devices_broker` never drops a claim on - **Releasing a claim.** There is no `dev_release`, and `devices_broker` never drops a claim on
exit — only IRQ bindings are released. A dead driver's device stays owned forever, exit — only IRQ bindings are released. A dead driver's device stays owned forever,
which blocks restart. which blocks restart.
+18
View File
@@ -90,3 +90,21 @@ pub fn msiBind(device_id: u64, endpoint: usize) ?Msi {
if (failed(rax)) return null; if (failed(rax)) return null;
return .{ .address = rax, .data = @intCast(rdx) }; return .{ .address = rax, .data = @intCast(rdx) };
} }
/// Read `width` bytes (1, 2, or 4) from a port in a claimed device's `io_port`
/// resource, at byte `offset` within it. Ring 3 has no direct `in`/`out`, so a legacy
/// driver (PS/2, 16550 UART) reaches its ports through this claim-gated call — each
/// access is a syscall, which is fine for the low-rate hardware that needs it. Returns
/// null if the capability check fails (device not claimed, wrong resource, out of
/// range). A device that decodes no data returns all-ones, which is a valid value, not
/// a failure.
pub fn ioRead(device_id: u64, resource_index: u64, offset: u64, width: u8) ?u32 {
const r = sc.systemCall4(.io_read, device_id, resource_index, offset, width);
return if (failed(r)) null else @intCast(r);
}
/// Write `value` (its low `width` bytes, 1/2/4) to a port in a claimed device's
/// `io_port` resource, at byte `offset`. Same capability gate as `ioRead`.
pub fn ioWrite(device_id: u64, resource_index: u64, offset: u64, width: u8, value: u32) bool {
return !failed(sc.systemCall5(.io_write, device_id, resource_index, offset, width, value));
}
+2
View File
@@ -40,6 +40,8 @@ pub const SystemCall = enum(u64) {
dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory
dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
_, _,
}; };
+45
View File
@@ -160,6 +160,8 @@ fn system_call(state: *architecture.CpuState) void {
.dma_alloc => systemDmaAlloc(state), .dma_alloc => systemDmaAlloc(state),
.dma_free => systemDmaFree(state), .dma_free => systemDmaFree(state),
.msi_bind => systemMsiBind(state), .msi_bind => systemMsiBind(state),
.io_read => systemIoRead(state),
.io_write => systemIoWrite(state),
_ => fail(state), _ => fail(state),
} }
} }
@@ -271,6 +273,49 @@ fn systemMmioMap(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, base_v + (r.start & (page_size - 1))); // register base architecture.setSystemCallResult(state, base_v + (r.start & (page_size - 1))); // register base
} }
/// Resolve a port-I/O access against the caller's claims. The device must be claimed by
/// `t`, `resource_index` must name one of its `io_port` resources, and the access
/// `[offset, offset+width)` must fall wholly inside it. Returns the absolute 16-bit
/// port, or null if the capability check fails. The claim plus the discovered `io_port`
/// resource are the capability — exactly like `mmio_map` for memory, so a driver can
/// only touch the ports its device actually owns, never a raw `in`/`out` to anywhere.
pub fn resolveIoPort(t: *scheduler.Task, device_id: u64, resource_index: u64, offset: u64, width: u64) ?u16 {
if (width != 1 and width != 2 and width != 4) return null;
const owner = devices_broker.ownerOf(device_id) orelse return null;
if (owner != t.id) return null; // not claimed by this process
const r = devices_broker.resourceOf(device_id, resource_index) orelse return null;
if (r.kind != @intFromEnum(device_abi.ResourceKind.io_port)) return null;
if (offset + width > r.len) return null; // access escapes the claimed port range
const port = r.start + offset;
if (port + width > 0x1_0000) return null; // I/O ports are 16-bit
return @intCast(port);
}
/// io_read(device_id, resource_index, offset, width) -> value: read `width` bytes (1/2/4)
/// from a port in a claimed device's `io_port` resource. Ring 3 has no direct `in`/`out`
/// (no TSS I/O bitmap, IOPL never raised), so a legacy driver (PS/2, 16550 UART) reaches
/// its ports through this claim-gated call — low-rate hardware, so a syscall per access
/// is fine. See docs/drivers.md.
fn systemIoRead(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const width = architecture.systemCallArg(state, 3);
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
architecture.setSystemCallResult(state, architecture.pioRead(@intCast(width), port));
}
/// io_write(device_id, resource_index, offset, width, value) -> 0: write `value` (low
/// `width` bytes) to a port in a claimed device's `io_port` resource. Same capability
/// gate as `io_read`.
fn systemIoWrite(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const width = architecture.systemCallArg(state, 3);
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
architecture.pioWrite(@intCast(width), port, @intCast(architecture.systemCallArg(state, 4)));
architecture.setSystemCallResult(state, 0);
}
/// dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): grant `len` bytes (rounded up to /// dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): grant `len` bytes (rounded up to
/// whole pages) of DMA-capable memory — physically contiguous, zeroed, pinned, and /// whole pages) of DMA-capable memory — physically contiguous, zeroed, pinned, and
/// strong-uncacheable (coherent) — mapping it into the caller's DMA arena and handing /// strong-uncacheable (coherent) — mapping it into the caller's DMA arena and handing
+49
View File
@@ -90,6 +90,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
msiTest(); msiTest();
} else if (eql(case, "iommu")) { } else if (eql(case, "iommu")) {
iommuTest(); iommuTest();
} else if (eql(case, "ioport")) {
ioPortTest();
} else if (eql(case, "smp")) { } else if (eql(case, "smp")) {
smpTest(); smpTest();
} else if (eql(case, "affinity")) { } else if (eql(case, "affinity")) {
@@ -1051,6 +1053,53 @@ fn iommuTest() void {
result(); result();
} }
/// Port I/O grants: ring 3 has no `in`/`out`, so a legacy driver reaches its ports
/// through `io_read`/`io_write`, gated by `device_claim` and the device's `io_port`
/// resource exactly like `mmio_map` gates memory. Target the PS/2 controller's status
/// port (0x64) — discovered on every PC and side-effect-free to read. Proves the
/// capability gate (`resolveIoPort` admits an in-range access, refuses out-of-range,
/// over-wide, and unclaimed) and that the kernel actually performs the `in`. The
/// `io_read`/`io_write` syscalls wrap this with the same ring-3 dispatch every device
/// driver already uses.
fn ioPortTest() void {
log("DANOS-TEST-BEGIN: ioport\n", .{});
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var found_id: ?u64 = null;
var found_res: u64 = 0;
outer: for (buffer[0..n]) |d| {
for (0..d.resource_count) |ri| {
const r = d.resources[ri];
if (r.kind == @intFromEnum(device_abi.ResourceKind.io_port) and r.start == 0x64 and r.len >= 1) {
found_id = d.id;
found_res = ri;
break :outer;
}
}
}
const id = found_id orelse {
check("discovered the PS/2 status port (io_port 0x64)", false);
result();
return;
};
check("discovered the PS/2 status port (io_port 0x64)", true);
const me = scheduler.current();
check("claimed the io_port device", devices_broker.claim(id, me.id));
check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0, 1) == 0x64);
check("an over-wide access is refused", process.resolveIoPort(me, id, found_res, 0, 2) == null);
check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 1, 1) == null);
check("an unclaimed device id is refused", process.resolveIoPort(me, 0xDEAD_BEEF, found_res, 0, 1) == null);
// The kernel actually issues the `in`. Reaching this line at all proves it didn't
// fault; a width-1 read must return a single byte.
const status = architecture.pioRead(1, 0x64);
check("reading the PS/2 status port returned a byte", status <= 0xFF);
log("DANOS-IOPORT: PS/2 status = 0x{x}\n", .{status});
result();
}
var proc_worker_run: bool = true; var proc_worker_run: bool = true;
var proc_worker_ran: bool = false; var proc_worker_ran: bool = false;
+5
View File
@@ -146,6 +146,11 @@ CASES = [
"qemu_extra": ["-device", "intel-iommu,intremap=off"], "qemu_extra": ["-device", "intel-iommu,intremap=off"],
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# Port I/O grants: a claimed device's io_port resource lets a driver read/write its
# ports (PS/2 status 0x64), gated by the claim; out-of-range/unclaimed is refused.
{"name": "ioport",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Parallelism: needs more than one core, so this case boots with -smp 4. # Parallelism: needs more than one core, so this case boots with -smp 4.
{"name": "smp", {"name": "smp",
"smp": 4, "smp": 4,