Port I/O grants (io_read / io_write) + refresh stale driver docs
Ring 3 still has no direct in/out (no TSS I/O bitmap, IOPL never raised — a #GP), but a
driver no longer needs it: io_read(device_id, resource_index, offset, width) and
io_write(..., value) grant port access the same way mmio_map grants memory. The claim
plus the device's discovered io_port resource are the capability — resolveIoPort checks
the device is claimed by the caller, the resource is io_port, and [offset, offset+width)
stays inside it, then issues the in/out via architecture.pioRead/pioWrite. So a PS/2 or
16550 driver is now writable; the low-rate legacy hardware that needs port I/O is fine
with a syscall per access. io_port resources were recorded by discovery and ignored —
now they're used. Runtime: device.ioRead/ioWrite.
New `ioport` test claims QEMU's PS/2 controller (io_port 0x64, discovered via ACPI) and
checks the gate admits an in-range access, refuses over-wide / out-of-range / unclaimed,
and that the kernel actually reads the status port (0x1c). Suite 41/41 plus host tests.
Docs refreshed for the whole M13-M16 + port-I/O reality: drivers.md "Limits" no longer
lists port I/O, DMA memory, or barriers as missing (they exist) and its "what's next"
reflects that; the FSH doc's character-device and block-device sections are corrected
("cannot host a block driver at all" is no longer true — writable now, not yet memory-
safe pending IOMMU enforcement); device-interrupts.md unblocks the keyboard and narrows
the interrupt gap to MSI-X.
This commit is contained in:
@@ -64,14 +64,15 @@ already provide — it claims its device, maps its registers with `mmio_map`, an
|
||||
on `replyWait` for either an interrupt or a client request. `system/drivers/hpet/hpet.zig` is already
|
||||
that program, minus the client half.
|
||||
|
||||
The obstacle is not the file type, it is which hardware a ring-3 driver can actually
|
||||
drive. Port I/O is unavailable to user space — the TSS I/O permission bitmap is absent
|
||||
and IOPL is never raised — so `in`/`out` from a driver is a #GP. That excludes the
|
||||
16550 UART at `0x3F8` and PS/2 at `0x60`/`0x64`, which is to say it excludes the
|
||||
obvious implementations of `/dev/tty`, `/dev/ttyS0` and a keyboard node. Until either
|
||||
port I/O grants or a memory-mapped UART exist, serial output stays a kernel service
|
||||
reached through the `write` system call rather than a file. A memory-mapped device such
|
||||
as the framebuffer has no such problem and is the more likely first real entry here.
|
||||
The obstacle was never the file type; it is which hardware a ring-3 driver can reach.
|
||||
Direct `in`/`out` from user space is still a #GP (no TSS I/O bitmap, IOPL never raised),
|
||||
but a driver no longer needs it: **`io_read`/`io_write`** grant port access the same way
|
||||
`mmio_map` grants memory — gated by `device_claim` and the device's discovered `io_port`
|
||||
resource. So the 16550 UART at `0x3F8` and the PS/2 controller at `0x60`/`0x64` (and thus
|
||||
`/dev/ttyS0` and a keyboard node) are now writable as ordinary ring-3 drivers; the
|
||||
low-rate legacy hardware that needs port I/O is fine with a syscall per access. A
|
||||
memory-mapped device such as the framebuffer, needing no port I/O at all, remains the
|
||||
easiest first entry.
|
||||
|
||||
### Block devices
|
||||
|
||||
@@ -79,20 +80,23 @@ A block device is addressed in fixed-size blocks and, unlike a character device,
|
||||
layer above is free to buffer, reorder, coalesce and retry requests against it. Disks
|
||||
and other persistent storage are the whole population of this class.
|
||||
|
||||
**danos cannot host a block driver at all today,** and the reason is worth stating
|
||||
plainly because it is not a matter of unwritten code. Every storage controller worth
|
||||
naming is a bus master: it is programmed by handing it the physical address of a
|
||||
descriptor ring and left to read and write memory on its own. A ring-3 driver cannot
|
||||
build such a ring, because `mmap` returns writeback-cached, physically discontiguous
|
||||
pages and never discloses their physical address. Nor should it be allowed to: a device
|
||||
programmed with an arbitrary physical address writes to arbitrary physical memory, and
|
||||
page tables do not sit between a device and RAM — an IOMMU does. Granting a DMA-capable
|
||||
device to a driver process, with no IOMMU programmed, is equivalent to granting ring 0,
|
||||
which would forfeit the isolation that motivates user-space drivers in the first place.
|
||||
A block driver is now **writable, but not yet memory-safe.** Every storage controller
|
||||
worth naming is a bus master: it is programmed by handing it the physical address of a
|
||||
descriptor ring and left to read and write memory on its own. That ring is exactly what
|
||||
**`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its
|
||||
physical address disclosed — and **`/lib/mmio`**'s barriers order the descriptor writes
|
||||
against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver
|
||||
can be written today (the M14/M15 work in [driver-model.md](driver-model.md); the earlier
|
||||
"cannot host a block driver at all" is no longer true).
|
||||
|
||||
Block devices therefore wait on DMA-capable memory, memory barriers, and VT-d/DMAR —
|
||||
the M14–M16 work in [driver-model.md](driver-model.md). A ramdisk over the initial ramdisk is
|
||||
the one block-shaped thing implementable now, and it needs no driver process.
|
||||
What is *not* yet true is that it is safe. A device programmed with an arbitrary physical
|
||||
address writes to arbitrary physical memory, and page tables do not sit between a device
|
||||
and RAM — an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
|
||||
are programmed, so granting a DMA-capable device to a driver process is still equivalent
|
||||
to granting ring 0. Until per-device domains confine a driver's DMA to the buffers it
|
||||
`dma_alloc`'d, a block driver works but forfeits the isolation that motivates user-space
|
||||
drivers — enforcement is the next step, and lands with that first driver. A ramdisk over
|
||||
the initial ramdisk remains the one block-shaped thing that needs no driver process at all.
|
||||
|
||||
### Pseudo-devices
|
||||
|
||||
|
||||
@@ -156,10 +156,11 @@ spinning in unrelated code — is the whole mechanism working end to end.
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and ring 3
|
||||
has no port I/O yet, so the first *input* device is blocked on either an I/O
|
||||
permission bitmap or `io_in`/`io_out` syscalls ([drivers.md](drivers.md)).
|
||||
- **MSI/MSI-X**: per-device vectors, edge-triggered and unshared, which retire the
|
||||
I/O APIC's mask/ack cycle and its 24-GSI ceiling.
|
||||
- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and port I/O is
|
||||
now available to ring 3 via the claim-gated `io_read`/`io_write` syscalls
|
||||
([drivers.md](drivers.md)) — so the first *input* device is unblocked; it just needs
|
||||
writing (claim the controller, `irq_bind` GSI 1, read scancodes from `0x60`).
|
||||
- **MSI-X**: `msi_bind` gives one per-device edge-triggered vector (M15); MSI-X's
|
||||
multi-vector table (many queues per device, e.g. NVMe) is the remaining extension.
|
||||
- **The LAPIC's own page** is still mapped writeback-cacheable like the rest of the
|
||||
identity map. QEMU tolerates it; real hardware wants it uncacheable.
|
||||
|
||||
@@ -152,6 +152,12 @@ If a class driver needs `mmio`, it has become an HCD and should be one.
|
||||
ack cycle. Legacy INTx (`_PRT` parsing + shared lines) is deliberately skipped — MSI
|
||||
is the real answer. QEMU's HPET has no MSI, so delivery is proven with a self-IPI; the
|
||||
first PCI driver is the first real consumer.
|
||||
- **Port I/O** — `io_read`/`io_write(device_id, resource_index, offset, width[, value])`:
|
||||
a claimed device's `io_port` resource lets a driver read/write its ports, gated exactly
|
||||
like `mmio_map` gates memory (direct ring-3 `in`/`out` stays a #GP). This is what makes
|
||||
a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with
|
||||
a syscall per access. `io_port` resources were recorded by discovery and ignored — now
|
||||
they're used.
|
||||
- **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table,
|
||||
maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the
|
||||
platform info). This is detection only — **no translation domains are programmed, so
|
||||
|
||||
+18
-22
@@ -266,26 +266,24 @@ controller drivers fit together.
|
||||
|
||||
Worth knowing before you write the second driver:
|
||||
|
||||
- **Ring 3 has no port I/O.** The TSS I/O permission bitmap is absent
|
||||
(`tss.zig`: `iomap_base = @sizeOf(Tss)`), and IOPL is never raised, so `in`/`out`
|
||||
from a driver is a #GP. That rules out a user-space 16550 UART (`0x3F8`), PS/2
|
||||
(`0x60`/`0x64`), and legacy PCI config (`0xCF8`/`0xCFC`). Everything must be MMIO.
|
||||
`io_port` resources are recorded by discovery and then ignored.
|
||||
Several things this list used to warn about are now available (see
|
||||
[driver-model.md](driver-model.md)): **port I/O** (`io_read`/`io_write`, claim-gated by
|
||||
the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so a PS/2 or
|
||||
16550 driver goes through these), **DMA memory** (`dma_alloc`: contiguous, pinned,
|
||||
uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
|
||||
`mb`/`rmb`/`wmb`). What remains:
|
||||
|
||||
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
|
||||
granting one grants the other. A `device_register`ed child's *resource* can be narrower
|
||||
than a page, but its *mapping* can't.
|
||||
- **No DMA memory.** `mmap` gives you writeback-cached, non-contiguous pages and never
|
||||
tells you their physical address, so you cannot build a descriptor ring. Any driver
|
||||
for a bus-mastering device is blocked on this.
|
||||
- **No memory barriers.** There are none in the tree, and `volatile` is not one — it
|
||||
won't stop the compiler sinking an ordinary store (your DMA descriptor) past a
|
||||
volatile MMIO store (your doorbell). On x86 you mostly get away with it; on ARM you
|
||||
will not. See [driver-model.md](driver-model.md#m14).
|
||||
- **DMA is not contained.** A driver that can program a bus-mastering device can make
|
||||
that device write to *any* physical address — page tables don't sit between a device
|
||||
and RAM; an IOMMU does. Until VT-d/DMAR is programmed, `device_claim` on a DMA-capable
|
||||
device is effectively equivalent to granting ring 0. This is the largest gap between
|
||||
the design's promise and what it delivers.
|
||||
and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
|
||||
are programmed, so `device_claim` on a DMA-capable device is still effectively
|
||||
equivalent to granting ring 0. This is the largest gap between the design's promise and
|
||||
what it delivers; enforcement lands with the first DMA driver.
|
||||
- **No `dev_release`.** A claim is never dropped (only IRQ/MSI bindings are, on exit), so
|
||||
a device stays owned for the life of its driver — which blocks restart.
|
||||
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
|
||||
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
|
||||
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
|
||||
@@ -341,14 +339,12 @@ Two companions cover what `hpet` can't, because it never exits:
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
The big ones — capability passing (class drivers), DMA + barriers and MSI (host
|
||||
controller drivers), and the IOMMU — have proposed signatures in
|
||||
[driver-model.md](driver-model.md). Smaller items:
|
||||
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
|
||||
and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as
|
||||
is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550
|
||||
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on
|
||||
the first DMA driver to protect and test against) and these smaller items:
|
||||
|
||||
- **Port I/O grants**, so a PS/2 or 16550 driver is possible: either a per-device TSS
|
||||
I/O permission bitmap swapped on context switch, or `io_in`/`io_out` syscalls gated
|
||||
by the same claim. The legacy devices that need it are all low-rate, so the syscall
|
||||
is likely fast enough.
|
||||
- **Releasing a claim.** There is no `dev_release`, and `devices_broker` never drops a claim on
|
||||
exit — only IRQ bindings are released. A dead driver's device stays owned forever,
|
||||
which blocks restart.
|
||||
|
||||
@@ -90,3 +90,21 @@ pub fn msiBind(device_id: u64, endpoint: usize) ?Msi {
|
||||
if (failed(rax)) return null;
|
||||
return .{ .address = rax, .data = @intCast(rdx) };
|
||||
}
|
||||
|
||||
/// Read `width` bytes (1, 2, or 4) from a port in a claimed device's `io_port`
|
||||
/// resource, at byte `offset` within it. Ring 3 has no direct `in`/`out`, so a legacy
|
||||
/// driver (PS/2, 16550 UART) reaches its ports through this claim-gated call — each
|
||||
/// access is a syscall, which is fine for the low-rate hardware that needs it. Returns
|
||||
/// null if the capability check fails (device not claimed, wrong resource, out of
|
||||
/// range). A device that decodes no data returns all-ones, which is a valid value, not
|
||||
/// a failure.
|
||||
pub fn ioRead(device_id: u64, resource_index: u64, offset: u64, width: u8) ?u32 {
|
||||
const r = sc.systemCall4(.io_read, device_id, resource_index, offset, width);
|
||||
return if (failed(r)) null else @intCast(r);
|
||||
}
|
||||
|
||||
/// Write `value` (its low `width` bytes, 1/2/4) to a port in a claimed device's
|
||||
/// `io_port` resource, at byte `offset`. Same capability gate as `ioRead`.
|
||||
pub fn ioWrite(device_id: u64, resource_index: u64, offset: u64, width: u8, value: u32) bool {
|
||||
return !failed(sc.systemCall5(.io_write, device_id, resource_index, offset, width, value));
|
||||
}
|
||||
|
||||
@@ -40,6 +40,8 @@ pub const SystemCall = enum(u64) {
|
||||
dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory
|
||||
dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc
|
||||
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
|
||||
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
|
||||
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
|
||||
_,
|
||||
};
|
||||
|
||||
|
||||
@@ -160,6 +160,8 @@ fn system_call(state: *architecture.CpuState) void {
|
||||
.dma_alloc => systemDmaAlloc(state),
|
||||
.dma_free => systemDmaFree(state),
|
||||
.msi_bind => systemMsiBind(state),
|
||||
.io_read => systemIoRead(state),
|
||||
.io_write => systemIoWrite(state),
|
||||
_ => fail(state),
|
||||
}
|
||||
}
|
||||
@@ -271,6 +273,49 @@ fn systemMmioMap(state: *architecture.CpuState) void {
|
||||
architecture.setSystemCallResult(state, base_v + (r.start & (page_size - 1))); // register base
|
||||
}
|
||||
|
||||
/// Resolve a port-I/O access against the caller's claims. The device must be claimed by
|
||||
/// `t`, `resource_index` must name one of its `io_port` resources, and the access
|
||||
/// `[offset, offset+width)` must fall wholly inside it. Returns the absolute 16-bit
|
||||
/// port, or null if the capability check fails. The claim plus the discovered `io_port`
|
||||
/// resource are the capability — exactly like `mmio_map` for memory, so a driver can
|
||||
/// only touch the ports its device actually owns, never a raw `in`/`out` to anywhere.
|
||||
pub fn resolveIoPort(t: *scheduler.Task, device_id: u64, resource_index: u64, offset: u64, width: u64) ?u16 {
|
||||
if (width != 1 and width != 2 and width != 4) return null;
|
||||
const owner = devices_broker.ownerOf(device_id) orelse return null;
|
||||
if (owner != t.id) return null; // not claimed by this process
|
||||
const r = devices_broker.resourceOf(device_id, resource_index) orelse return null;
|
||||
if (r.kind != @intFromEnum(device_abi.ResourceKind.io_port)) return null;
|
||||
if (offset + width > r.len) return null; // access escapes the claimed port range
|
||||
const port = r.start + offset;
|
||||
if (port + width > 0x1_0000) return null; // I/O ports are 16-bit
|
||||
return @intCast(port);
|
||||
}
|
||||
|
||||
/// io_read(device_id, resource_index, offset, width) -> value: read `width` bytes (1/2/4)
|
||||
/// from a port in a claimed device's `io_port` resource. Ring 3 has no direct `in`/`out`
|
||||
/// (no TSS I/O bitmap, IOPL never raised), so a legacy driver (PS/2, 16550 UART) reaches
|
||||
/// its ports through this claim-gated call — low-rate hardware, so a syscall per access
|
||||
/// is fine. See docs/drivers.md.
|
||||
fn systemIoRead(state: *architecture.CpuState) void {
|
||||
const t = scheduler.current();
|
||||
if (t.aspace == 0) return fail(state);
|
||||
const width = architecture.systemCallArg(state, 3);
|
||||
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
|
||||
architecture.setSystemCallResult(state, architecture.pioRead(@intCast(width), port));
|
||||
}
|
||||
|
||||
/// io_write(device_id, resource_index, offset, width, value) -> 0: write `value` (low
|
||||
/// `width` bytes) to a port in a claimed device's `io_port` resource. Same capability
|
||||
/// gate as `io_read`.
|
||||
fn systemIoWrite(state: *architecture.CpuState) void {
|
||||
const t = scheduler.current();
|
||||
if (t.aspace == 0) return fail(state);
|
||||
const width = architecture.systemCallArg(state, 3);
|
||||
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
|
||||
architecture.pioWrite(@intCast(width), port, @intCast(architecture.systemCallArg(state, 4)));
|
||||
architecture.setSystemCallResult(state, 0);
|
||||
}
|
||||
|
||||
/// dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): grant `len` bytes (rounded up to
|
||||
/// whole pages) of DMA-capable memory — physically contiguous, zeroed, pinned, and
|
||||
/// strong-uncacheable (coherent) — mapping it into the caller's DMA arena and handing
|
||||
|
||||
@@ -90,6 +90,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
||||
msiTest();
|
||||
} else if (eql(case, "iommu")) {
|
||||
iommuTest();
|
||||
} else if (eql(case, "ioport")) {
|
||||
ioPortTest();
|
||||
} else if (eql(case, "smp")) {
|
||||
smpTest();
|
||||
} else if (eql(case, "affinity")) {
|
||||
@@ -1051,6 +1053,53 @@ fn iommuTest() void {
|
||||
result();
|
||||
}
|
||||
|
||||
/// Port I/O grants: ring 3 has no `in`/`out`, so a legacy driver reaches its ports
|
||||
/// through `io_read`/`io_write`, gated by `device_claim` and the device's `io_port`
|
||||
/// resource exactly like `mmio_map` gates memory. Target the PS/2 controller's status
|
||||
/// port (0x64) — discovered on every PC and side-effect-free to read. Proves the
|
||||
/// capability gate (`resolveIoPort` admits an in-range access, refuses out-of-range,
|
||||
/// over-wide, and unclaimed) and that the kernel actually performs the `in`. The
|
||||
/// `io_read`/`io_write` syscalls wrap this with the same ring-3 dispatch every device
|
||||
/// driver already uses.
|
||||
fn ioPortTest() void {
|
||||
log("DANOS-TEST-BEGIN: ioport\n", .{});
|
||||
var buffer: [64]device_abi.DeviceDescriptor = undefined;
|
||||
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
|
||||
|
||||
var found_id: ?u64 = null;
|
||||
var found_res: u64 = 0;
|
||||
outer: for (buffer[0..n]) |d| {
|
||||
for (0..d.resource_count) |ri| {
|
||||
const r = d.resources[ri];
|
||||
if (r.kind == @intFromEnum(device_abi.ResourceKind.io_port) and r.start == 0x64 and r.len >= 1) {
|
||||
found_id = d.id;
|
||||
found_res = ri;
|
||||
break :outer;
|
||||
}
|
||||
}
|
||||
}
|
||||
const id = found_id orelse {
|
||||
check("discovered the PS/2 status port (io_port 0x64)", false);
|
||||
result();
|
||||
return;
|
||||
};
|
||||
check("discovered the PS/2 status port (io_port 0x64)", true);
|
||||
|
||||
const me = scheduler.current();
|
||||
check("claimed the io_port device", devices_broker.claim(id, me.id));
|
||||
check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0, 1) == 0x64);
|
||||
check("an over-wide access is refused", process.resolveIoPort(me, id, found_res, 0, 2) == null);
|
||||
check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 1, 1) == null);
|
||||
check("an unclaimed device id is refused", process.resolveIoPort(me, 0xDEAD_BEEF, found_res, 0, 1) == null);
|
||||
|
||||
// The kernel actually issues the `in`. Reaching this line at all proves it didn't
|
||||
// fault; a width-1 read must return a single byte.
|
||||
const status = architecture.pioRead(1, 0x64);
|
||||
check("reading the PS/2 status port returned a byte", status <= 0xFF);
|
||||
log("DANOS-IOPORT: PS/2 status = 0x{x}\n", .{status});
|
||||
result();
|
||||
}
|
||||
|
||||
var proc_worker_run: bool = true;
|
||||
var proc_worker_ran: bool = false;
|
||||
|
||||
|
||||
@@ -146,6 +146,11 @@ CASES = [
|
||||
"qemu_extra": ["-device", "intel-iommu,intremap=off"],
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# Port I/O grants: a claimed device's io_port resource lets a driver read/write its
|
||||
# ports (PS/2 status 0x64), gated by the claim; out-of-range/unclaimed is refused.
|
||||
{"name": "ioport",
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# Parallelism: needs more than one core, so this case boots with -smp 4.
|
||||
{"name": "smp",
|
||||
"smp": 4,
|
||||
|
||||
Reference in New Issue
Block a user