re-org docs
This commit is contained in:
@@ -63,7 +63,7 @@ zig build release-x86-64
|
|||||||
|
|
||||||
Produces `zig-out/danos-x86-64.iso`, a hybrid ISO that boots flashed raw to a
|
Produces `zig-out/danos-x86-64.iso`, a hybrid ISO that boots flashed raw to a
|
||||||
USB stick (balenaEtcher, dd) or burned to optical media — see
|
USB stick (balenaEtcher, dd) or burned to optical media — see
|
||||||
[docs/release-iso.md](docs/release-iso.md). `zig build check-iso-image`
|
[docs/release-iso.md](docs/os-development-guide/release-iso.md). `zig build check-iso-image`
|
||||||
validates it without booting.
|
validates it without booting.
|
||||||
|
|
||||||
## Run
|
## Run
|
||||||
|
|||||||
+66
-66
@@ -3,102 +3,102 @@
|
|||||||
Notes on how danos boots and draws, written to explain the *why* behind the code
|
Notes on how danos boots and draws, written to explain the *why* behind the code
|
||||||
rather than restate it. Roughly in the order things happen at runtime:
|
rather than restate it. Roughly in the order things happen at runtime:
|
||||||
|
|
||||||
1. **[efi.md](efi.md) — EFI / the boot process.** How UEFI firmware finds and
|
1. **[efi.md](os-development-guide/efi.md) — EFI / the boot process.** How UEFI firmware finds and
|
||||||
runs the bootloader, what the loader gathers before `ExitBootServices`, how it
|
runs the bootloader, what the loader gathers before `ExitBootServices`, how it
|
||||||
loads the kernel ELF, and the ABI contract for the jump into the kernel. Start
|
loads the kernel ELF, and the ABI contract for the jump into the kernel. Start
|
||||||
here.
|
here.
|
||||||
2. **[system-image.md](system-image.md) — system.img, the boot capsule.** The
|
2. **[system-image.md](os-development-guide/system-image.md) — system.img, the boot capsule.** The
|
||||||
bundled user binaries packed into one file in the initial-ramdisk wire
|
bundled user binaries packed into one file in the initial-ramdisk wire
|
||||||
format, because one open + one sequential read is the only file I/O shape
|
format, because one open + one sequential read is the only file I/O shape
|
||||||
firmware is fast at. The trivial container format, the three artifacts one
|
firmware is fast at. The trivial container format, the three artifacts one
|
||||||
build list derives (tree, manifest, capsule), the loader's three-strategy
|
build list derives (tree, manifest, capsule), the loader's three-strategy
|
||||||
fallback chain, and the capsule's kernel-side life as both the spawn table
|
fallback chain, and the capsule's kernel-side life as both the spawn table
|
||||||
and the read-only `/system` mount.
|
and the read-only `/system` mount.
|
||||||
3. **[gop.md](gop.md) — the Graphics Output Protocol.** How UEFI exposes graphics
|
3. **[gop.md](os-development-guide/gop.md) — the Graphics Output Protocol.** How UEFI exposes graphics
|
||||||
modes (unlike fixed VGA modes), how we detect the monitor's native resolution
|
modes (unlike fixed VGA modes), how we detect the monitor's native resolution
|
||||||
from EDID and switch to it, and the pixel formats we accept or reject.
|
from EDID and switch to it, and the pixel formats we accept or reject.
|
||||||
4. **[framebuffer.md](framebuffer.md) — the framebuffer.** What the linear
|
4. **[framebuffer.md](os-development-guide/framebuffer.md) — the framebuffer.** What the linear
|
||||||
framebuffer the loader hands over actually is, and what **pitch** (stride)
|
framebuffer the loader hands over actually is, and what **pitch** (stride)
|
||||||
means versus width — the detail you have to get right to avoid a skewed image.
|
means versus width — the detail you have to get right to avoid a skewed image.
|
||||||
5. **[memory-map.md](memory-map.md) — the memory map.** How the loader learns what
|
5. **[memory-map.md](os-development-guide/memory-map.md) — the memory map.** How the loader learns what
|
||||||
physical RAM exists and hands it to the kernel in danos's own neutral format,
|
physical RAM exists and hands it to the kernel in danos's own neutral format,
|
||||||
rather than leaking UEFI's memory descriptors across the boundary.
|
rather than leaking UEFI's memory descriptors across the boundary.
|
||||||
6. **[frame-allocator.md](frame-allocator.md) — the physical frame allocator.** The
|
6. **[frame-allocator.md](os-development-guide/frame-allocator.md) — the physical frame allocator.** The
|
||||||
bitmap allocator that hands out and reclaims 4 KiB physical frames from that
|
bitmap allocator that hands out and reclaims 4 KiB physical frames from that
|
||||||
map — the primitive page tables and the heap are built on.
|
map — the primitive page tables and the heap are built on.
|
||||||
7. **[interrupts.md](interrupts.md) — interrupts and exceptions.** The GDT, IDT and
|
7. **[interrupts.md](os-development-guide/interrupts.md) — interrupts and exceptions.** The GDT, IDT and
|
||||||
TSS, the exception stubs, and the handler that reports a CPU fault in red instead
|
TSS, the exception stubs, and the handler that reports a CPU fault in red instead
|
||||||
of letting it triple-fault into a silent reset.
|
of letting it triple-fault into a silent reset.
|
||||||
8. **[paging.md](paging.md) — the kernel's page tables.** Building our own 4-level
|
8. **[paging.md](os-development-guide/paging.md) — the kernel's page tables.** Building our own 4-level
|
||||||
page tables, identity-mapping the low 4 GiB, and switching CR3 off the firmware's
|
page tables, identity-mapping the low 4 GiB, and switching CR3 off the firmware's
|
||||||
tables onto ours.
|
tables onto ours.
|
||||||
9. **[device-interrupts.md](device-interrupts.md) — device interrupts.** The Local
|
9. **[device-interrupts.md](device-driver-development-guide/device-interrupts.md) — device interrupts.** The Local
|
||||||
APIC and its timer — the kernel's first interrupt that is *handled and returned
|
APIC and its timer — the kernel's first interrupt that is *handled and returned
|
||||||
from*, giving it a heartbeat.
|
from*, giving it a heartbeat.
|
||||||
10. **[heap.md](heap.md) — the kernel heap.** A growable free-list allocator built on
|
10. **[heap.md](os-development-guide/heap.md) — the kernel heap.** A growable free-list allocator built on
|
||||||
the VMM, exposed as a `std.mem.Allocator` so std containers work — dynamic
|
the VMM, exposed as a `std.mem.Allocator` so std containers work — dynamic
|
||||||
allocation for the kernel.
|
allocation for the kernel.
|
||||||
11. **[scheduling.md](scheduling.md) — the scheduler.** Fixed-priority preemptive
|
11. **[scheduling.md](os-development-guide/scheduling.md) — the scheduler.** Fixed-priority preemptive
|
||||||
multitasking: kernel threads, the context switch, O(1) priority selection, and
|
multitasking: kernel threads, the context switch, O(1) priority selection, and
|
||||||
blocking (sleep, wait queues) — the leap to a running system.
|
blocking (sleep, wait queues) — the leap to a running system.
|
||||||
12. **[ipc.md](ipc.md) — inter-process communication.** Bounded blocking
|
12. **[ipc.md](device-driver-development-guide/ipc.md) — inter-process communication.** Bounded blocking
|
||||||
message-passing channels, then synchronous call/reply between *processes* over
|
message-passing channels, then synchronous call/reply between *processes* over
|
||||||
endpoints — the backbone the microkernel's isolated servers talk over.
|
endpoints — the backbone the microkernel's isolated servers talk over.
|
||||||
13. **[syscall.md](syscall.md) — system calls.** How ring 3 asks the kernel for
|
13. **[syscall.md](os-development-guide/syscall.md) — system calls.** How ring 3 asks the kernel for
|
||||||
something: the `syscall`/`sysret` fast path, the trap frame, and why the table is
|
something: the `syscall`/`sysret` fast path, the trap frame, and why the table is
|
||||||
deliberately tiny. The numbers are a **private** ABI — [vdso.md](vdso.md) designs
|
deliberately tiny. The numbers are a **private** ABI — [vdso.md](os-development-guide/vdso.md) designs
|
||||||
the public boundary that will hide them.
|
the public boundary that will hide them.
|
||||||
14. **[vfs-protocol.md](vfs-protocol.md) — the VFS wire protocol.** The language-neutral
|
14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral
|
||||||
byte-level spec of the file protocol spoken over IPC: request/reply headers,
|
byte-level spec of the file protocol spoken over IPC: request/reply headers,
|
||||||
the operation table, mount routing, and the append-only evolution rules — the
|
the operation table, mount routing, and the append-only evolution rules — the
|
||||||
first IPC protocol documented as public ABI.
|
first IPC protocol documented as public ABI.
|
||||||
15. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an
|
15. **[drivers.md](device-driver-development-guide/drivers.md) — writing a driver.** The payoff: a driver is an
|
||||||
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
||||||
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
||||||
unmask.
|
unmask.
|
||||||
16. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
|
16. **[driver-model.md](device-driver-development-guide/driver-model.md) — buses, classes and host controllers.** How
|
||||||
real driver stacks factor into three shapes and how families share code. The
|
real driver stacks factor into three shapes and how families share code. The
|
||||||
three primitives it proposed are long since built (M13 capability passing,
|
three primitives it proposed are long since built (M13 capability passing,
|
||||||
M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them —
|
M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them —
|
||||||
hello, supervision, restart — is built too (device-manager.md, M18).
|
hello, supervision, restart — is built too (device-manager.md, M18).
|
||||||
17. **[usb-hub.md](usb-hub.md) — USB hubs.** Built (M22): why hub topology is handled
|
17. **[usb-hub.md](device-driver-development-guide/usb-hub.md) — USB hubs.** Built (M22): why hub topology is handled
|
||||||
*inside* the `usb-xhci-bus` driver rather than a separate hub class driver — a
|
*inside* the `usb-xhci-bus` driver rather than a separate hub class driver — a
|
||||||
device behind a hub is reached by the **controller**, programmed with a route
|
device behind a hub is reached by the **controller**, programmed with a route
|
||||||
string in its slot context — plus the compound-hub reality (a USB 3.0 hub is
|
string in its slot context — plus the compound-hub reality (a USB 3.0 hub is
|
||||||
physically two hubs) and detection via the hub's status-change interrupt endpoint.
|
physically two hubs) and detection via the hub's status-change interrupt endpoint.
|
||||||
18. **[process-management.md](process-management.md) — process management.** The
|
18. **[process-management.md](os-development-guide/process-management.md) — process management.** The
|
||||||
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
|
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
|
||||||
supervision link as the kill authority, and child-exit notifications over the
|
supervision link as the kill authority, and child-exit notifications over the
|
||||||
same endpoints IRQs arrive on.
|
same endpoints IRQs arrive on.
|
||||||
19. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
|
19. **[process-lifecycle.md](os-development-guide/process-lifecycle.md) — the process lifecycle.** Built
|
||||||
(M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
|
(M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
|
||||||
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
|
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
|
||||||
`process` module interface, exit reasons, published exit events any stateful
|
`process` module interface, exit reasons, published exit events any stateful
|
||||||
service can subscribe to (the VFS releasing dead clients' handles), and the two
|
service can subscribe to (the VFS releasing dead clients' handles), and the two
|
||||||
iron rules (cleanup is the kernel's job; kill is not a signal).
|
iron rules (cleanup is the kernel's job; kill is not a signal).
|
||||||
20. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
|
20. **[device-manager.md](device-driver-development-guide/device-manager.md) — the device manager.** Built (M18,
|
||||||
through the app surface): the
|
through the app surface): the
|
||||||
tree, the matcher, and the supervisor. Tree structure lives in the manager,
|
tree, the matcher, and the supervisor. Tree structure lives in the manager,
|
||||||
authority stays in the kernel; bus drivers report what they see; drivers are
|
authority stays in the kernel; bus drivers report what they see; drivers are
|
||||||
restarted through the lifecycle vocabulary — the plan that turns
|
restarted through the lifecycle vocabulary — the plan that turns
|
||||||
[resilience.md](resilience.md)'s restart goal into increments.
|
[resilience.md](os-development-guide/resilience.md)'s restart goal into increments.
|
||||||
21. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard,
|
21. **[input.md](device-driver-development-guide/input.md) — the input module.** Broadcasting input events (keyboard,
|
||||||
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
|
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
|
||||||
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
|
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
|
||||||
service layered on top.
|
service layered on top.
|
||||||
22. **[display.md](display.md) — the display service.** The display half of the GUI
|
22. **[display.md](device-driver-development-guide/display.md) — the display service.** The display half of the GUI
|
||||||
track: a user-space compositor that owns the framebuffer, composes a layer stack into
|
track: a user-space compositor that owns the framebuffer, composes a layer stack into
|
||||||
a double buffer, and presents it. Why GOP and the PCI display device are two views of
|
a double buffer, and presents it. Why GOP and the PCI display device are two views of
|
||||||
one controller, the device-node + write-combining handoff, and what flicker-free buys
|
one controller, the device-node + write-combining handoff, and what flicker-free buys
|
||||||
that tear-free doesn't. Plan: [display-plan.md](display-plan.md). **v2** (complete) makes
|
that tear-free doesn't. Plan: [display-plan.md](device-driver-development-guide/display-plan.md). **v2** (complete) makes
|
||||||
scanout a pluggable backend — GOP floor + a native virtio-gpu driver, hot-attached, with
|
scanout a pluggable backend — GOP floor + a native virtio-gpu driver, hot-attached, with
|
||||||
runtime mode-set, EDID, fenced vsync presents, and restart re-attach:
|
runtime mode-set, EDID, fenced vsync presents, and restart re-attach:
|
||||||
[display-v2.md](display-v2.md), plan [display-v2-plan.md](display-v2-plan.md). Looking
|
[display-v2.md](device-driver-development-guide/display-v2.md), plan [display-v2-plan.md](device-driver-development-guide/display-v2-plan.md). Looking
|
||||||
further out, three research snapshots survey what a *native* driver for real GPU silicon
|
further out, three research snapshots survey what a *native* driver for real GPU silicon
|
||||||
would take as another `.scanout` backend: [nvidia-gpus.md](nvidia-gpus.md) (RTX 3060 /
|
would take as another `.scanout` backend: [nvidia-gpus.md](device-driver-development-guide/nvidia-gpus.md) (RTX 3060 /
|
||||||
Ampere), [amd-gpus.md](amd-gpus.md) (RX 6600 / RDNA2), and [intel-igpu.md](intel-igpu.md)
|
Ampere), [amd-gpus.md](device-driver-development-guide/amd-gpus.md) (RX 6600 / RDNA2), and [intel-igpu.md](device-driver-development-guide/intel-igpu.md)
|
||||||
(Intel iGPU).
|
(Intel iGPU).
|
||||||
23. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
23. **[halting.md](os-development-guide/halting.md) — halting.** Why a kernel can't just "exit", and
|
||||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||||
|
|
||||||
Start with the north star:
|
Start with the north star:
|
||||||
@@ -108,7 +108,7 @@ Start with the north star:
|
|||||||
**resilience** (restartable components). Win condition: runs on the author's PC and
|
**resilience** (restartable components). Win condition: runs on the author's PC and
|
||||||
both Raspberry Pis, ideally with a GUI. Real-time is an option to explore, not a
|
both Raspberry Pis, ideally with a GUI. Real-time is an option to explore, not a
|
||||||
requirement. The *why* that shapes everything below.
|
requirement. The *why* that shapes everything below.
|
||||||
- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on
|
- **[resilience.md](os-development-guide/resilience.md) — resilience.** A design note (not built yet) on
|
||||||
fault isolation + live restart — the reincarnation-server + capability model that
|
fault isolation + live restart — the reincarnation-server + capability model that
|
||||||
makes "if I break it, I can restart it" real. danos's core motivation.
|
makes "if I break it, I can restart it" real. danos's core motivation.
|
||||||
- **[zig-self-hosting.md](zig-self-hosting.md) — running Zig on danos.** A design note
|
- **[zig-self-hosting.md](zig-self-hosting.md) — running Zig on danos.** A design note
|
||||||
@@ -117,19 +117,19 @@ Start with the north star:
|
|||||||
port to **one seam** (`std.os.danos`), so we build an `os` seam module (→ that seam) plus
|
port to **one seam** (`std.os.danos`), so we build an `os` seam module (→ that seam) plus
|
||||||
the thin `file-system` module, retire the `posix` shim, and follow a phased path to
|
the thin `file-system` module, retire the `posix` shim, and follow a phased path to
|
||||||
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
|
`zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation.
|
||||||
- **[threading.md](threading.md) — threads, the std-shaped way.** **Built** (M1–M6):
|
- **[threading.md](os-development-guide/threading.md) — threads, the std-shaped way.** **Built** (M1–M6):
|
||||||
the `thread` module's `Thread` mirrors `std.Thread`'s API (spawn/join/detach, Mutex/Condition/
|
the `thread` module's `Thread` mirrors `std.Thread`'s API (spawn/join/detach, Mutex/Condition/
|
||||||
Semaphore) over a **private** thread ABI — several tasks sharing one address space via
|
Semaphore) over a **private** thread ABI — several tasks sharing one address space via
|
||||||
a `thread_spawn` syscall, futex-backed blocking, address-space refcounting. Why it's the
|
a `thread_spawn` syscall, futex-backed blocking, address-space refcounting. Why it's the
|
||||||
native type and not literal `std.Thread` (the [private ABI](syscall.md)), and why
|
native type and not literal `std.Thread` (the [private ABI](os-development-guide/syscall.md)), and why
|
||||||
threads stay a narrow opt-in against the [resilience](resilience.md) default. Build
|
threads stay a narrow opt-in against the [resilience](os-development-guide/resilience.md) default. Build
|
||||||
plan + gates: [threading-plan.md](threading-plan.md).
|
plan + gates: [threading-plan.md](os-development-guide/threading-plan.md).
|
||||||
- **[vdso.md](vdso.md) — the vDSO, the public system-call boundary.** A design note
|
- **[vdso.md](os-development-guide/vdso.md) — the vDSO, the public system-call boundary.** A design note
|
||||||
(not built yet) on keeping `abi.zig` genuinely private: a kernel-supplied, C-ABI
|
(not built yet) on keeping `abi.zig` genuinely private: a kernel-supplied, C-ABI
|
||||||
entry blob mapped into every process as the *only* way into the kernel — so the
|
entry blob mapped into every process as the *only* way into the kernel — so the
|
||||||
syscall numbers can be renumbered or randomised at will, and Rust/C binaries get a
|
syscall numbers can be renumbered or randomised at will, and Rust/C binaries get a
|
||||||
stable boundary without danos growing a dynamic linker. danos's public ABI = the
|
stable boundary without danos growing a dynamic linker. danos's public ABI = the
|
||||||
vDSO + the documented IPC wire protocols ([vfs-protocol.md](vfs-protocol.md) first).
|
vDSO + the documented IPC wire protocols ([vfs-protocol.md](file-system-development/vfs-protocol.md) first).
|
||||||
|
|
||||||
Cutting across all of these:
|
Cutting across all of these:
|
||||||
|
|
||||||
@@ -137,69 +137,69 @@ Cutting across all of these:
|
|||||||
hardware needed to run danos: minimum specs (UEFI x86-64, ACPI, PCIe ECAM,
|
hardware needed to run danos: minimum specs (UEFI x86-64, ACPI, PCIe ECAM,
|
||||||
xHCI, ~128 MiB RAM) grounded in what the boot path actually assumes, plus a
|
xHCI, ~128 MiB RAM) grounded in what the boot path actually assumes, plus a
|
||||||
plain-language guide matching Intel/AMD CPU generations by name.
|
plain-language guide matching Intel/AMD CPU generations by name.
|
||||||
- **[release-iso.md](release-iso.md) — the release ISO.** The flashable boot
|
- **[release-iso.md](os-development-guide/release-iso.md) — the release ISO.** The flashable boot
|
||||||
media: `zig build release-x86-64` wraps the FAT32 boot volume in a hybrid ISO
|
media: `zig build release-x86-64` wraps the FAT32 boot volume in a hybrid ISO
|
||||||
(MBR ESP partition + El Torito EFI entry, one embedded image) that Etcher/dd
|
(MBR ESP partition + El Torito EFI entry, one embedded image) that Etcher/dd
|
||||||
flash to USB or a burner writes to disc — built by an in-repo pure-Python
|
flash to USB or a burner writes to disc — built by an in-repo pure-Python
|
||||||
tool, like the FAT image itself.
|
tool, like the FAT image itself.
|
||||||
- **[architecture.md](architecture.md) — the architecture split.** How CPU-specific code is kept
|
- **[architecture.md](os-development-guide/architecture.md) — the architecture split.** How CPU-specific code is kept
|
||||||
behind a build-time `arch` module so the generic kernel never names x86_64,
|
behind a build-time `arch` module so the generic kernel never names x86_64,
|
||||||
leaving room for other systems (e.g. an AArch64 Raspberry Pi) later.
|
leaving room for other systems (e.g. an AArch64 Raspberry Pi) later.
|
||||||
- **[arm.md](arm.md) — ARM targets.** The Raspberry Pi landscape the arch split is
|
- **[arm.md](os-development-guide/arm.md) — ARM targets.** The Raspberry Pi landscape the arch split is
|
||||||
aiming at: `arm` (32-bit, Pi Zero W) vs `aarch64` (64-bit, Pi 3-5), UEFI vs
|
aiming at: `arm` (32-bit, Pi Zero W) vs `aarch64` (64-bit, Pi 3-5), UEFI vs
|
||||||
device-tree boot, and what each layer needs.
|
device-tree boot, and what each layer needs.
|
||||||
- **[discovery.md](discovery.md) — device discovery.** A design note on learning what
|
- **[discovery.md](os-development-guide/discovery.md) — device discovery.** A design note on learning what
|
||||||
hardware exists via ACPI (x86) or device tree (ARM) behind one neutral device model —
|
hardware exists via ACPI (x86) or device tree (ARM) behind one neutral device model —
|
||||||
when to build it, and how to keep it architecture-agnostic.
|
when to build it, and how to keep it architecture-agnostic.
|
||||||
- **[acpi.md](acpi.md) — finding the ACPI tables.** The concrete x86 locator chain:
|
- **[acpi.md](os-development-guide/acpi.md) — finding the ACPI tables.** The concrete x86 locator chain:
|
||||||
how the loader captures the **RSDP**, hands its physical address across in `BootInformation`,
|
how the loader captures the **RSDP**, hands its physical address across in `BootInformation`,
|
||||||
and how the platform derives the **RSDT/XSDT** from it and walks the SDTs — plus the
|
and how the platform derives the **RSDT/XSDT** from it and walks the SDTs — plus the
|
||||||
live event side (the SCI, the power button, GPE/Notify) the ring-3 acpi service runs.
|
live event side (the SCI, the power button, GPE/Notify) the ring-3 acpi service runs.
|
||||||
- **[power.md](power.md) — the power service.** System power as a domain-named
|
- **[power.md](os-development-guide/power.md) — the power service.** System power as a domain-named
|
||||||
service: button/lid/battery events published to subscribers, and init's orderly
|
service: button/lid/battery events published to subscribers, and init's orderly
|
||||||
shutdown composing the [lifecycle](process-lifecycle.md) stop sequence with an ACPI
|
shutdown composing the [lifecycle](os-development-guide/process-lifecycle.md) stop sequence with an ACPI
|
||||||
S5 write. Firmware-neutral — a PSCI backend drops in on ARM.
|
S5 write. Firmware-neutral — a PSCI backend drops in on ARM.
|
||||||
- **[timers.md](timers.md) — timers and time.** The ring-3 surface for reading the
|
- **[timers.md](os-development-guide/timers.md) — timers and time.** The ring-3 surface for reading the
|
||||||
clock and waiting: why `now()` is a syscall rather than a service, and the one-shot
|
clock and waiting: why `now()` is a syscall rather than a service, and the one-shot
|
||||||
timer notification (`timer_bind`) that gives supervisors a timed wait — built on the
|
timer notification (`timer_bind`) that gives supervisors a timed wait — built on the
|
||||||
LAPIC heartbeat and calibrated TSC of [device-interrupts.md](device-interrupts.md).
|
LAPIC heartbeat and calibrated TSC of [device-interrupts.md](device-driver-development-guide/device-interrupts.md).
|
||||||
- **[smp.md](smp.md) — multiple cores.** A design/research note on how microkernels
|
- **[smp.md](os-development-guide/smp.md) — multiple cores.** A design/research note on how microkernels
|
||||||
(L4, seL4) handle SMP — big kernel lock vs per-CPU vs multikernel — and how the
|
(L4, seL4) handle SMP — big kernel lock vs per-CPU vs multikernel — and how the
|
||||||
right choice depends on whether danos is chasing real-time or resilience.
|
right choice depends on whether danos is chasing real-time or resilience.
|
||||||
- **[coding-standards.md](coding-standards.md) — coding standards.** The naming rule the
|
- **[coding-standards.md](coding-standards.md) — coding standards.** The naming rule the
|
||||||
tree follows: non-acronyms are spelled out in full (`message`, not `msg`), files are
|
tree follows: non-acronyms are spelled out in full (`message`, not `msg`), files are
|
||||||
`kebab-case`, code follows Zig's case conventions, and the handful of exceptions
|
`kebab-case`, code follows Zig's case conventions, and the handful of exceptions
|
||||||
(POSIX/C ABI names, `init`/`len`/`ptr`, acronyms).
|
(POSIX/C ABI names, `init`/`len`/`ptr`, acronyms).
|
||||||
- **[sysv.md](sysv.md) — the calling convention.** What "the kernel is SysV" means,
|
- **[sysv.md](os-development-guide/sysv.md) — the calling convention.** What "the kernel is SysV" means,
|
||||||
and why the loader→kernel boundary has to pin it (the RDI-vs-RCX handoff).
|
and why the loader→kernel boundary has to pin it (the RDI-vs-RCX handoff).
|
||||||
- **[testing.md](testing.md) — testing.** How the kernel is tested by booting it in
|
- **[testing.md](testing.md) — testing.** How the kernel is tested by booting it in
|
||||||
QEMU and asserting on its serial output — reproducibly, and structured so the
|
QEMU and asserting on its serial output — reproducibly, and structured so the
|
||||||
same tests run across architectures.
|
same tests run across architectures.
|
||||||
- **[logging.md](logging.md) — logging.** The multi-sink diagnostic log (serial,
|
- **[logging.md](os-development-guide/logging.md) — logging.** The multi-sink diagnostic log (serial,
|
||||||
0xE9 debugcon, file later) kept separate from the framebuffer display, plus the
|
0xE9 debugcon, file later) kept separate from the framebuffer display, plus the
|
||||||
robustness path: optional framebuffer, POST-code checkpoints, and a persistent
|
robustness path: optional framebuffer, POST-code checkpoints, and a persistent
|
||||||
panic breadcrumb so the kernel survives — and can be diagnosed — with no output.
|
panic breadcrumb so the kernel survives — and can be diagnosed — with no output.
|
||||||
|
|
||||||
## How the pieces relate
|
## How the pieces relate
|
||||||
|
|
||||||
The boot flow ties them together: UEFI runs the loader ([efi.md](efi.md)), which
|
The boot flow ties them together: UEFI runs the loader ([efi.md](os-development-guide/efi.md)), which
|
||||||
queries the **GOP** to pick a graphics mode ([gop.md](gop.md)), hands the kernel a
|
queries the **GOP** to pick a graphics mode ([gop.md](os-development-guide/gop.md)), hands the kernel a
|
||||||
**framebuffer** to draw into ([framebuffer.md](framebuffer.md)) and a **memory
|
**framebuffer** to draw into ([framebuffer.md](os-development-guide/framebuffer.md)) and a **memory
|
||||||
map** of physical RAM ([memory-map.md](memory-map.md)); the kernel turns that map
|
map** of physical RAM ([memory-map.md](os-development-guide/memory-map.md)); the kernel turns that map
|
||||||
into a **frame allocator** ([frame-allocator.md](frame-allocator.md)), installs
|
into a **frame allocator** ([frame-allocator.md](os-development-guide/frame-allocator.md)), installs
|
||||||
its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.md)),
|
its **descriptor tables** so CPU faults are caught ([interrupts.md](os-development-guide/interrupts.md)),
|
||||||
builds its own **page tables** and switches onto them ([paging.md](paging.md)),
|
builds its own **page tables** and switches onto them ([paging.md](os-development-guide/paging.md)),
|
||||||
brings up the **heap** for dynamic allocation ([heap.md](heap.md)), starts the
|
brings up the **heap** for dynamic allocation ([heap.md](os-development-guide/heap.md)), starts the
|
||||||
**scheduler** ([scheduling.md](scheduling.md)) and the **timer** that preempts it
|
**scheduler** ([scheduling.md](os-development-guide/scheduling.md)) and the **timer** that preempts it
|
||||||
([device-interrupts.md](device-interrupts.md)) — with tasks blocking, sleeping and
|
([device-interrupts.md](device-driver-development-guide/device-interrupts.md)) — with tasks blocking, sleeping and
|
||||||
passing messages over **[IPC](ipc.md)** channels — runs, its CPU-specific bits
|
passing messages over **[IPC](device-driver-development-guide/ipc.md)** channels — runs, its CPU-specific bits
|
||||||
behind the [architecture](architecture.md) boundary, and when idle, or on a panic, it **halts**
|
behind the [architecture](os-development-guide/architecture.md) boundary, and when idle, or on a panic, it **halts**
|
||||||
([halting.md](halting.md)).
|
([halting.md](os-development-guide/halting.md)).
|
||||||
|
|
||||||
Above that line the microkernel proper begins: **discovery** ([discovery.md](discovery.md),
|
Above that line the microkernel proper begins: **discovery** ([discovery.md](os-development-guide/discovery.md),
|
||||||
[acpi.md](acpi.md)) learns what hardware exists, ring-3 processes ask the kernel for
|
[acpi.md](os-development-guide/acpi.md)) learns what hardware exists, ring-3 processes ask the kernel for
|
||||||
things through the small **[syscall](syscall.md)** table, isolated servers reach each
|
things through the small **[syscall](os-development-guide/syscall.md)** table, isolated servers reach each
|
||||||
other over IPC **endpoints** ([ipc.md](ipc.md)), and a **[driver](drivers.md)** claims
|
other over IPC **endpoints** ([ipc.md](device-driver-development-guide/ipc.md)), and a **[driver](device-driver-development-guide/drivers.md)** claims
|
||||||
a device, maps its registers, and sleeps until the hardware interrupts it — which is
|
a device, maps its registers, and sleeps until the hardware interrupts it — which is
|
||||||
the whole reason for the arrangement ([vision.md](vision.md)).
|
the whole reason for the arrangement ([vision.md](vision.md)).
|
||||||
|
|
||||||
@@ -208,7 +208,7 @@ the whole reason for the arrangement ([vision.md](vision.md)).
|
|||||||
danos is a **monorepo of sub-projects**. Each service or driver is a directory that is
|
danos is a **monorepo of sub-projects**. Each service or driver is a directory that is
|
||||||
its own Zig module — it can hold as many files as it needs, and other sub-projects
|
its own Zig module — it can hold as many files as it needs, and other sub-projects
|
||||||
reach it *by module name*, never by a path into its files. The source tree deliberately
|
reach it *by module name*, never by a path into its files. The source tree deliberately
|
||||||
**mirrors the runtime FHS** ([danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md)):
|
**mirrors the runtime FHS** ([danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md)):
|
||||||
what you see under `system/` in the source is what a running danos represents under
|
what you see under `system/` in the source is what a running danos represents under
|
||||||
`/system`.
|
`/system`.
|
||||||
|
|
||||||
|
|||||||
+10
-10
@@ -1,6 +1,6 @@
|
|||||||
# Device interrupts
|
# Device interrupts
|
||||||
|
|
||||||
CPU exceptions ([interrupts.md](interrupts.md)) are the kernel reacting to its own
|
CPU exceptions ([interrupts.md](../os-development-guide/interrupts.md)) are the kernel reacting to its own
|
||||||
mistakes. **Device interrupts** are the opposite: hardware asking for attention —
|
mistakes. **Device interrupts** are the opposite: hardware asking for attention —
|
||||||
a timer firing, a key pressed, a packet arriving. They share the IDT, but differ
|
a timer firing, a key pressed, a packet arriving. They share the IDT, but differ
|
||||||
in one fundamental way: an exception here is terminal (we report and halt), while a
|
in one fundamental way: an exception here is terminal (we report and halt), while a
|
||||||
@@ -10,7 +10,7 @@ back — the same mechanism a scheduler will later use to preempt tasks.
|
|||||||
|
|
||||||
The first device we bring up is the **timer**, because it's the simplest: it lives
|
The first device we bring up is the **timer**, because it's the simplest: it lives
|
||||||
entirely on the CPU's local interrupt controller, needing no external routing.
|
entirely on the CPU's local interrupt controller, needing no external routing.
|
||||||
It's all x86_64-specific, behind the [architecture](architecture.md) boundary.
|
It's all x86_64-specific, behind the [architecture](../os-development-guide/architecture.md) boundary.
|
||||||
|
|
||||||
## The APIC, not the PIC
|
## The APIC, not the PIC
|
||||||
|
|
||||||
@@ -40,7 +40,7 @@ count that becomes the reload value. From then on it fires vector 32 repeatedly,
|
|||||||
its own, forever.
|
its own, forever.
|
||||||
|
|
||||||
The reload count isn't picked arbitrarily — it's **calibrated to real time**,
|
The reload count isn't picked arbitrarily — it's **calibrated to real time**,
|
||||||
which the [real-time](vision.md) scheduling guarantees depend on. Since the LAPIC
|
which the [real-time](../vision.md) scheduling guarantees depend on. Since the LAPIC
|
||||||
timer's raw rate is bus-clock dependent and unknown up front, `calibrate` runs the
|
timer's raw rate is bus-clock dependent and unknown up front, `calibrate` runs the
|
||||||
LAPIC timer one-shot from its maximum count while a **reference clock** counts out a
|
LAPIC timer one-shot from its maximum count while a **reference clock** counts out a
|
||||||
known 10 ms, then sees how far the LAPIC got — its counts-per-millisecond, from which
|
known 10 ms, then sees how far the LAPIC got — its counts-per-millisecond, from which
|
||||||
@@ -53,7 +53,7 @@ a missing PIT would hang the boot):
|
|||||||
|
|
||||||
1. **CPUID leaf 0x15** — the CPU's TSC frequency directly, needing no external timer
|
1. **CPUID leaf 0x15** — the CPU's TSC frequency directly, needing no external timer
|
||||||
at all (the LAPIC is then measured against the TSC).
|
at all (the LAPIC is then measured against the TSC).
|
||||||
2. The **HPET**, discovered via ACPI (see [discovery](discovery.md) / [acpi](acpi.md)).
|
2. The **HPET**, discovered via ACPI (see [discovery](../os-development-guide/discovery.md) / [acpi](../os-development-guide/acpi.md)).
|
||||||
3. The **ACPI PM timer** (a fixed 3.579545 MHz counter from the FADT).
|
3. The **ACPI PM timer** (a fixed 3.579545 MHz counter from the FADT).
|
||||||
4. The **PIT** (legacy 8254, 1.193182 MHz) — last resort, and bounded so it can't hang.
|
4. The **PIT** (legacy 8254, 1.193182 MHz) — last resort, and bounded so it can't hang.
|
||||||
|
|
||||||
@@ -100,7 +100,7 @@ values (a second socket, some firmware), so a thread migrating from a core readi
|
|||||||
check** as each application processor comes online (`checkWarpSource`, adapted from
|
check** as each application processor comes online (`checkWarpSource`, adapted from
|
||||||
Linux's): the waking core and the BSP hammer a shared "highest seen" TSC under a lock,
|
Linux's): the waking core and the BSP hammer a shared "highest seen" TSC under a lock,
|
||||||
and if either ever reads below it, the cores' TSCs are skewed. It's pairwise because APs
|
and if either ever reads below it, the cores' TSCs are skewed. It's pairwise because APs
|
||||||
come up one at a time ([smp.md](smp.md)).
|
come up one at a time ([smp.md](../os-development-guide/smp.md)).
|
||||||
|
|
||||||
**The fallback.** When the TSC fails either test — non-invariant (a bare VM such as the
|
**The fallback.** When the TSC fails either test — non-invariant (a bare VM such as the
|
||||||
default qemu64), or warped between cores — danos moves the monotonic clock onto the
|
default qemu64), or warped between cores — danos moves the monotonic clock onto the
|
||||||
@@ -160,7 +160,7 @@ A device handler is a plain `fn () void` — a timer or keyboard handler doesn't
|
|||||||
the interrupted registers. (The stubs originally didn't save the SSE/vector
|
the interrupted registers. (The stubs originally didn't save the SSE/vector
|
||||||
registers, so a handler couldn't use them; `isr_common` now does an
|
registers, so a handler couldn't use them; `isr_common` now does an
|
||||||
`fxsave`/`fxrstor` of the full SSE/x87 state around dispatch — see
|
`fxsave`/`fxrstor` of the full SSE/x87 state around dispatch — see
|
||||||
[interrupts.md](interrupts.md).)
|
[interrupts.md](../os-development-guide/interrupts.md).)
|
||||||
|
|
||||||
## Turning them on
|
## Turning them on
|
||||||
|
|
||||||
@@ -168,11 +168,11 @@ Exceptions can't be masked, which is why they worked all along. Maskable device
|
|||||||
interrupts don't fire until the CPU's interrupt flag is set — so the final step is
|
interrupts don't fire until the CPU's interrupt flag is set — so the final step is
|
||||||
`sti` (`arch.enableInterrupts()`), after the APIC and timer are configured. From
|
`sti` (`arch.enableInterrupts()`), after the APIC and timer are configured. From
|
||||||
that instant the kernel has a heartbeat, and its idle `hlt` loop
|
that instant the kernel has a heartbeat, and its idle `hlt` loop
|
||||||
([halting.md](halting.md)) wakes on every tick and dozes off again.
|
([halting.md](../os-development-guide/halting.md)) wakes on every tick and dozes off again.
|
||||||
|
|
||||||
## Verifying it
|
## Verifying it
|
||||||
|
|
||||||
The `timer` test (see [testing.md](testing.md)) is the proof that an interrupt both
|
The `timer` test (see [testing.md](../testing.md)) is the proof that an interrupt both
|
||||||
*fires* and *returns*: it records the tick count, busy-waits, and checks the count
|
*fires* and *returns*: it records the tick count, busy-waits, and checks the count
|
||||||
advanced on its own.
|
advanced on its own.
|
||||||
|
|
||||||
@@ -188,13 +188,13 @@ spinning in unrelated code — is the whole mechanism working end to end.
|
|||||||
## Since (done elsewhere)
|
## Since (done elsewhere)
|
||||||
|
|
||||||
- **Preemption**: the timer handler is where the scheduler decides to switch — the
|
- **Preemption**: the timer handler is where the scheduler decides to switch — the
|
||||||
reason a *returning* interrupt matters. See [scheduling.md](scheduling.md).
|
reason a *returning* interrupt matters. See [scheduling.md](../os-development-guide/scheduling.md).
|
||||||
- **`sleep()` / timeouts** built on the calibrated clock.
|
- **`sleep()` / timeouts** built on the calibrated clock.
|
||||||
- **The I/O APIC, routed**: external device lines now reach a vector, and the
|
- **The I/O APIC, routed**: external device lines now reach a vector, and the
|
||||||
interrupt is delivered onward to a *user-space* driver as an IPC message. See
|
interrupt is delivered onward to a *user-space* driver as an IPC message. See
|
||||||
[drivers.md](drivers.md).
|
[drivers.md](drivers.md).
|
||||||
- **Uncacheable MMIO**: device grants are mapped `PCD|PWT` (strong-uncacheable) for
|
- **Uncacheable MMIO**: device grants are mapped `PCD|PWT` (strong-uncacheable) for
|
||||||
user drivers — see [paging.md](paging.md).
|
user drivers — see [paging.md](../os-development-guide/paging.md).
|
||||||
|
|
||||||
## What's next (partly done since)
|
## What's next (partly done since)
|
||||||
|
|
||||||
@@ -10,18 +10,18 @@ mirrors them and prunes a dead reporter's children, and the `usb-report`
|
|||||||
scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13):
|
scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13):
|
||||||
`enumerate` and `subscribe` over IPC, with `device-list` as the first client —
|
`enumerate` and `subscribe` over IPC, with `device-list` as the first client —
|
||||||
the manager is now the one answer to "what devices exist" for applications.
|
the manager is now the one answer to "what devices exist" for applications.
|
||||||
The primitives underneath are real ([process-management.md](process-management.md):
|
The primitives underneath are real ([process-management.md](../os-development-guide/process-management.md):
|
||||||
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
|
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
|
||||||
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
|
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
|
||||||
per-device driver spawn works (the device manager matches the xHCI controller by PCI
|
per-device driver spawn works (the device manager matches the xHCI controller by PCI
|
||||||
class and spawns `usb-xhci-bus` with the device id as argv[1]). This document designs
|
class and spawns `usb-xhci-bus` with the device id as argv[1]). This document designs
|
||||||
the rest: the device manager as **the tree, the matcher, and the supervisor** — the
|
the rest: the device manager as **the tree, the matcher, and the supervisor** — the
|
||||||
policy process that turns [resilience.md](resilience.md)'s restart goal into practice
|
policy process that turns [resilience.md](../os-development-guide/resilience.md)'s restart goal into practice
|
||||||
for drivers.
|
for drivers.
|
||||||
|
|
||||||
How processes stop, reload, and report their deaths is deliberately **not** in this
|
How processes stop, reload, and report their deaths is deliberately **not** in this
|
||||||
document: that is the universal lifecycle every danos process speaks —
|
document: that is the universal lifecycle every danos process speaks —
|
||||||
[process-lifecycle.md](process-lifecycle.md), signals over IPC and the stable
|
[process-lifecycle.md](../os-development-guide/process-lifecycle.md), signals over IPC and the stable
|
||||||
`process` interface. The device manager is that design's first serious
|
`process` interface. The device manager is that design's first serious
|
||||||
customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a
|
customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a
|
||||||
driver is stopped, health-checked, and buried exactly like any other process.
|
driver is stopped, health-checked, and buried exactly like any other process.
|
||||||
@@ -35,7 +35,7 @@ The device tree is two things fused: *information* (what exists, how it nests) a
|
|||||||
claims, resource containment on `device_register`, the
|
claims, resource containment on `device_register`, the
|
||||||
`mmio_map`/`irq_bind`/`msi_bind` gates — and **cleans all of it up when a process
|
`mmio_map`/`irq_bind`/`msi_bind` gates — and **cleans all of it up when a process
|
||||||
dies** (settled; it is increment 1 of
|
dies** (settled; it is increment 1 of
|
||||||
[process-lifecycle.md](process-lifecycle.md)). The three invariants in
|
[process-lifecycle.md](../os-development-guide/process-lifecycle.md)). The three invariants in
|
||||||
[driver-model.md](driver-model.md) stay exactly where they are. A device manager
|
[driver-model.md](driver-model.md) stay exactly where they are. A device manager
|
||||||
that could mint MMIO mappings by its own say-so would be a second kernel, and a
|
that could mint MMIO mappings by its own say-so would be a second kernel, and a
|
||||||
buggy one would un-earn everything the microkernel bought.
|
buggy one would un-earn everything the microkernel bought.
|
||||||
@@ -51,7 +51,7 @@ enumeration is a **pci-bus driver**: the manager spawns it against the host brid
|
|||||||
like any bus reports children. ACPI becomes an **acpi service** that interprets the
|
like any bus reports children. ACPI becomes an **acpi service** that interprets the
|
||||||
tables and reports the namespace. The manager only orchestrates and merges. Moving
|
tables and reports the namespace. The manager only orchestrates and merges. Moving
|
||||||
AML interpretation out of ring 0 is its own project on its own track; nothing here
|
AML interpretation out of ring 0 is its own project on its own track; nothing here
|
||||||
depends on when it lands. (It landed: [discovery.md](discovery.md), M19–M20.)
|
depends on when it lands. (It landed: [discovery.md](../os-development-guide/discovery.md), M19–M20.)
|
||||||
|
|
||||||
`device_register` is **idempotent on exact match**: a re-registration with an
|
`device_register` is **idempotent on exact match**: a re-registration with an
|
||||||
identical (parent, class, identity, resources) tuple returns the existing id
|
identical (parent, class, identity, resources) tuple returns the existing id
|
||||||
@@ -82,7 +82,7 @@ one world.
|
|||||||
deadline means wrong binary, wrong protocol version, or wedged before main — apply
|
deadline means wrong binary, wrong protocol version, or wedged before main — apply
|
||||||
the stop sequence and the restart policy. Everything else lifecycle-shaped
|
the stop sequence and the restart policy. Everything else lifecycle-shaped
|
||||||
(terminate, the common `ping` liveness call, exit reasons) arrives through
|
(terminate, the common `ping` liveness call, exit reasons) arrives through
|
||||||
[process-lifecycle.md](process-lifecycle.md)'s vocabulary, not this protocol.
|
[process-lifecycle.md](../os-development-guide/process-lifecycle.md)'s vocabulary, not this protocol.
|
||||||
|
|
||||||
Assignment stays argv (`usb-xhci-bus <device id>`) for now — simple, and it works.
|
Assignment stays argv (`usb-xhci-bus <device id>`) for now — simple, and it works.
|
||||||
The step after `hello` exists is delegation: the manager claims (or is granted) the
|
The step after `hello` exists is delegation: the manager claims (or is granted) the
|
||||||
@@ -97,7 +97,7 @@ from usb-ids.zig — each bus's native language, decoded by the shared ids modul
|
|||||||
Every driver is spawned with the manager's exit endpoint (`spawnSupervised` — built).
|
Every driver is spawned with the manager's exit endpoint (`spawnSupervised` — built).
|
||||||
On a death notification:
|
On a death notification:
|
||||||
|
|
||||||
1. **Read the reason** ([process-lifecycle.md](process-lifecycle.md) increment 2).
|
1. **Read the reason** ([process-lifecycle.md](../os-development-guide/process-lifecycle.md) increment 2).
|
||||||
Clean exit → it meant to; don't restart. Fault or missed `hello` deadline →
|
Clean exit → it meant to; don't restart. Fault or missed `hello` deadline →
|
||||||
restart with **backoff**, and a crash-loop cap (three fast deaths → mark failed,
|
restart with **backoff**, and a crash-loop cap (three fast deaths → mark failed,
|
||||||
stop respawning, log loudly; a later `reload` to the manager can retry).
|
stop respawning, log loudly; a later `reload` to the manager can retry).
|
||||||
@@ -136,7 +136,7 @@ way.
|
|||||||
## Increments
|
## Increments
|
||||||
|
|
||||||
Increments 1–4 are the lifecycle prerequisites and live in
|
Increments 1–4 are the lifecycle prerequisites and live in
|
||||||
[process-lifecycle.md](process-lifecycle.md) (claim cleanup on death, exit reasons,
|
[process-lifecycle.md](../os-development-guide/process-lifecycle.md) (claim cleanup on death, exit reasons,
|
||||||
published exit events, signals + `process`). On top of those:
|
published exit events, signals + `process`). On top of those:
|
||||||
|
|
||||||
5. **device-manager-protocol**: `hello`, supervised spawn with restart policy;
|
5. **device-manager-protocol**: `hello`, supervised spawn with restart policy;
|
||||||
@@ -147,7 +147,7 @@ published exit events, signals + `process`). On top of those:
|
|||||||
to a manager-internal seam.
|
to a manager-internal seam.
|
||||||
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): enumeration moved to
|
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): enumeration moved to
|
||||||
ring 3 as swappable per-firmware discoverers — the pci-bus driver (M19) then
|
ring 3 as swappable per-firmware discoverers — the pci-bus driver (M19) then
|
||||||
the acpi service (M20), see [discovery.md](discovery.md); of the enumerable
|
the acpi service (M20), see [discovery.md](../os-development-guide/discovery.md); of the enumerable
|
||||||
devices, the kernel seeds only the host bridge and the acpi-tables node (the
|
devices, the kernel seeds only the host bridge and the acpi-tables node (the
|
||||||
non-enumerable platform nodes — processors, interrupt controllers, the HPET,
|
non-enumerable platform nodes — processors, interrupt controllers, the HPET,
|
||||||
the loader's framebuffer — stay kernel-seeded too). Matching moved with it:
|
the loader's framebuffer — stay kernel-seeded too). Matching moved with it:
|
||||||
@@ -18,9 +18,9 @@ Read [display.md](display.md) first for the *why*; this is the *what* and the *o
|
|||||||
|
|
||||||
## Conventions
|
## Conventions
|
||||||
|
|
||||||
Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations in
|
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations in
|
||||||
full, kebab-case file names, no `Co-Authored-By` trailers on commits. New user binaries
|
full, kebab-case file names, no `Co-Authored-By` trailers on commits. New user binaries
|
||||||
go through `addUserBinary` in [build.zig](../build.zig) and get packed into the
|
go through `addUserBinary` in [build.zig](../../build.zig) and get packed into the
|
||||||
initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported into the
|
initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported into the
|
||||||
`runtime` module.
|
`runtime` module.
|
||||||
|
|
||||||
@@ -29,7 +29,7 @@ initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported i
|
|||||||
- `zig build test` — host unit tests (compositor math: layer clipping, damage merge,
|
- `zig build test` — host unit tests (compositor math: layer clipping, damage merge,
|
||||||
pitch/format blits are all host-testable with a fake framebuffer).
|
pitch/format blits are all host-testable with a fake framebuffer).
|
||||||
- `python3 test/qemu_test.py <case>` — boots the real kernel in QEMU; assert on the
|
- `python3 test/qemu_test.py <case>` — boots the real kernel in QEMU; assert on the
|
||||||
serial log ([tests.zig](../system/kernel/tests.zig) is the registry).
|
serial log ([tests.zig](../../system/kernel/tests.zig) is the registry).
|
||||||
- The `run-efi` target renders to QEMU's display (`-device VGA,edid=on,xres=1280,yres=720`)
|
- The `run-efi` target renders to QEMU's display (`-device VGA,edid=on,xres=1280,yres=720`)
|
||||||
— a screenshot confirms pixels for the milestones whose gate is visual.
|
— a screenshot confirms pixels for the milestones whose gate is visual.
|
||||||
|
|
||||||
@@ -39,22 +39,22 @@ initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported i
|
|||||||
|
|
||||||
Make the boot framebuffer reachable and mappable **write-combining** from user space.
|
Make the boot framebuffer reachable and mappable **write-combining** from user space.
|
||||||
|
|
||||||
- [x] [device-abi.zig](../library/device/model/device-abi.zig): added `DeviceClass.display`; a
|
- [x] [device-abi.zig](../../library/device/model/device-abi.zig): added `DeviceClass.display`; a
|
||||||
`DisplayInfo{ width, height, pitch, format }` carried on the descriptor; a
|
`DisplayInfo{ width, height, pitch, format }` carried on the descriptor; a
|
||||||
`flags` field on `ResourceDescriptor` + `resource_flag_write_combining`.
|
`flags` field on `ResourceDescriptor` + `resource_flag_write_combining`.
|
||||||
- [x] [devices-broker.zig](../system/kernel/devices-broker.zig): `seedDisplay(base, w, h,
|
- [x] [devices-broker.zig](../../system/kernel/devices-broker.zig): `seedDisplay(base, w, h,
|
||||||
pitch, format)` publishes a root-level `display` node with one WC-flagged `memory`
|
pitch, format)` publishes a root-level `display` node with one WC-flagged `memory`
|
||||||
resource `[base, height*pitch]` + the `DisplayInfo`; `displayDevice()` /
|
resource `[base, height*pitch]` + the `DisplayInfo`; `displayDevice()` /
|
||||||
`displayClaimed()`. Seeded from `kmain` after `devices_broker.init`.
|
`displayClaimed()`. Seeded from `kmain` after `devices_broker.init`.
|
||||||
- [x] [process.zig](../system/kernel/process.zig) `systemMmioMap` + paging
|
- [x] [process.zig](../../system/kernel/process.zig) `systemMmioMap` + paging
|
||||||
(`mapUserDeviceInto` gains a `write_combining` bool): a resource's WC flag maps it
|
(`mapUserDeviceInto` gains a `write_combining` bool): a resource's WC flag maps it
|
||||||
through the WC PAT slot (`setupPat`) instead of strong-uncacheable.
|
through the WC PAT slot (`setupPat`) instead of strong-uncacheable.
|
||||||
- [x] [console.zig](../system/kernel/console.zig): `setSuppressed` quiesces `write` while
|
- [x] [console.zig](../../system/kernel/console.zig): `setSuppressed` quiesces `write` while
|
||||||
the display device is claimed (driven from `systemDeviceClaim` / release); the
|
the display device is claimed (driven from `systemDeviceClaim` / release); the
|
||||||
terminal panic + exception paths clear it first so a dying machine still draws.
|
terminal panic + exception paths clear it first so a dying machine still draws.
|
||||||
|
|
||||||
**Gate (met, automated):** the `display` kernel test (`python3 test/qemu_test.py display`,
|
**Gate (met, automated):** the `display` kernel test (`python3 test/qemu_test.py display`,
|
||||||
`displayTest` in [tests.zig](../system/kernel/tests.zig)) asserts the seeded node's shape
|
`displayTest` in [tests.zig](../../system/kernel/tests.zig)) asserts the seeded node's shape
|
||||||
and geometry, then walks the real claim + `mmio_map` path into a throwaway address space
|
and geometry, then walks the real claim + `mmio_map` path into a throwaway address space
|
||||||
and verifies the leaf is **write-combining** (PAT entry 4: PAT bit set, PCD/PWT clear) —
|
and verifies the leaf is **write-combining** (PAT entry 4: PAT bit set, PCD/PWT clear) —
|
||||||
with an uncacheable-still-uncacheable regression guard. Chosen over the original
|
with an uncacheable-still-uncacheable regression guard. Chosen over the original
|
||||||
@@ -70,7 +70,7 @@ Stand up the named service and the double-buffer, no layers yet.
|
|||||||
- [x] `library/protocol/display/display-protocol.zig`: `Operation{ info, create_layer,
|
- [x] `library/protocol/display/display-protocol.zig`: `Operation{ info, create_layer,
|
||||||
configure_layer, destroy_layer, fill_rect, blit_tile, damage, present }`; `extern`
|
configure_layer, destroy_layer, fill_rect, blit_tile, damage, present }`; `extern`
|
||||||
`Request`/`Reply`; size + `maximum_payload` consts. (Model: block/protocol.zig.)
|
`Request`/`Reply`; size + `maximum_payload` consts. (Model: block/protocol.zig.)
|
||||||
- [x] [abi.zig](../system/abi.zig): `ServiceId.display = 9`.
|
- [x] [abi.zig](../../system/abi.zig): `ServiceId.display = 9`.
|
||||||
- [x] `system/services/display/display.zig`: `main` → enumerate + claim + WC-map the LFB
|
- [x] `system/services/display/display.zig`: `main` → enumerate + claim + WC-map the LFB
|
||||||
(front) → `mmap` a cacheable back buffer of `height*pitch` → `runtime.service.run`.
|
(front) → `mmap` a cacheable back buffer of `height*pitch` → `runtime.service.run`.
|
||||||
`info` and a whole-screen `present` (back → front) are live; layer ops fail-stub
|
`info` and a whole-screen `present` (back → front) are live; layer ops fail-stub
|
||||||
@@ -78,8 +78,8 @@ Stand up the named service and the double-buffer, no layers yet.
|
|||||||
- [x] [library/runtime/display.zig](../library/runtime/runtime.zig) (+ barrel export of
|
- [x] [library/runtime/display.zig](../library/runtime/runtime.zig) (+ barrel export of
|
||||||
`display` and `display_protocol`): `info()` and `present()`, cached `.display`
|
`display` and `display_protocol`): `info()` and `present()`, cached `.display`
|
||||||
lookup with retry (model: block.zig).
|
lookup with retry (model: block.zig).
|
||||||
- [x] [init.zig](../system/services/init/init.zig): `"display"` added to `boot_services`.
|
- [x] [init.zig](../../system/services/init/init.zig): `"display"` added to `boot_services`.
|
||||||
- [x] [build.zig](../build.zig): `display-protocol` module on the runtime; `display` exe
|
- [x] [build.zig](../../build.zig): `display-protocol` module on the runtime; `display` exe
|
||||||
via `addUserBinary`; packed into the initial-ramdisk; installed to
|
via `addUserBinary`; packed into the initial-ramdisk; installed to
|
||||||
`/system/services/display`.
|
`/system/services/display`.
|
||||||
- [x] **Kernel fix the back buffer surfaced:** `mmap` was capped at 256 pages (1 MiB) by
|
- [x] **Kernel fix the back buffer surfaced:** `mmap` was capped at 256 pages (1 MiB) by
|
||||||
@@ -105,7 +105,7 @@ The heart: composite an ordered layer stack, present only what changed.
|
|||||||
- [x] `create_layer` / `configure_layer` (damages old + new footprints) / `destroy_layer`,
|
- [x] `create_layer` / `configure_layer` (damages old + new footprints) / `destroy_layer`,
|
||||||
`fill_rect`, `blit_tile` (reads the inline tile from the IPC payload, unaligned-safe),
|
`fill_rect`, `blit_tile` (reads the inline tile from the IPC payload, unaligned-safe),
|
||||||
`damage`, `present`.
|
`damage`, `present`.
|
||||||
- [x] Pure, host-tested [compositor.zig](../system/services/display/compositor.zig): `Rect`
|
- [x] Pure, host-tested [compositor.zig](../../system/services/display/compositor.zig): `Rect`
|
||||||
(intersect/unite), `Surface`, `fillRect`, `composite` (opaque, clipped to a damage
|
(intersect/unite), `Surface`, `fillRect`, `composite` (opaque, clipped to a damage
|
||||||
rect), `blitTile`. `present` clears the damaged region to the wallpaper, paints the
|
rect), `blitTile`. `present` clears the damaged region to the wallpaper, paints the
|
||||||
visible layers bottom-to-top (z-sorted), and flushes just that rect back → front (WC).
|
visible layers bottom-to-top (z-sorted), and flushes just that rect back → front (WC).
|
||||||
@@ -148,10 +148,10 @@ still pass, and the default `zig build` is clean.
|
|||||||
- [x] The three integration cases exist and pass: `display` (D1 handoff, kernel),
|
- [x] The three integration cases exist and pass: `display` (D1 handoff, kernel),
|
||||||
`display-service` (D2/D3 compositor + self-check), and `display-demo` (D4 full
|
`display-service` (D2/D3 compositor + self-check), and `display-demo` (D4 full
|
||||||
pipeline: spawn `display` + `display-demo`, match `display-demo: ok`) —
|
pipeline: spawn `display` + `display-demo`, match `display-demo: ok`) —
|
||||||
[tests.zig](../system/kernel/tests.zig) + [qemu_test.py](../test/qemu_test.py). Plus
|
[tests.zig](../../system/kernel/tests.zig) + [qemu_test.py](../../test/qemu_test.py). Plus
|
||||||
the pure host tests (`zig build test`).
|
the pure host tests (`zig build test`).
|
||||||
- [x] [display.md](display.md) updated to the built state (the "Verifying it" section names
|
- [x] [display.md](display.md) updated to the built state (the "Verifying it" section names
|
||||||
the real cases); [README index](README.md) entry present (#19); the `display-track`
|
the real cases); [README index](../README.md) entry present (#19); the `display-track`
|
||||||
memory marked DONE with the commits.
|
memory marked DONE with the commits.
|
||||||
|
|
||||||
**Gate (met):** `python3 test/qemu_test.py display display-service display-demo` all pass,
|
**Gate (met):** `python3 test/qemu_test.py display display-service display-demo` all pass,
|
||||||
@@ -16,11 +16,11 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run,
|
|||||||
|
|
||||||
## Conventions
|
## Conventions
|
||||||
|
|
||||||
Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations,
|
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations,
|
||||||
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
|
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
|
||||||
`addUserBinary` and get packed into the initial-ramdisk; protocols are
|
`addUserBinary` and get packed into the initial-ramdisk; protocols are
|
||||||
`b.addModule("…-protocol", …)` imported into `runtime`; new syscalls extend
|
`b.addModule("…-protocol", …)` imported into `runtime`; new syscalls extend
|
||||||
[abi.zig](../system/abi.zig) `SystemCall` + a `library/runtime` wrapper.
|
[abi.zig](../../system/abi.zig) `SystemCall` + a `library/runtime` wrapper.
|
||||||
|
|
||||||
## How to verify along the way
|
## How to verify along the way
|
||||||
|
|
||||||
@@ -59,7 +59,7 @@ is the only backend), and `zig build test` stays green.
|
|||||||
|
|
||||||
## V2 — The shared-memory cross-process capability (kernel) ✅
|
## V2 — The shared-memory cross-process capability (kernel) ✅
|
||||||
|
|
||||||
- [x] [abi.zig](../system/abi.zig): `shared_memory_create` (34) / `shared_memory_map` (35) syscalls + a
|
- [x] [abi.zig](../../system/abi.zig): `shared_memory_create` (34) / `shared_memory_map` (35) syscalls + a
|
||||||
`shared_memory_test` service id. Handlers in process.zig: `shared_memory_create(len)` allocates contiguous,
|
`shared_memory_test` service id. Handlers in process.zig: `shared_memory_create(len)` allocates contiguous,
|
||||||
zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability
|
zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability
|
||||||
handle, maps them into the caller's shared-memory arena → returns virtual_address + handle; `shared_memory_map(cap)`
|
handle, maps them into the caller's shared-memory arena → returns virtual_address + handle; `shared_memory_map(cap)`
|
||||||
@@ -144,4 +144,4 @@ path in VMs**, where danos development happens. The framebuffer floor never goes
|
|||||||
- [display.md](display.md) — v1: the compositor, the GOP-vs-device split, the WC discipline.
|
- [display.md](display.md) — v1: the compositor, the GOP-vs-device split, the WC discipline.
|
||||||
- [display-v2-plan.md](display-v2-plan.md) — the ordered build-out.
|
- [display-v2-plan.md](display-v2-plan.md) — the ordered build-out.
|
||||||
- [driver-model.md](driver-model.md) — claim / `mmio_map` / MSI / capability passing (M13).
|
- [driver-model.md](driver-model.md) — claim / `mmio_map` / MSI / capability passing (M13).
|
||||||
- [resilience.md](resilience.md) — the restart machinery the hot-attach leans on.
|
- [resilience.md](../os-development-guide/resilience.md) — the restart machinery the hot-attach leans on.
|
||||||
@@ -1,12 +1,12 @@
|
|||||||
# The display service: a framebuffer compositor
|
# The display service: a framebuffer compositor
|
||||||
|
|
||||||
The [framebuffer](framebuffer.md) the loader hands over is a flat block of pixel
|
The [framebuffer](../os-development-guide/framebuffer.md) the loader hands over is a flat block of pixel
|
||||||
memory, and the kernel's [bootstrap console](../system/kernel/console.zig) draws text
|
memory, and the kernel's [bootstrap console](../../system/kernel/console.zig) draws text
|
||||||
into it directly. That console is a stop-gap. The **display service**
|
into it directly. That console is a stop-gap. The **display service**
|
||||||
(`system/services/display/`) is the real thing: an ordinary ring-3 process that *owns*
|
(`system/services/display/`) is the real thing: an ordinary ring-3 process that *owns*
|
||||||
the framebuffer, composes a stack of **layers** into an off-screen back buffer, and
|
the framebuffer, composes a stack of **layers** into an off-screen back buffer, and
|
||||||
**presents** finished frames to the screen — the display half of the GUI track
|
**presents** finished frames to the screen — the display half of the GUI track
|
||||||
([vision.md](vision.md)), the sibling of the [input service](input.md).
|
([vision.md](../vision.md)), the sibling of the [input service](input.md).
|
||||||
|
|
||||||
This note is the architecture and the reasoning behind it. The concrete build order
|
This note is the architecture and the reasoning behind it. The concrete build order
|
||||||
lives in [display-plan.md](display-plan.md).
|
lives in [display-plan.md](display-plan.md).
|
||||||
@@ -20,20 +20,20 @@ which one you're holding decides what you can do.
|
|||||||
|
|
||||||
- **GOP is firmware's *temporary* driver** for the display controller. It gives you a
|
- **GOP is firmware's *temporary* driver** for the display controller. It gives you a
|
||||||
linear framebuffer pointer and can set video modes — but only until
|
linear framebuffer pointer and can set video modes — but only until
|
||||||
`ExitBootServices`. The loader already leans on this: [`queryFramebuffer`](../boot/efi.zig)
|
`ExitBootServices`. The loader already leans on this: [`queryFramebuffer`](../../boot/efi.zig)
|
||||||
reads the monitor's EDID, picks the native mode, and calls `set_mode` **before**
|
reads the monitor's EDID, picks the native mode, and calls `set_mode` **before**
|
||||||
exiting ([gop.md](gop.md)). Once the kernel runs, GOP is **gone** — no `set_mode`, no
|
exiting ([gop.md](../os-development-guide/gop.md)). Once the kernel runs, GOP is **gone** — no `set_mode`, no
|
||||||
mode list, no EDID. What survives is the frozen snapshot in
|
mode list, no EDID. What survives is the frozen snapshot in
|
||||||
[`BootInformation.framebuffer`](../system/boot-handoff.zig): `{base, width, height,
|
[`BootInformation.framebuffer`](../../system/boot-handoff.zig): `{base, width, height,
|
||||||
pitch, format, refresh_hz}`, and nothing more.
|
pitch, format, refresh_hz}`, and nothing more.
|
||||||
|
|
||||||
- **The PCI class-0x03 device is the raw controller** — BARs, config space, registers,
|
- **The PCI class-0x03 device is the raw controller** — BARs, config space, registers,
|
||||||
IO ports. It is what you actually *own* after boot. On QEMU's emulated adapter
|
IO ports. It is what you actually *own* after boot. On QEMU's emulated adapter
|
||||||
([`-device VGA,edid=on`](../build.zig), the Bochs VBE/DISPI model) the `base` GOP handed
|
([`-device VGA,edid=on`](../../build.zig), the Bochs VBE/DISPI model) the `base` GOP handed
|
||||||
you *is* that device's linear-framebuffer BAR — the same physical memory, seen through
|
you *is* that device's linear-framebuffer BAR — the same physical memory, seen through
|
||||||
a different door. On a real discrete GPU, GOP's `base` is an aperture inside the GPU's
|
a different door. On a real discrete GPU, GOP's `base` is an aperture inside the GPU's
|
||||||
VRAM BAR. danos already decodes this device
|
VRAM BAR. danos already decodes this device
|
||||||
([pci-class.zig](../library/device/pci/pci-class.zig) has the full `display` namespace, and
|
([pci-class.zig](../../library/device/pci/pci-class.zig) has the full `display` namespace, and
|
||||||
`pci-bus` already reports it to the [device manager](device-manager.md) with its class
|
`pci-bus` already reports it to the [device manager](device-manager.md) with its class
|
||||||
triple) — but nothing binds it yet.
|
triple) — but nothing binds it yet.
|
||||||
|
|
||||||
@@ -63,16 +63,16 @@ rest of the system hasn't had to face:
|
|||||||
|
|
||||||
1. **The framebuffer is kernel-only today.** It arrives through the boot handoff, is
|
1. **The framebuffer is kernel-only today.** It arrives through the boot handoff, is
|
||||||
mapped into the kernel's physmap, and is touched only by
|
mapped into the kernel's physmap, and is touched only by
|
||||||
[`console.zig`](../system/kernel/console.zig). It is *not* a
|
[`console.zig`](../../system/kernel/console.zig). It is *not* a
|
||||||
[devices-broker](../system/kernel/devices-broker.zig) node, so
|
[devices-broker](../../system/kernel/devices-broker.zig) node, so
|
||||||
`device.claim`/`mmio_map` cannot reach it, and there is no framebuffer
|
`device.claim`/`mmio_map` cannot reach it, and there is no framebuffer
|
||||||
[syscall](syscall.md). A user-space display service needs a **new mechanism just to
|
[syscall](../os-development-guide/syscall.md). A user-space display service needs a **new mechanism just to
|
||||||
touch the pixels**. (See "The handoff" below — this is built.)
|
touch the pixels**. (See "The handoff" below — this is built.)
|
||||||
|
|
||||||
2. **danos had no cross-process shared memory.** At v1 the memory syscalls were `mmap`
|
2. **danos had no cross-process shared memory.** At v1 the memory syscalls were `mmap`
|
||||||
(private, zeroed), `mmio_map` (a *claimed device's* MMIO), and `dma_alloc` (new
|
(private, zeroed), `mmio_map` (a *claimed device's* MMIO), and `dma_alloc` (new
|
||||||
pinned physical). The block driver's "pass a buffer by physical address" trick
|
pinned physical). The block driver's "pass a buffer by physical address" trick
|
||||||
([block/protocol.zig](../library/protocol/block/block-protocol.zig)) works *only because its
|
([block/protocol.zig](../../library/protocol/block/block-protocol.zig)) works *only because its
|
||||||
consumer is DMA hardware*. A compositor that CPU-reads and blends client layers can't
|
consumer is DMA hardware*. A compositor that CPU-reads and blends client layers can't
|
||||||
use it — it would have to *map* another process's memory, which nothing allowed. v1
|
use it — it would have to *map* another process's memory, which nothing allowed. v1
|
||||||
sidesteps it entirely (see "What v1 does not do"); v2 has since built the primitive
|
sidesteps it entirely (see "What v1 does not do"); v2 has since built the primitive
|
||||||
@@ -103,10 +103,10 @@ rest of the system hasn't had to face:
|
|||||||
```
|
```
|
||||||
|
|
||||||
The bring-up sequence mirrors a hardware driver's — it is the
|
The bring-up sequence mirrors a hardware driver's — it is the
|
||||||
[`usb-xhci-bus` `initialise`](../system/drivers/usb-xhci-bus/usb-xhci-bus.zig) shape
|
[`usb-xhci-bus` `initialise`](../../system/drivers/usb-xhci-bus/usb-xhci-bus.zig) shape
|
||||||
(claim → `mmio_map` → run loop) — and the request/reply service shell is the
|
(claim → `mmio_map` → run loop) — and the request/reply service shell is the
|
||||||
[FAT](../system/services/fat/fat.zig) / [input](../system/services/input/input.zig) shape
|
[FAT](../../system/services/fat/fat.zig) / [input](../../system/services/input/input.zig) shape
|
||||||
([`service.run`](../library/kernel/service.zig) with a `protocol.zig` of
|
([`service.run`](../../library/kernel/service.zig) with a `protocol.zig` of
|
||||||
`extern struct` messages and an `Operation` tag).
|
`extern struct` messages and an `Operation` tag).
|
||||||
|
|
||||||
**One process, for now.** v1 is a *single* service that both owns the framebuffer and
|
**One process, for now.** v1 is a *single* service that both owns the framebuffer and
|
||||||
@@ -119,12 +119,12 @@ second backend or a second monitor appears; until then it is complexity with no
|
|||||||
|
|
||||||
The framebuffer crosses into user space through the machinery that already exists for
|
The framebuffer crosses into user space through the machinery that already exists for
|
||||||
every other device, rather than a bespoke syscall — so it inherits ownership,
|
every other device, rather than a bespoke syscall — so it inherits ownership,
|
||||||
release-on-death, and re-claim-on-restart for free (the [resilience](resilience.md)
|
release-on-death, and re-claim-on-restart for free (the [resilience](../os-development-guide/resilience.md)
|
||||||
story: a crashed display service returns the LFB to the kernel, and its restart
|
story: a crashed display service returns the LFB to the kernel, and its restart
|
||||||
re-claims it).
|
re-claims it).
|
||||||
|
|
||||||
- The kernel seeds a synthetic **display-class** node into the
|
- The kernel seeds a synthetic **display-class** node into the
|
||||||
[devices-broker](../system/kernel/devices-broker.zig) at init (`seedDisplay`), from
|
[devices-broker](../../system/kernel/devices-broker.zig) at init (`seedDisplay`), from
|
||||||
`BootInformation.framebuffer`: one `ResourceKind.memory` resource spanning
|
`BootInformation.framebuffer`: one `ResourceKind.memory` resource spanning
|
||||||
`[base, height*pitch]`, tagged **write-combining**, plus a small
|
`[base, height*pitch]`, tagged **write-combining**, plus a small
|
||||||
`DisplayInfo{width, height, pitch, format, refresh_hz}` (the memory resource says *where*
|
`DisplayInfo{width, height, pitch, format, refresh_hz}` (the memory resource says *where*
|
||||||
@@ -135,7 +135,7 @@ re-claims it).
|
|||||||
- The service `device.claim`s it and `mmio_map`s the resource. The map is
|
- The service `device.claim`s it and `mmio_map`s the resource. The map is
|
||||||
**write-combining**, not the strong-uncacheable that `mmio_map` uses for register
|
**write-combining**, not the strong-uncacheable that `mmio_map` uses for register
|
||||||
MMIO. The kernel already programs a WC PAT slot for its own console
|
MMIO. The kernel already programs a WC PAT slot for its own console
|
||||||
([`setupPat`](../system/kernel/architecture/x86_64/paging.zig)); this reaches it from
|
([`setupPat`](../../system/kernel/architecture/x86_64/paging.zig)); this reaches it from
|
||||||
the user mapping path. **This matters:** an uncacheable framebuffer makes the
|
the user mapping path. **This matters:** an uncacheable framebuffer makes the
|
||||||
back→front blit unusably slow.
|
back→front blit unusably slow.
|
||||||
- On `claim`, the kernel's bootstrap console goes quiet, so the two never fight over the
|
- On `claim`, the kernel's bootstrap console goes quiet, so the two never fight over the
|
||||||
@@ -143,7 +143,7 @@ re-claims it).
|
|||||||
panic on screen wins.
|
panic on screen wins.
|
||||||
|
|
||||||
The display service is a **named boot service**: `init` spawns it by name alongside
|
The display service is a **named boot service**: `init` spawns it by name alongside
|
||||||
`input`/`device-manager`/`fat` ([init.zig](../system/services/init/init.zig)), and it
|
`input`/`device-manager`/`fat` ([init.zig](../../system/services/init/init.zig)), and it
|
||||||
self-discovers the display node with `device.enumerate` (matching on `DeviceClass.display`). The [device manager](device-manager.md)
|
self-discovers the display node with `device.enumerate` (matching on `DeviceClass.display`). The [device manager](device-manager.md)
|
||||||
matching path (PCI class 0x03 → a driver) is reserved for the future *native* backend, not
|
matching path (PCI class 0x03 → a driver) is reserved for the future *native* backend, not
|
||||||
this singleton synthetic node.
|
this singleton synthetic node.
|
||||||
@@ -160,8 +160,8 @@ Two buffers, with deliberately different memory types:
|
|||||||
|
|
||||||
So a frame is: compose every dirty layer into the cacheable back buffer, then **present**
|
So a frame is: compose every dirty layer into the cacheable back buffer, then **present**
|
||||||
— copy the changed regions back→front in sequential, WC-friendly writes. Two details the
|
— copy the changed regions back→front in sequential, WC-friendly writes. Two details the
|
||||||
[framebuffer](framebuffer.md) note already establishes carry over: step rows by `pitch`,
|
[framebuffer](../os-development-guide/framebuffer.md) note already establishes carry over: step rows by `pitch`,
|
||||||
not `width*4`; and handle both `rgbx` and `bgrx` [pixel formats](gop.md).
|
not `width*4`; and handle both `rgbx` and `bgrx` [pixel formats](../os-development-guide/gop.md).
|
||||||
|
|
||||||
## Flicker vs. tearing — what double buffering does and doesn't buy
|
## Flicker vs. tearing — what double buffering does and doesn't buy
|
||||||
|
|
||||||
@@ -188,10 +188,10 @@ The compositor holds an **ordered stack of layers**. Each layer has a rectangle,
|
|||||||
z-order, a visibility flag, and a surface. Presenting walks the stack bottom-to-top,
|
z-order, a visibility flag, and a surface. Presenting walks the stack bottom-to-top,
|
||||||
painting each dirty layer into the back buffer, then flushes the damage to the front.
|
painting each dirty layer into the back buffer, then flushes the damage to the front.
|
||||||
Damage is tracked by one of two interchangeable trackers behind a compile-time
|
Damage is tracked by one of two interchangeable trackers behind a compile-time
|
||||||
`damage_mode` A/B switch ([display.zig](../system/services/display/display.zig)): a
|
`damage_mode` A/B switch ([display.zig](../../system/services/display/display.zig)): a
|
||||||
free-form dirty-rectangle **list** (tight bounds, heuristic merging) or a fixed 64-px
|
free-form dirty-rectangle **list** (tight bounds, heuristic merging) or a fixed 64-px
|
||||||
**tile grid** (exact O(1) merging, tile-quantized repaints) — the grid is the default;
|
**tile grid** (exact O(1) merging, tile-quantized repaints) — the grid is the default;
|
||||||
[compositor.zig](../system/services/display/compositor.zig) has both, with the trade-off
|
[compositor.zig](../../system/services/display/compositor.zig) has both, with the trade-off
|
||||||
discussion.
|
discussion.
|
||||||
|
|
||||||
In v1 the surfaces are **server-owned**, and clients draw into them with a small
|
In v1 the surfaces are **server-owned**, and clients draw into them with a small
|
||||||
@@ -210,7 +210,7 @@ shell, a terminal, a cursor, and a wallpaper:
|
|||||||
| `present` | request a repaint: composited at the next frame-clock tick |
|
| `present` | request a repaint: composited at the next frame-clock tick |
|
||||||
|
|
||||||
Text is intentionally *not* an operation — a client renders glyphs by blitting tiles
|
Text is intentionally *not* an operation — a client renders glyphs by blitting tiles
|
||||||
(the [PSF font](../system/kernel/font.psf) path the console already uses can move into a
|
(the [PSF font](../../system/kernel/font.psf) path the console already uses can move into a
|
||||||
client). Keeping the protocol to rectangles and tiles keeps the compositor small and the
|
client). Keeping the protocol to rectangles and tiles keeps the compositor small and the
|
||||||
policy in the client.
|
policy in the client.
|
||||||
|
|
||||||
@@ -224,8 +224,8 @@ pixels on screen synchronously (initialisation, the self-checks) bypass the cloc
|
|||||||
|
|
||||||
## `display`
|
## `display`
|
||||||
|
|
||||||
Clients speak the protocol through a new [`library/client/display/display.zig`](../library/client/display/display.zig),
|
Clients speak the protocol through a new [`library/client/display/display.zig`](../../library/client/display/display.zig),
|
||||||
the [`block`](../library/device/block/block.zig) shape (a cached `.display` lookup
|
the [`block`](../../library/device/block/block.zig) shape (a cached `.display` lookup
|
||||||
with a boot-race retry): `display.info()`, a `Layer` handle with `fill` / `blitTile` /
|
with a boot-race retry): `display.info()`, a `Layer` handle with `fill` / `blitTile` /
|
||||||
`damage`, and `present()`. Application code never issues the raw syscalls — it calls the
|
`damage`, and `present()`. Application code never issues the raw syscalls — it calls the
|
||||||
client module, as with every other danos service.
|
client module, as with every other danos service.
|
||||||
@@ -234,7 +234,7 @@ client module, as with every other danos service.
|
|||||||
|
|
||||||
The compositor is the single owner of the framebuffer — only the main `service.run` loop
|
The compositor is the single owner of the framebuffer — only the main `service.run` loop
|
||||||
touches the backend and the layer stack. Tracking the mouse without breaking that
|
touches the backend and the layer stack. Tracking the mouse without breaking that
|
||||||
ownership is the display's first use of [threads](threading.md): the service is built
|
ownership is the display's first use of [threads](../os-development-guide/threading.md): the service is built
|
||||||
multi-threaded (`addThreadedUserBinary`) and, at startup, spawns a **mouse-listener
|
multi-threaded (`addThreadedUserBinary`) and, at startup, spawns a **mouse-listener
|
||||||
thread** beside the compositor loop.
|
thread** beside the compositor loop.
|
||||||
|
|
||||||
@@ -242,7 +242,7 @@ thread** beside the compositor loop.
|
|||||||
(`input.subscribeMouse()`), accumulates the relative `dx`/`dy` motion into an absolute
|
(`input.subscribeMouse()`), accumulates the relative `dx`/`dy` motion into an absolute
|
||||||
cursor position clamped to the screen, and hands it to the compositor. It never touches
|
cursor position clamped to the screen, and hands it to the compositor. It never touches
|
||||||
the compositor — so no lock guards the framebuffer. A parked `next()` leaves its core
|
the compositor — so no lock guards the framebuffer. A parked `next()` leaves its core
|
||||||
free to halt ([halting.md](halting.md)).
|
free to halt ([halting.md](../os-development-guide/halting.md)).
|
||||||
- **The channel.** A single-slot *latest-value* cell (`CursorChannel`) guarded by a
|
- **The channel.** A single-slot *latest-value* cell (`CursorChannel`) guarded by a
|
||||||
`Thread.Mutex`: the renderer wants where the cursor *is now*, not a replay of
|
`Thread.Mutex`: the renderer wants where the cursor *is now*, not a replay of
|
||||||
every delta, so a new position overwrites the old. The listener also **pokes** the
|
every delta, so a new position overwrites the old. The listener also **pokes** the
|
||||||
@@ -254,13 +254,13 @@ thread** beside the compositor loop.
|
|||||||
which is just a top-z compositor layer — with the existing `configure` + `present` path
|
which is just a top-z compositor layer — with the existing `configure` + `present` path
|
||||||
(it damages the old and new footprints, so only those two rectangles repaint).
|
(it damages the old and new footprints, so only those two rectangles repaint).
|
||||||
|
|
||||||
Two threading facts shape this (both in [threading.md](threading.md)). IPC **handles do
|
Two threading facts shape this (both in [threading.md](../os-development-guide/threading.md)). IPC **handles do
|
||||||
not cross threads**, so the listener can't reuse the main loop's endpoint handle — it
|
not cross threads**, so the listener can't reuse the main loop's endpoint handle — it
|
||||||
`ipc.lookup(.display)`s its *own* handle to the same endpoint to poke through. And a
|
`ipc.lookup(.display)`s its *own* handle to the same endpoint to poke through. And a
|
||||||
multi-threaded service doing concurrent IPC is why the kernel's endpoint-create / register
|
multi-threaded service doing concurrent IPC is why the kernel's endpoint-create / register
|
||||||
/ lookup syscalls now serialize under the big kernel lock. Shared fate applies: a fault in
|
/ lookup syscalls now serialize under the big kernel lock. Shared fate applies: a fault in
|
||||||
the listener takes the whole display down, and the supervisor restarts the process
|
the listener takes the whole display down, and the supervisor restarts the process
|
||||||
([resilience.md](resilience.md)).
|
([resilience.md](../os-development-guide/resilience.md)).
|
||||||
|
|
||||||
## What v1 does not do (and why that's fine)
|
## What v1 does not do (and why that's fine)
|
||||||
|
|
||||||
@@ -284,7 +284,7 @@ both are clean additions behind the interfaces v1 establishes.
|
|||||||
|
|
||||||
## Verifying it
|
## Verifying it
|
||||||
|
|
||||||
Four QEMU test cases ([tests.zig](../system/kernel/tests.zig), `python3
|
Four QEMU test cases ([tests.zig](../../system/kernel/tests.zig), `python3
|
||||||
test/qemu_test.py <case>`), each layering on the last:
|
test/qemu_test.py <case>`), each layering on the last:
|
||||||
|
|
||||||
- **`display`** — the kernel handoff: the seeded `display` device is shaped correctly and
|
- **`display`** — the kernel handoff: the seeded `display` device is shaped correctly and
|
||||||
@@ -317,8 +317,8 @@ packing are additionally covered by pure host unit tests under `zig build test`.
|
|||||||
|
|
||||||
## See also
|
## See also
|
||||||
|
|
||||||
- [framebuffer.md](framebuffer.md) — the linear framebuffer, pitch vs. width, `volatile`.
|
- [framebuffer.md](../os-development-guide/framebuffer.md) — the linear framebuffer, pitch vs. width, `volatile`.
|
||||||
- [gop.md](gop.md) — GOP, and why only linear RGBX/BGRX modes are paintable.
|
- [gop.md](../os-development-guide/gop.md) — GOP, and why only linear RGBX/BGRX modes are paintable.
|
||||||
- [input.md](input.md) — the sibling service; the async `ipc_send` fan-out.
|
- [input.md](input.md) — the sibling service; the async `ipc_send` fan-out.
|
||||||
- [driver-model.md](driver-model.md) — claim / `mmio_map`, capability passing, the trust model.
|
- [driver-model.md](driver-model.md) — claim / `mmio_map`, capability passing, the trust model.
|
||||||
- [device-manager.md](device-manager.md) — matching and supervision (the native backend's route).
|
- [device-manager.md](device-manager.md) — matching and supervision (the native backend's route).
|
||||||
@@ -34,7 +34,7 @@ plain bus driver with no controller — a USB hub — is also a real thing.
|
|||||||
|
|
||||||
danos already has the right central structure. `system/kernel/devices-broker.zig` holds a table of
|
danos already has the right central structure. `system/kernel/devices-broker.zig` holds a table of
|
||||||
`DeviceDescriptor`, each with a parent, a class, and a set of resources. Firmware discovery
|
`DeviceDescriptor`, each with a parent, a class, and a set of resources. Firmware discovery
|
||||||
seeds it ([discovery.md](discovery.md)); `device_register` grows it.
|
seeds it ([discovery.md](../os-development-guide/discovery.md)); `device_register` grows it.
|
||||||
|
|
||||||
Three invariants make it a capability system rather than a directory:
|
Three invariants make it a capability system rather than a directory:
|
||||||
|
|
||||||
@@ -97,7 +97,7 @@ A "family" is two modules, not one:
|
|||||||
|
|
||||||
danos already has one of each: `library/device/pci/pci.zig` is a logic module (the
|
danos already has one of each: `library/device/pci/pci.zig` is a logic module (the
|
||||||
`Function` view of a claimed PCI function),
|
`Function` view of a claimed PCI function),
|
||||||
[`library/protocol/vfs/vfs-protocol.zig`](../library/protocol/vfs/vfs-protocol.zig) is a
|
[`library/protocol/vfs/vfs-protocol.zig`](../../library/protocol/vfs/vfs-protocol.zig) is a
|
||||||
protocol module shared by the mount backends (today the fat server) and their clients.
|
protocol module shared by the mount backends (today the fat server) and their clients.
|
||||||
(The user-space VFS server it was originally written against has since retired — path
|
(The user-space VFS server it was originally written against has since retired — path
|
||||||
routing moved into the kernel, `system/kernel/vfs.zig`'s `fs_resolve` — but the protocol
|
routing moved into the kernel, `system/kernel/vfs.zig`'s `fs_resolve` — but the protocol
|
||||||
@@ -199,7 +199,7 @@ class driver, the device manager, or the kernel may share them freely.
|
|||||||
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
|
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
|
||||||
fresh ring-3 process; `name` becomes the child's argv[0] and the optional
|
fresh ring-3 process; `name` becomes the child's argv[0] and the optional
|
||||||
NUL-separated `arguments` blob its argv[1..], delivered on a SysV entry stack
|
NUL-separated `arguments` blob its argv[1..], delivered on a SysV entry stack
|
||||||
([sysv.md](sysv.md)). This is what
|
([sysv.md](../os-development-guide/sysv.md)). This is what
|
||||||
turned the device manager from "log the match" into "run the driver": the kernel now
|
turned the device manager from "log the match" into "run the driver": the kernel now
|
||||||
spawns only `init`, `init` spawns the services, and the **device-manager** discovers
|
spawns only `init`, `init` spawns the services, and the **device-manager** discovers
|
||||||
the hardware and spawns each driver ([drivers.md](drivers.md)). Ungated for now — a
|
the hardware and spawns each driver ([drivers.md](drivers.md)). Ungated for now — a
|
||||||
@@ -321,7 +321,7 @@ sprinkling of `asm volatile`:
|
|||||||
| DMA cache coherency | coherent; nothing to do | **not guaranteed**; needs non-cacheable buffers or cache maintenance |
|
| DMA cache coherency | coherent; nothing to do | **not guaranteed**; needs non-cacheable buffers or cache maintenance |
|
||||||
|
|
||||||
x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away
|
x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away
|
||||||
with a compiler barrier alone. ARM is not, and [vision.md](vision.md) makes ARM the win
|
with a compiler barrier alone. ARM is not, and [vision.md](../vision.md) makes ARM the win
|
||||||
condition. Build the abstraction while there is one caller to fix.
|
condition. Build the abstraction while there is one caller to fix.
|
||||||
|
|
||||||
(Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full
|
(Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full
|
||||||
@@ -398,6 +398,6 @@ from hand-rolling `*volatile` and getting ARM wrong.
|
|||||||
## See also
|
## See also
|
||||||
|
|
||||||
- [drivers.md](drivers.md) — how to write one, concretely.
|
- [drivers.md](drivers.md) — how to write one, concretely.
|
||||||
- [discovery.md](discovery.md) / [acpi.md](acpi.md) — where the device table comes from.
|
- [discovery.md](../os-development-guide/discovery.md) / [acpi.md](../os-development-guide/acpi.md) — where the device table comes from.
|
||||||
- [ipc.md](ipc.md) — endpoints, badges, and the notification path an IRQ arrives on.
|
- [ipc.md](ipc.md) — endpoints, badges, and the notification path an IRQ arrives on.
|
||||||
- [resilience.md](resilience.md) — restart, the reason any of this is worth the trouble.
|
- [resilience.md](../os-development-guide/resilience.md) — restart, the reason any of this is worth the trouble.
|
||||||
@@ -4,13 +4,13 @@ In a monolithic kernel a driver is a function call away from everything: it runs
|
|||||||
ring 0, dereferences any physical address, and its interrupt handler *is* the ISR. In
|
ring 0, dereferences any physical address, and its interrupt handler *is* the ISR. In
|
||||||
danos a driver is **an ordinary ring-3 process**. It has its own address space, it
|
danos a driver is **an ordinary ring-3 process**. It has its own address space, it
|
||||||
can crash without taking the kernel with it, and — the point of this document — it
|
can crash without taking the kernel with it, and — the point of this document — it
|
||||||
can be restarted ([resilience](resilience.md)).
|
can be restarted ([resilience](../os-development-guide/resilience.md)).
|
||||||
|
|
||||||
That leaves three questions the kernel has to answer, because a process can't answer
|
That leaves three questions the kernel has to answer, because a process can't answer
|
||||||
them for itself:
|
them for itself:
|
||||||
|
|
||||||
1. **What hardware exists?** → `device_enumerate`, over the device table discovery built
|
1. **What hardware exists?** → `device_enumerate`, over the device table discovery built
|
||||||
([discovery](discovery.md), [acpi](acpi.md)).
|
([discovery](../os-development-guide/discovery.md), [acpi](../os-development-guide/acpi.md)).
|
||||||
2. **How do I touch its registers?** → `device_claim` + `mmio_map`: the kernel maps the
|
2. **How do I touch its registers?** → `device_claim` + `mmio_map`: the kernel maps the
|
||||||
device's physical MMIO window into your address space, and from then on it's plain
|
device's physical MMIO window into your address space, and from then on it's plain
|
||||||
memory. No syscall per register access.
|
memory. No syscall per register access.
|
||||||
@@ -51,7 +51,7 @@ the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Every
|
|||||||
is the **driver supervisor**. It does the three steps a monolithic kernel would do in
|
is the **driver supervisor**. It does the three steps a monolithic kernel would do in
|
||||||
its probe path, entirely from ring 3:
|
its probe path, entirely from ring 3:
|
||||||
1. **Discover** — `device_enumerate` snapshots the device table the kernel built from
|
1. **Discover** — `device_enumerate` snapshots the device table the kernel built from
|
||||||
ACPI/PCI ([discovery](discovery.md)).
|
ACPI/PCI ([discovery](../os-development-guide/discovery.md)).
|
||||||
2. **Match** — for each device it looks up a driver. The match policy is code, a few
|
2. **Match** — for each device it looks up a driver. The match policy is code, a few
|
||||||
small per-bus tables: from the boot snapshot only the PCI host bridge matches
|
small per-bus tables: from the boot snapshot only the PCI host bridge matches
|
||||||
(→ `pci-bus`); everything else arrives later as bus reports and matches on
|
(→ `pci-bus`); everything else arrives later as bus reports and matches on
|
||||||
@@ -269,7 +269,7 @@ already owns.
|
|||||||
A device with **no resources** is legal and common. A USB device is reached through its
|
A device with **no resources** is legal and common. A USB device is reached through its
|
||||||
controller, not by MMIO, so it gets `resource_count = 0`.
|
controller, not by MMIO, so it gets `resource_count = 0`.
|
||||||
|
|
||||||
See [`system/drivers/pci-bus/pci-bus.zig`](../system/drivers/pci-bus/pci-bus.zig) for a
|
See [`system/drivers/pci-bus/pci-bus.zig`](../../system/drivers/pci-bus/pci-bus.zig) for a
|
||||||
real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function
|
real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function
|
||||||
it finds as a child — and [driver-model.md](driver-model.md) for how bus drivers, class
|
it finds as a child — and [driver-model.md](driver-model.md) for how bus drivers, class
|
||||||
drivers and host controller drivers fit together.
|
drivers and host controller drivers fit together.
|
||||||
@@ -293,7 +293,7 @@ the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so
|
|||||||
uncacheable, physical address exposed), **memory barriers** (`library/device/mmio`'s
|
uncacheable, physical address exposed), **memory barriers** (`library/device/mmio`'s
|
||||||
`memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier`, imported as the `mmio` module), **fault isolation** (a ring-3 fault kills only the faulting
|
`memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier`, imported as the `mmio` module), **fault isolation** (a ring-3 fault kills only the faulting
|
||||||
process — `killCurrentProcess` — and the machine keeps running,
|
process — `killCurrentProcess` — and the machine keeps running,
|
||||||
[resilience](resilience.md)), and **reclaim + restart on death** (every path out of a
|
[resilience](../os-development-guide/resilience.md)), and **reclaim + restart on death** (every path out of a
|
||||||
process releases its claims and IRQ/MSI bindings — `releaseAllOwnedBy`,
|
process releases its claims and IRQ/MSI bindings — `releaseAllOwnedBy`,
|
||||||
`irq.releaseOwner` — and the device manager respawns the driver with backoff,
|
`irq.releaseOwner` — and the device manager respawns the driver with backoff,
|
||||||
[device-manager.md](device-manager.md)). What remains:
|
[device-manager.md](device-manager.md)). What remains:
|
||||||
@@ -394,7 +394,7 @@ Claiming and mapping is half of being a danos driver; the other half is the
|
|||||||
- Build on `service.run` — one replyWait loop folding protocol
|
- Build on `service.run` — one replyWait loop folding protocol
|
||||||
requests, signals, and notifications into callbacks. The harness answers the
|
requests, signals, and notifications into callbacks. The harness answers the
|
||||||
universal zero-length ping and turns `terminate` into a clean exit for you
|
universal zero-length ping and turns `terminate` into a clean exit for you
|
||||||
([process-lifecycle.md](process-lifecycle.md)).
|
([process-lifecycle.md](../os-development-guide/process-lifecycle.md)).
|
||||||
- A driver spawned with an assignment (its device id as argv[1]) sends the
|
- A driver spawned with an assignment (its device id as argv[1]) sends the
|
||||||
versioned `hello` to the device manager inside the deadline, and a **bus**
|
versioned `hello` to the device manager inside the deadline, and a **bus**
|
||||||
driver reports what it discovers with `child_added`
|
driver reports what it discovers with `child_added`
|
||||||
@@ -5,14 +5,14 @@ window server, a logger. None of them owns the hardware, and the driver should n
|
|||||||
who is listening. So between the drivers and the listeners sits the **input service**
|
who is listening. So between the drivers and the listeners sits the **input service**
|
||||||
(`system/services/input/`): drivers **publish** events to it, programs **subscribe**, and
|
(`system/services/input/`): drivers **publish** events to it, programs **subscribe**, and
|
||||||
it fans each event out to every interested subscriber. It is an ordinary ring-3 process
|
it fans each event out to every interested subscriber. It is an ordinary ring-3 process
|
||||||
reached over IPC, like the [FAT server](../system/services/fat/fat.zig) — no kernel knows
|
reached over IPC, like the [FAT server](../../system/services/fat/fat.zig) — no kernel knows
|
||||||
what a key is.
|
what a key is.
|
||||||
|
|
||||||
## One service, several device classes
|
## One service, several device classes
|
||||||
|
|
||||||
The service carries three device classes today — **keyboard**, **mouse**, and
|
The service carries three device classes today — **keyboard**, **mouse**, and
|
||||||
**joystick/gamepad** — and is built to take more
|
**joystick/gamepad** — and is built to take more
|
||||||
([protocol.zig](../library/protocol/input/input-protocol.zig)). Each class has its own typed
|
([protocol.zig](../../library/protocol/input/input-protocol.zig)). Each class has its own typed
|
||||||
event:
|
event:
|
||||||
|
|
||||||
- `KeyEvent` — `key_down`/`key_up` (physical make/break) and `key_press` (a character was
|
- `KeyEvent` — `key_down`/`key_up` (physical make/break) and `key_press` (a character was
|
||||||
@@ -45,7 +45,7 @@ consequences decide the whole design:
|
|||||||
`ipc_call`, it would block until each subscriber replied. `ipc_call` has no timeout, and
|
`ipc_call`, it would block until each subscriber replied. `ipc_call` has no timeout, and
|
||||||
a subscriber's endpoint is an *unregistered* capability the kernel's death path cannot
|
a subscriber's endpoint is an *unregistered* capability the kernel's death path cannot
|
||||||
reach (since display v2's V6, `killOwnedEndpointsLocked` in
|
reach (since display v2's V6, `killOwnedEndpointsLocked` in
|
||||||
[ipc-synchronous.zig](../system/kernel/ipc-synchronous.zig) marks a dead owner's
|
[ipc-synchronous.zig](../../system/kernel/ipc-synchronous.zig) marks a dead owner's
|
||||||
*registered* endpoints dead and wakes parked callers with `-EPEER` — but unregistered
|
*registered* endpoints dead and wakes parked callers with `-EPEER` — but unregistered
|
||||||
ones just drop with the task's handle table). One subscriber that exits mid-delivery
|
ones just drop with the task's handle table). One subscriber that exits mid-delivery
|
||||||
would wedge input for everyone. That is the opposite of the resilience the microkernel
|
would wedge input for everyone. That is the opposite of the resilience the microkernel
|
||||||
@@ -66,7 +66,7 @@ the badge (distinguishing it from a bare IRQ/child-exit notification), the sende
|
|||||||
in the low bits, and the payload in the receive buffer, with no reply owed. The queue holds
|
in the low bits, and the payload in the receive buffer, with no reply owed. The queue holds
|
||||||
16 messages per endpoint; a full queue **drops the oldest**, because a buffered message is
|
16 messages per endpoint; a full queue **drops the oldest**, because a buffered message is
|
||||||
discrete data, not a coalescing "level" like an interrupt. See
|
discrete data, not a coalescing "level" like an interrupt. See
|
||||||
[ipc-synchronous.zig](../system/kernel/ipc-synchronous.zig) (`sendLocked`, `popPost`, and
|
[ipc-synchronous.zig](../../system/kernel/ipc-synchronous.zig) (`sendLocked`, `popPost`, and
|
||||||
the `replyWait` receive loop).
|
the `replyWait` receive loop).
|
||||||
|
|
||||||
This is the async counterpart of `ipc_call`, and the input service is its first consumer.
|
This is the async counterpart of `ipc_call`, and the input service is its first consumer.
|
||||||
@@ -89,7 +89,7 @@ This is the async counterpart of `ipc_call`, and the input service is its first
|
|||||||
- A **subscriber** calls `input.subscribe(mask)` — or a typed helper: `subscribeKeyboard()`,
|
- A **subscriber** calls `input.subscribe(mask)` — or a typed helper: `subscribeKeyboard()`,
|
||||||
`subscribeMouse()`, `subscribeJoystick()` (one class, `next()` returns the decoded event),
|
`subscribeMouse()`, `subscribeJoystick()` (one class, `next()` returns the decoded event),
|
||||||
or `subscribeAll()` (every class, `next()` returns a tagged `InputEvent`)
|
or `subscribeAll()` (every class, `next()` returns a tagged `InputEvent`)
|
||||||
([library/client/input/input.zig](../library/client/input/input.zig)). It creates its own endpoint
|
([library/client/input/input.zig](../../library/client/input/input.zig)). It creates its own endpoint
|
||||||
and hands it to the service as a **capability** (M13 capability passing — the input
|
and hands it to the service as a **capability** (M13 capability passing — the input
|
||||||
service is that feature's first real user), along with its `device_mask`. Then it loops on
|
service is that feature's first real user), along with its `device_mask`. Then it loops on
|
||||||
`next()`, a `replyWait` on that endpoint returning each pushed event.
|
`next()`, a `replyWait` on that endpoint returning each pushed event.
|
||||||
@@ -98,7 +98,7 @@ This is the async counterpart of `ipc_call`, and the input service is its first
|
|||||||
`publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at
|
`publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at
|
||||||
once; the service's own fan-out is asynchronous, so publishing never blocks on a slow
|
once; the service's own fan-out is asynchronous, so publishing never blocks on a slow
|
||||||
subscriber.
|
subscriber.
|
||||||
- The **service** ([input.zig](../system/services/input/input.zig)) keeps a small subscriber
|
- The **service** ([input.zig](../../system/services/input/input.zig)) keeps a small subscriber
|
||||||
table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the
|
table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the
|
||||||
event to every subscriber whose mask includes the event's device class. On `subscribe` it
|
event to every subscriber whose mask includes the event's device class. On `subscribe` it
|
||||||
stores the passed capability and mask and, as housekeeping, prunes any slot whose owning
|
stores the passed capability and mask and, as housekeeping, prunes any slot whose owning
|
||||||
@@ -113,18 +113,18 @@ the service delivers to its endpoint, which only the same thread could receive).
|
|||||||
|
|
||||||
- **The keyboard is real.** The `ps2-bus` driver owns PNP0303, which carries *both* the
|
- **The keyboard is real.** The `ps2-bus` driver owns PNP0303, which carries *both* the
|
||||||
0x60/0x64 ports and IRQ1, so reading the hardware lives in the bus, not in
|
0x60/0x64 ports and IRQ1, so reading the hardware lives in the bus, not in
|
||||||
[keyboard.zig](../system/drivers/ps2-bus/keyboard.zig): the bus binds IRQ1 and, on each
|
[keyboard.zig](../../system/drivers/ps2-bus/keyboard.zig): the bus binds IRQ1 and, on each
|
||||||
interrupt, drains port 0x60, routing every byte by the status register's
|
interrupt, drains port 0x60, routing every byte by the status register's
|
||||||
auxiliary-output bit to whichever child driver **attached** for that device (an
|
auxiliary-output bit to whichever child driver **attached** for that device (an
|
||||||
`AttachRequest` to the well-known `ps2_bus` service, carrying the child's endpoint as a
|
`AttachRequest` to the well-known `ps2_bus` service, carrying the child's endpoint as a
|
||||||
capability; the bytes then arrive as asynchronous `ForwardedByte` messages, so the IRQ
|
capability; the bytes then arrive as asynchronous `ForwardedByte` messages, so the IRQ
|
||||||
path never blocks on a child). The keyboard driver decodes the stream — scancode **set 2**,
|
path never blocks on a child). The keyboard driver decodes the stream — scancode **set 2**,
|
||||||
what the keyboard sends with the 8042's legacy translation off, decoded by
|
what the keyboard sends with the 8042's legacy translation off, decoded by
|
||||||
[scancode.zig](../system/drivers/ps2-bus/scancode.zig) into USB HID usage keycodes with
|
[scancode.zig](../../system/drivers/ps2-bus/scancode.zig) into USB HID usage keycodes with
|
||||||
make/break, typematic-repeat, and modifier tracking (host-tested under `zig build test`) —
|
make/break, typematic-repeat, and modifier tracking (host-tested under `zig build test`) —
|
||||||
and publishes real `key_down`/`key_press`/`key_up` events.
|
and publishes real `key_down`/`key_press`/`key_up` events.
|
||||||
- **Keycode → character** is wired in: the keyboard driver fills a `key_press` event's
|
- **Keycode → character** is wired in: the keyboard driver fills a `key_press` event's
|
||||||
`character` through [`library/xkeyboard-config`](../library/xkeyboard-config/README.md)
|
`character` through [`library/xkeyboard-config`](../../library/xkeyboard-config/README.md)
|
||||||
(`xkb.map(layout, keycode, mods)` → keysym + Unicode character), synthesizing the ASCII
|
(`xkb.map(layout, keycode, mods)` → keysym + Unicode character), synthesizing the ASCII
|
||||||
control characters for Enter/Tab/Backspace/Escape, whose keysyms map to no Unicode. The
|
control characters for Enter/Tab/Backspace/Escape, whose keysyms map to no Unicode. The
|
||||||
layout defaults to `us`; the bus can pass another as the driver's argv[2] — the seam for
|
layout defaults to `us`; the bus can pass another as the driver's argv[2] — the seam for
|
||||||
@@ -132,9 +132,9 @@ the service delivers to its endpoint, which only the same thread could receive).
|
|||||||
- **The mouse is real too.** IRQ12 is enumerated on the auxiliary device's own ACPI node
|
- **The mouse is real too.** IRQ12 is enumerated on the auxiliary device's own ACPI node
|
||||||
(PNP0F13), so the bus claims that node alongside the controller and routes both IRQs to
|
(PNP0F13), so the bus claims that node alongside the controller and routes both IRQs to
|
||||||
its one endpoint, acking whichever line the notification's badge names.
|
its one endpoint, acking whichever line the notification's badge names.
|
||||||
[mouse.zig](../system/drivers/ps2-bus/mouse.zig) attaches the way the keyboard does and
|
[mouse.zig](../../system/drivers/ps2-bus/mouse.zig) attaches the way the keyboard does and
|
||||||
assembles the forwarded bytes with
|
assembles the forwarded bytes with
|
||||||
[mouse-packet.zig](../system/drivers/ps2-bus/mouse-packet.zig) (three-byte stream-mode
|
[mouse-packet.zig](../../system/drivers/ps2-bus/mouse-packet.zig) (three-byte stream-mode
|
||||||
packets: sync/overflow handling, nine-bit movement, screen-convention `dy` — host-tested
|
packets: sync/overflow handling, nine-bit movement, screen-convention `dy` — host-tested
|
||||||
under `zig build test`) into `button_down`/`button_up` transitions and `motion` events.
|
under `zig build test`) into `button_down`/`button_up` transitions and `motion` events.
|
||||||
**Follow-up:** the IntelliMouse magic-knock for a scroll wheel (four-byte packets) and
|
**Follow-up:** the IntelliMouse magic-knock for a scroll wheel (four-byte packets) and
|
||||||
@@ -150,7 +150,7 @@ the service delivers to its endpoint, which only the same thread could receive).
|
|||||||
## Verifying it
|
## Verifying it
|
||||||
|
|
||||||
The `input` case (`python3 test/qemu_test.py input`, in
|
The `input` case (`python3 test/qemu_test.py input`, in
|
||||||
[tests.zig](../system/kernel/tests.zig) `inputTest`) boots the real kernel and spawns the
|
[tests.zig](../../system/kernel/tests.zig) `inputTest`) boots the real kernel and spawns the
|
||||||
service, the synthetic source (which cycles keyboard, mouse, and joystick events), and a
|
service, the synthetic source (which cycles keyboard, mouse, and joystick events), and a
|
||||||
subscriber that took all three classes. It passes only when the subscriber heartbeats
|
subscriber that took all three classes. It passes only when the subscriber heartbeats
|
||||||
`input-test: ok` — proof that an event travelled source → service → subscriber over IPC,
|
`input-test: ok` — proof that an event travelled source → service → subscriber over IPC,
|
||||||
@@ -160,5 +160,5 @@ serial line names the class received, so the log shows all three arriving on one
|
|||||||
## See also
|
## See also
|
||||||
|
|
||||||
- [ipc.md](ipc.md) — the synchronous rendezvous and the notification path `ipc_send` extends.
|
- [ipc.md](ipc.md) — the synchronous rendezvous and the notification path `ipc_send` extends.
|
||||||
- [syscall.md](syscall.md) — the system-call surface, including `ipc_send`.
|
- [syscall.md](../os-development-guide/syscall.md) — the system-call surface, including `ipc_send`.
|
||||||
- [driver-model.md](driver-model.md) — class drivers, capability passing (M13), the trust model.
|
- [driver-model.md](driver-model.md) — class drivers, capability passing (M13), the trust model.
|
||||||
@@ -1,7 +1,7 @@
|
|||||||
# IPC: message-passing channels
|
# IPC: message-passing channels
|
||||||
|
|
||||||
Inter-process communication is the **backbone of a microkernel**. Once drivers and
|
Inter-process communication is the **backbone of a microkernel**. Once drivers and
|
||||||
services run isolated in their own address spaces ([vision](vision.md)), they can't
|
services run isolated in their own address spaces ([vision](../vision.md)), they can't
|
||||||
just call each other — a request becomes a **message**. In a microkernel, whatever
|
just call each other — a request becomes a **message**. In a microkernel, whatever
|
||||||
was a function call across a monolithic kernel is IPC, so it's a first-class
|
was a function call across a monolithic kernel is IPC, so it's a first-class
|
||||||
concern, not an afterthought.
|
concern, not an afterthought.
|
||||||
@@ -18,7 +18,7 @@ There are two layers, built a milestone apart:
|
|||||||
|
|
||||||
The first form is a **bounded blocking channel** (`system/kernel/ipc.zig`): a fixed-size
|
The first form is a **bounded blocking channel** (`system/kernel/ipc.zig`): a fixed-size
|
||||||
ring buffer of messages with a producer/consumer rendezvous, built on the
|
ring buffer of messages with a producer/consumer rendezvous, built on the
|
||||||
scheduler's [wait queues](scheduling.md).
|
scheduler's [wait queues](../os-development-guide/scheduling.md).
|
||||||
|
|
||||||
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
|
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
|
||||||
ring buffer, a count, and two wait queues:
|
ring buffer, a count, and two wait queues:
|
||||||
@@ -37,7 +37,7 @@ Two details make it correct:
|
|||||||
rather than assuming the slot is still available — another waiter may have taken
|
rather than assuming the slot is still available — another waiter may have taken
|
||||||
it first. This is the standard guard against spurious or racing wakeups.
|
it first. This is the standard guard against spurious or racing wakeups.
|
||||||
- **One critical section.** `send`/`receive` run under the [big kernel
|
- **One critical section.** `send`/`receive` run under the [big kernel
|
||||||
lock](smp.md) (`sync.enter` / `sync.leave`), which disables interrupts on this
|
lock](../os-development-guide/smp.md) (`sync.enter` / `sync.leave`), which disables interrupts on this
|
||||||
core *and* takes the kernel's one spinlock — since SMP, the interrupt flag alone
|
core *and* takes the kernel's one spinlock — since SMP, the interrupt flag alone
|
||||||
is not atomicity, because `cli` on one core does nothing to another. So checking
|
is not atomicity, because `cli` on one core does nothing to another. So checking
|
||||||
the condition and committing the block/enqueue happen atomically both with respect
|
the condition and committing the block/enqueue happen atomically both with respect
|
||||||
@@ -47,7 +47,7 @@ Two details make it correct:
|
|||||||
|
|
||||||
## Verifying it
|
## Verifying it
|
||||||
|
|
||||||
The `ipc` test (see [testing.md](testing.md)) runs a producer and a consumer passing
|
The `ipc` test (see [testing.md](../testing.md)) runs a producer and a consumer passing
|
||||||
**100 messages through a 4-slot channel**. The small buffer means the channel goes
|
**100 messages through a 4-slot channel**. The small buffer means the channel goes
|
||||||
full and empty over and over, so both the blocking-send and blocking-receive paths are
|
full and empty over and over, so both the blocking-send and blocking-receive paths are
|
||||||
exercised heavily. The messages arrive intact and in order (their sum is the
|
exercised heavily. The messages arrive intact and in order (their sum is the
|
||||||
@@ -114,7 +114,7 @@ This is what makes a user-space driver possible at all, and it's the subject of
|
|||||||
|
|
||||||
## Lifecycle conventions over IPC (M17)
|
## Lifecycle conventions over IPC (M17)
|
||||||
|
|
||||||
Three conventions from [process-lifecycle.md](process-lifecycle.md) ride the
|
Three conventions from [process-lifecycle.md](../os-development-guide/process-lifecycle.md) ride the
|
||||||
notification mechanism:
|
notification mechanism:
|
||||||
|
|
||||||
- **Signals** arrive as notifications on the endpoint a process nominated with
|
- **Signals** arrive as notifications on the endpoint a process nominated with
|
||||||
+3
-3
@@ -49,7 +49,7 @@ addressed by device id. `/dev` is the much smaller set of devices that have a dr
|
|||||||
willing to serve them, addressed by name.
|
willing to serve them, addressed by name.
|
||||||
|
|
||||||
A device node is not a file the VFS can read. The bytes live in a driver process
|
A device node is not a file the VFS can read. The bytes live in a driver process
|
||||||
([drivers.md](drivers.md)), so opening a `/dev` name has to resolve to that driver's
|
([drivers.md](../device-driver-development-guide/drivers.md)), so opening a `/dev` name has to resolve to that driver's
|
||||||
IPC endpoint, and subsequent reads and writes are calls against it. Resolve-to-endpoint
|
IPC endpoint, and subsequent reads and writes are calls against it. Resolve-to-endpoint
|
||||||
is exactly what the kernel's `fs_resolve` already does for any mounted backend, and
|
is exactly what the kernel's `fs_resolve` already does for any mounted backend, and
|
||||||
`FileStatus.kind` is the field that marks a device node; **what is not implemented today
|
`FileStatus.kind` is the field that marks a device node; **what is not implemented today
|
||||||
@@ -92,7 +92,7 @@ descriptor ring and left to read and write memory on its own. That ring is exact
|
|||||||
**`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its
|
**`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its
|
||||||
physical address disclosed — and **`/lib/device/mmio`**'s barriers order the descriptor writes
|
physical address disclosed — and **`/lib/device/mmio`**'s barriers order the descriptor writes
|
||||||
against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver
|
against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver
|
||||||
can be written today (the M14/M15 work in [driver-model.md](driver-model.md); the earlier
|
can be written today (the M14/M15 work in [driver-model.md](../device-driver-development-guide/driver-model.md); the earlier
|
||||||
"cannot host a block driver at all" is no longer true).
|
"cannot host a block driver at all" is no longer true).
|
||||||
|
|
||||||
What is *not* yet true is that it is safe. A device programmed with an arbitrary physical
|
What is *not* yet true is that it is safe. A device programmed with an arbitrary physical
|
||||||
@@ -125,4 +125,4 @@ caller can tell a character device from a regular file.
|
|||||||
options on this kernel are `RDRAND`/`RDSEED` where CPUID advertises them, and the HPET
|
options on this kernel are `RDRAND`/`RDSEED` where CPUID advertises them, and the HPET
|
||||||
counter's low bits as a poor fallback. Neither is a seeded CSPRNG, and a `/dev/random`
|
counter's low bits as a poor fallback. Neither is a seeded CSPRNG, and a `/dev/random`
|
||||||
that is merely unpredictable-looking is worse than none — nothing should be keyed from
|
that is merely unpredictable-looking is worse than none — nothing should be keyed from
|
||||||
it until it is a real one.
|
it until it is a real one.
|
||||||
@@ -9,7 +9,7 @@
|
|||||||
> backend, unchanged. The Zig source of truth is `library/protocol/vfs/vfs-protocol.zig`
|
> backend, unchanged. The Zig source of truth is `library/protocol/vfs/vfs-protocol.zig`
|
||||||
> (the `vfs-protocol` module), whose unit test pins a sample of the sizes
|
> (the `vfs-protocol` module), whose unit test pins a sample of the sizes
|
||||||
> and values below. This page is the **language-neutral wire specification**
|
> and values below. This page is the **language-neutral wire specification**
|
||||||
> of that contract — what a Rust or C client implements ([vdso.md](vdso.md)
|
> of that contract — what a Rust or C client implements ([vdso.md](../os-development-guide/vdso.md)
|
||||||
> explains why the IPC protocols, not the syscall numbers, are danos's
|
> explains why the IPC protocols, not the syscall numbers, are danos's
|
||||||
> public ABI).
|
> public ABI).
|
||||||
|
|
||||||
@@ -75,9 +75,9 @@ There are really two independent questions, and it's worth not conflating them:
|
|||||||
- **`system/kernel/architecture/x86_64/paging.zig`** — the kernel's page tables and address-space
|
- **`system/kernel/architecture/x86_64/paging.zig`** — the kernel's page tables and address-space
|
||||||
management (see [paging.md](paging.md)).
|
management (see [paging.md](paging.md)).
|
||||||
- **`system/kernel/architecture/x86_64/apic.zig`** / **`ioapic.zig`** — the Local APIC, its timer,
|
- **`system/kernel/architecture/x86_64/apic.zig`** / **`ioapic.zig`** — the Local APIC, its timer,
|
||||||
and the I/O APIC for device interrupts (see [device-interrupts.md](device-interrupts.md)).
|
and the I/O APIC for device interrupts (see [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)).
|
||||||
- **`system/kernel/architecture/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's
|
- **`system/kernel/architecture/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's
|
||||||
machine-readable log channel, see [testing.md](testing.md)) and the shared port-I/O + MSR primitives.
|
machine-readable log channel, see [testing.md](../testing.md)) and the shared port-I/O + MSR primitives.
|
||||||
- **`system/kernel/architecture/x86_64/smp.zig`** / **`per-cpu.zig`** — application-processor bring-up
|
- **`system/kernel/architecture/x86_64/smp.zig`** / **`per-cpu.zig`** — application-processor bring-up
|
||||||
and per-CPU state (GS base, system-call entry point, see [scheduling.md](scheduling.md)).
|
and per-CPU state (GS base, system-call entry point, see [scheduling.md](scheduling.md)).
|
||||||
- **`system/kernel/architecture/x86_64/isr.s`** — the exception stubs, the `lgdt`/`lidt`/`ltr` load
|
- **`system/kernel/architecture/x86_64/isr.s`** — the exception stubs, the `lgdt`/`lidt`/`ltr` load
|
||||||
@@ -87,7 +87,7 @@ The Pi is not a "standard" ARM platform — expect Broadcom-specific peripherals
|
|||||||
Pi 5. Everything below is an offset from it.
|
Pi 5. Everything below is an offset from it.
|
||||||
- **UART**: a **PL011** (at base + `0x20_1000`) plus a mini-UART; on some boards the
|
- **UART**: a **PL011** (at base + `0x20_1000`) plus a mini-UART; on some boards the
|
||||||
PL011 is wired to Bluetooth, so which one is the console varies. This is the
|
PL011 is wired to Bluetooth, so which one is the console varies. This is the
|
||||||
`aarch64`/`arm` equivalent of our x86 [COM1 serial](testing.md).
|
`aarch64`/`arm` equivalent of our x86 [COM1 serial](../testing.md).
|
||||||
- **Interrupt controller**: *not* a standard ARM GIC on the older parts — the Zero W
|
- **Interrupt controller**: *not* a standard ARM GIC on the older parts — the Zero W
|
||||||
and Pi 3 use Broadcom's own ARMCTRL controller (Pi 3 adds a per-core "local"
|
and Pi 3 use Broadcom's own ARMCTRL controller (Pi 3 adds a per-core "local"
|
||||||
controller for timers/mailboxes). The **Pi 4 and 5 do have a GIC-400**. So the
|
controller for timers/mailboxes). The **Pi 4 and 5 do have a GIC-400**. So the
|
||||||
@@ -122,4 +122,4 @@ Two routes, mirroring how we test x86-64 with OVMF:
|
|||||||
- [architecture.md](architecture.md) — the arch-module boundary these targets plug into, and the
|
- [architecture.md](architecture.md) — the arch-module boundary these targets plug into, and the
|
||||||
CPU-arch vs boot-protocol "two axes".
|
CPU-arch vs boot-protocol "two axes".
|
||||||
- [efi.md](efi.md) — the UEFI loader that carries over to aarch64-UEFI.
|
- [efi.md](efi.md) — the UEFI loader that carries over to aarch64-UEFI.
|
||||||
- [vision.md](vision.md) — why isolated, portable-across-architectures is the goal.
|
- [vision.md](../vision.md) — why isolated, portable-across-architectures is the goal.
|
||||||
@@ -36,7 +36,7 @@ Two things drive the need, and they set the timing:
|
|||||||
**aarch64 it's required to boot at all**. The [aarch64 port](arm.md) is what
|
**aarch64 it's required to boot at all**. The [aarch64 port](arm.md) is what
|
||||||
forces the issue.
|
forces the issue.
|
||||||
|
|
||||||
2. **Isolated user-space drivers need it.** In the [microkernel vision](vision.md),
|
2. **Isolated user-space drivers need it.** In the [microkernel vision](../vision.md),
|
||||||
drivers live in user space — but something has to enumerate the hardware and hand
|
drivers live in user space — but something has to enumerate the hardware and hand
|
||||||
each driver its MMIO regions and IRQs. That enumeration *is* device discovery. So
|
each driver its MMIO regions and IRQs. That enumeration *is* device discovery. So
|
||||||
discovery is a prerequisite for real drivers, **not** for user mode itself.
|
discovery is a prerequisite for real drivers, **not** for user mode itself.
|
||||||
@@ -132,7 +132,7 @@ when*:
|
|||||||
- **User-space enumeration: a device-manager server.** Everything else — PCI devices,
|
- **User-space enumeration: a device-manager server.** Everything else — PCI devices,
|
||||||
peripherals — is parsed (or queried from the kernel's parse) by a privileged
|
peripherals — is parsed (or queried from the kernel's parse) by a privileged
|
||||||
user-space server that hands each driver process its MMIO regions and IRQ rights
|
user-space server that hands each driver process its MMIO regions and IRQ rights
|
||||||
over [IPC](ipc.md). Combined with **interrupts-as-messages** (an IRQ delivered to a
|
over [IPC](../device-driver-development-guide/ipc.md). Combined with **interrupts-as-messages** (an IRQ delivered to a
|
||||||
driver as a message on a channel — a natural extension of the wait queues and
|
driver as a message on a channel — a natural extension of the wait queues and
|
||||||
channels already built), that's what makes drivers genuinely isolated.
|
channels already built), that's what makes drivers genuinely isolated.
|
||||||
|
|
||||||
@@ -144,7 +144,7 @@ slice is unavoidably in-kernel.
|
|||||||
On ARMv8 the generic timer exposes its frequency directly via the `CNTFRQ` register —
|
On ARMv8 the generic timer exposes its frequency directly via the `CNTFRQ` register —
|
||||||
no calibration needed. That's cleaner than the x86 side, where we measure the LAPIC
|
no calibration needed. That's cleaner than the x86 side, where we measure the LAPIC
|
||||||
and TSC against the PIT because nothing tells us their frequency (see
|
and TSC against the PIT because nothing tells us their frequency (see
|
||||||
[device-interrupts.md](device-interrupts.md)). Discovery on ARM hands you more for
|
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). Discovery on ARM hands you more for
|
||||||
free; discovery on x86 is partly about *finding* what ARM just tells you.
|
free; discovery on x86 is partly about *finding* what ARM just tells you.
|
||||||
|
|
||||||
## Suggested ordering
|
## Suggested ordering
|
||||||
@@ -166,18 +166,18 @@ free; discovery on x86 is partly about *finding* what ARM just tells you.
|
|||||||
- [arm.md](arm.md) — the aarch64 target that forces genuine discovery (DTB, GIC).
|
- [arm.md](arm.md) — the aarch64 target that forces genuine discovery (DTB, GIC).
|
||||||
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam,
|
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam,
|
||||||
and the note about grabbing the RSDP before `ExitBootServices`.
|
and the note about grabbing the RSDP before `ExitBootServices`.
|
||||||
- [device-interrupts.md](device-interrupts.md) — the LAPIC/timer bring-up that
|
- [device-interrupts.md](../device-driver-development-guide/device-interrupts.md) — the LAPIC/timer bring-up that
|
||||||
discovery will eventually feed (IOAPIC, real IRQ routing).
|
discovery will eventually feed (IOAPIC, real IRQ routing).
|
||||||
- [ipc.md](ipc.md) — the channels that interrupts-as-messages and the device manager
|
- [ipc.md](../device-driver-development-guide/ipc.md) — the channels that interrupts-as-messages and the device manager
|
||||||
will ride on.
|
will ride on.
|
||||||
- [vision.md](vision.md) — why drivers belong in isolated user space at all.
|
- [vision.md](../vision.md) — why drivers belong in isolated user space at all.
|
||||||
|
|
||||||
## Update (M19.3, 2026-07-13): PCI enumeration left the kernel
|
## Update (M19.3, 2026-07-13): PCI enumeration left the kernel
|
||||||
|
|
||||||
The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO
|
The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO
|
||||||
apertures derived from the memory map's holes, bus range, and the 16-bit I/O
|
apertures derived from the memory map's holes, bus range, and the 16-bit I/O
|
||||||
window). The per-function walk moved to the ring-3 `pci-bus` driver
|
window). The per-function walk moved to the ring-3 `pci-bus` driver
|
||||||
([device-manager.md](device-manager.md)): it claims the bridge, repeats the
|
([device-manager.md](../device-driver-development-guide/device-manager.md)): it claims the bridge, repeats the
|
||||||
ECAM scan through its mmio grant, and `device_register`s what it finds, which
|
ECAM scan through its mmio grant, and `device_register`s what it finds, which
|
||||||
the device manager mirrors and matches. The ACPI namespace walk follows in M20;
|
the device manager mirrors and matches. The ACPI namespace walk follows in M20;
|
||||||
the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side.
|
the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side.
|
||||||
@@ -190,7 +190,7 @@ for the host bridge, FADT); at this point it also still built the AML namespace
|
|||||||
but only to read the `\\_S5` sleep type for poweroff. (That remnant is gone too:
|
but only to read the `\\_S5` sleep type for poweroff. (That remnant is gone too:
|
||||||
the kernel now runs no AML at all — soft-off belongs to the acpi service, and the
|
the kernel now runs no AML at all — soft-off belongs to the acpi service, and the
|
||||||
kernel keeps only the AML-free reboot path.) Device discovery is the ring-3 **acpi
|
kernel keeps only the AML-free reboot path.) Device discovery is the ring-3 **acpi
|
||||||
service** ([device-manager.md](device-manager.md)): it claims the `acpi-tables`
|
service** ([device-manager.md](../device-driver-development-guide/device-manager.md)): it claims the `acpi-tables`
|
||||||
node the kernel publishes (the AML blobs, a broad io_port grant, the SCI),
|
node the kernel publishes (the AML blobs, a broad io_port grant, the SCI),
|
||||||
re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`,
|
re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`,
|
||||||
and registers + reports each `_HID` device — the device manager matches drivers
|
and registers + reports each `_HID` device — the device manager matches drivers
|
||||||
@@ -206,7 +206,7 @@ ring 0.)
|
|||||||
Moving PCI and ACPI enumeration out of ring 0 was not just a relocation — it
|
Moving PCI and ACPI enumeration out of ring 0 was not just a relocation — it
|
||||||
made discovery **firmware-neutral by construction**, which is the whole reason
|
made discovery **firmware-neutral by construction**, which is the whole reason
|
||||||
to do it before the second architecture rather than after. Everything at and
|
to do it before the second architecture rather than after. Everything at and
|
||||||
above the [device-manager](device-manager.md) protocol — descriptors,
|
above the [device-manager](../device-driver-development-guide/device-manager.md) protocol — descriptors,
|
||||||
containment, reports, matching, supervision — is generic and may never become
|
containment, reports, matching, supervision — is generic and may never become
|
||||||
x86-specific. Discovery is the single firmware-specific piece, and it is
|
x86-specific. Discovery is the single firmware-specific piece, and it is
|
||||||
isolated as **one swappable process per firmware**:
|
isolated as **one swappable process per firmware**:
|
||||||
@@ -24,7 +24,7 @@ EFI/BOOT/BOOTX64.efi <- the "removable media" default for x86-64
|
|||||||
```
|
```
|
||||||
|
|
||||||
The boot volume is **FHS-shaped** (see the repository-layout note in
|
The boot volume is **FHS-shaped** (see the repository-layout note in
|
||||||
[README.md](README.md)): `build.zig` installs `boot/efi.zig` (built for the `uefi`
|
[README.md](../README.md)): `build.zig` installs `boot/efi.zig` (built for the `uefi`
|
||||||
target) at `EFI/BOOT/BOOTX64.efi` — the one path UEFI firmware fixes — and lays
|
target) at `EFI/BOOT/BOOTX64.efi` — the one path UEFI firmware fixes — and lays
|
||||||
the rest out by FHS path: the kernel at `system/kernel`, init at
|
the rest out by FHS path: the kernel at `system/kernel`, init at
|
||||||
`system/services/init`, the pre-packed boot capsule at `boot/system.img`
|
`system/services/init`, the pre-packed boot capsule at `boot/system.img`
|
||||||
@@ -51,7 +51,7 @@ rest — works directly on the kernel heap, no bespoke containers required.
|
|||||||
|
|
||||||
## Verifying it
|
## Verifying it
|
||||||
|
|
||||||
The `heap` test (see [testing.md](testing.md)) exercises the allocator end to end:
|
The `heap` test (see [testing.md](../testing.md)) exercises the allocator end to end:
|
||||||
|
|
||||||
```
|
```
|
||||||
[PASS] alloc 4096 bytes
|
[PASS] alloc 4096 bytes
|
||||||
@@ -133,10 +133,10 @@ TSS/IST is wired up: the handler survived a completely broken stack.
|
|||||||
|
|
||||||
Both items originally deferred here have landed:
|
Both items originally deferred here have landed:
|
||||||
|
|
||||||
- **The IO-APIC**: [ioapic.zig](../system/kernel/architecture/x86_64/ioapic.zig)
|
- **The IO-APIC**: [ioapic.zig](../../system/kernel/architecture/x86_64/ioapic.zig)
|
||||||
routes external device lines onto vectors — discovered via ACPI's MADT, every
|
routes external device lines onto vectors — discovered via ACPI's MADT, every
|
||||||
input masked at init, lines unmasked one at a time as user-space drivers bind
|
input masked at init, lines unmasked one at a time as user-space drivers bind
|
||||||
them (see [device-interrupts.md](device-interrupts.md)). The keyboard followed
|
them (see [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). The keyboard followed
|
||||||
exactly as predicted: the PS/2 bus driver (`system/drivers/ps2-bus/`) claims
|
exactly as predicted: the PS/2 bus driver (`system/drivers/ps2-bus/`) claims
|
||||||
the 8042 controller and binds its IRQ 1 (and the aux mouse's IRQ 12) through
|
the 8042 controller and binds its IRQ 1 (and the aux mouse's IRQ 12) through
|
||||||
this routing. USB HID keyboards arrive over xHCI instead, which interrupts via
|
this routing. USB HID keyboards arrive over xHCI instead, which interrupts via
|
||||||
@@ -112,7 +112,7 @@ kernel heap will build on to map pages on demand.
|
|||||||
|
|
||||||
## Verifying it
|
## Verifying it
|
||||||
|
|
||||||
Four tests (see [testing.md](testing.md)) pin down the guarantees:
|
Four tests (see [testing.md](../testing.md)) pin down the guarantees:
|
||||||
|
|
||||||
- **`vmm`** — map a fresh frame at an unused virtual address, write and read it
|
- **`vmm`** — map a fresh frame at an unused virtual address, write and read it
|
||||||
back. Proves `map` works end to end.
|
back. Proves `map` works end to end.
|
||||||
@@ -138,8 +138,8 @@ Four tests (see [testing.md](testing.md)) pin down the guarantees:
|
|||||||
processes own the low half.
|
processes own the low half.
|
||||||
- **Per-address-space tables** — done: each user process gets its own root with
|
- **Per-address-space tables** — done: each user process gets its own root with
|
||||||
the kernel half shared, and refcounted shared-memory mappings exist
|
the kernel half shared, and refcounted shared-memory mappings exist
|
||||||
([ipc.md](ipc.md)). Copy-on-write remains unbuilt — nothing has needed it yet.
|
([ipc.md](../device-driver-development-guide/ipc.md)). Copy-on-write remains unbuilt — nothing has needed it yet.
|
||||||
- **Uncacheable MMIO** — half done: user-space device and DMA mappings are
|
- **Uncacheable MMIO** — half done: user-space device and DMA mappings are
|
||||||
strong-uncacheable and the framebuffer is write-combining via the PAT, but the
|
strong-uncacheable and the framebuffer is write-combining via the PAT, but the
|
||||||
kernel's own `mapMmio` path is still writeback — the LAPIC included (see
|
kernel's own `mapMmio` path is still writeback — the LAPIC included (see
|
||||||
[device-interrupts.md](device-interrupts.md)).
|
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md)).
|
||||||
@@ -7,7 +7,7 @@ them owns the hardware that reported the event, and the reporter should not know
|
|||||||
who is listening. So system power is a **service**: an event source **publishes**
|
who is listening. So system power is a **service**: an event source **publishes**
|
||||||
button/lid/battery/AC events, interested processes **subscribe**, and one
|
button/lid/battery/AC events, interested processes **subscribe**, and one
|
||||||
privileged caller — init — can ask it to power the machine off. It is the same
|
privileged caller — init — can ask it to power the machine off. It is the same
|
||||||
publish/subscribe shape as the [input service](input.md), applied to power.
|
publish/subscribe shape as the [input service](../device-driver-development-guide/input.md), applied to power.
|
||||||
|
|
||||||
## Why a service, and why it is named for the domain, not the firmware
|
## Why a service, and why it is named for the domain, not the firmware
|
||||||
|
|
||||||
@@ -27,7 +27,7 @@ unchanged.
|
|||||||
|
|
||||||
## The protocol
|
## The protocol
|
||||||
|
|
||||||
The `power-protocol` module ([library/protocol/power/power-protocol.zig](../library/protocol/power/power-protocol.zig))
|
The `power-protocol` module ([library/protocol/power/power-protocol.zig](../../library/protocol/power/power-protocol.zig))
|
||||||
follows the vfs-protocol pattern — extern-struct messages, a version, reserved
|
follows the vfs-protocol pattern — extern-struct messages, a version, reserved
|
||||||
fields. Three operations:
|
fields. Three operations:
|
||||||
|
|
||||||
@@ -126,5 +126,5 @@ until laptop sleep), and thermal zones.
|
|||||||
firmware neutrality that makes a PSCI backend drop-in on ARM.
|
firmware neutrality that makes a PSCI backend drop-in on ARM.
|
||||||
- [process-lifecycle.md](process-lifecycle.md) — the stop sequence
|
- [process-lifecycle.md](process-lifecycle.md) — the stop sequence
|
||||||
(`terminate → deadline → kill`) and signals init composes into shutdown.
|
(`terminate → deadline → kill`) and signals init composes into shutdown.
|
||||||
- [device-manager.md](device-manager.md) — the supervision model init mirrors for
|
- [device-manager.md](../device-driver-development-guide/device-manager.md) — the supervision model init mirrors for
|
||||||
its own children.
|
its own children.
|
||||||
@@ -9,7 +9,7 @@ the layer above them — the standard vocabulary a danos process speaks about it
|
|||||||
life, and the stable `process` interface that carries it. Nothing here is
|
life, and the stable `process` interface that carries it. Nothing here is
|
||||||
device- or driver-specific: a driver, the VFS, and a user application all stop,
|
device- or driver-specific: a driver, the VFS, and a user application all stop,
|
||||||
reload, and die the same way. The device manager is simply this design's first
|
reload, and die the same way. The device manager is simply this design's first
|
||||||
serious customer ([device-manager.md](device-manager.md)).
|
serious customer ([device-manager.md](../device-driver-development-guide/device-manager.md)).
|
||||||
|
|
||||||
**"POSIX" in this document means the concepts, never the letter of the standard.**
|
**"POSIX" in this document means the concepts, never the letter of the standard.**
|
||||||
danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE
|
danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE
|
||||||
@@ -19,7 +19,7 @@ rule is danos's own and it is strict: plain words that communicate intent
|
|||||||
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
|
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
|
||||||
a concept that already has one. Literal POSIX arrives later and lives elsewhere: the
|
a concept that already has one. Literal POSIX arrives later and lives elsewhere: the
|
||||||
`std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C
|
`std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C
|
||||||
layer** on the same native surface (see [zig-self-hosting.md](zig-self-hosting.md)) —
|
layer** on the same native surface (see [zig-self-hosting.md](../zig-self-hosting.md)) —
|
||||||
musl's syscall surface retargeted at danos system calls and IPC protocols (files onto
|
musl's syscall surface retargeted at danos system calls and IPC protocols (files onto
|
||||||
the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever
|
the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever
|
||||||
networking becomes). Ported programs see POSIX; the system underneath never does.
|
networking becomes). Ported programs see POSIX; the system underneath never does.
|
||||||
@@ -172,7 +172,7 @@ zombie state or privileged snooping:
|
|||||||
the server's reply with `-EPEER`; a server that dies fails its waiting clients
|
the server's reply with `-EPEER`; a server that dies fails its waiting clients
|
||||||
the same way. This covers the *synchronous* case only.
|
the same way. This covers the *synchronous* case only.
|
||||||
3. **The subscribers** — the new piece, and it is the input service's
|
3. **The subscribers** — the new piece, and it is the input service's
|
||||||
publish/subscribe shape ([input.md](input.md)) applied to exits. A stateful
|
publish/subscribe shape ([input.md](../device-driver-development-guide/input.md)) applied to exits. A stateful
|
||||||
service accumulates per-client state across many requests: a filesystem server
|
service accumulates per-client state across many requests: a filesystem server
|
||||||
(FAT today) holds a dead client's open file handles, the input service holds
|
(FAT today) holds a dead client's open file handles, the input service holds
|
||||||
its subscriptions, a future network stack holds its sockets. None of these
|
its subscriptions, a future network stack holds its sockets. None of these
|
||||||
@@ -320,7 +320,7 @@ get POSIX; danos-native programs never pay for it.
|
|||||||
`process` grows the interface above; the service harness handles
|
`process` grows the interface above; the service harness handles
|
||||||
`terminate` and answers the common `ping`; `stop()` for supervisors.
|
`terminate` and answers the common `ping`; `stop()` for supervisors.
|
||||||
|
|
||||||
[device-manager.md](device-manager.md) builds directly on all four.
|
[device-manager.md](../device-driver-development-guide/device-manager.md) builds directly on all four.
|
||||||
|
|
||||||
## Settled questions (2026-07-12)
|
## Settled questions (2026-07-12)
|
||||||
|
|
||||||
@@ -44,7 +44,7 @@ tables of contents, both pointing at the same embedded FAT image:
|
|||||||
same `BOOTX64.efi` off it.
|
same `BOOTX64.efi` off it.
|
||||||
|
|
||||||
Neither path involves the legacy BIOS boot-sector machinery: danos is
|
Neither path involves the legacy BIOS boot-sector machinery: danos is
|
||||||
UEFI-only ([system-requirements.md](system-requirements.md)), so the MBR holds
|
UEFI-only ([system-requirements.md](../system-requirements.md)), so the MBR holds
|
||||||
no boot code, just the partition entry, and the El Torito entry is EFI-class,
|
no boot code, just the partition entry, and the El Torito entry is EFI-class,
|
||||||
not floppy emulation.
|
not floppy emulation.
|
||||||
|
|
||||||
@@ -57,7 +57,7 @@ allocated from the front) either way. The USB path has no such cap.
|
|||||||
## The builder
|
## The builder
|
||||||
|
|
||||||
`tools/make-iso-image.py` follows the house rule of
|
`tools/make-iso-image.py` follows the house rule of
|
||||||
[make-fat-image.py](../tools/make-fat-image.py): pure Python 3 standard
|
[make-fat-image.py](../../tools/make-fat-image.py): pure Python 3 standard
|
||||||
library, no external tools (no xorriso, mkisofs, or isohybrid), with a
|
library, no external tools (no xorriso, mkisofs, or isohybrid), with a
|
||||||
`--verify` mode the `check-iso-image` step runs — it checks that the MBR
|
`--verify` mode the `check-iso-image` step runs — it checks that the MBR
|
||||||
partition and the El Torito catalog agree on where the FAT image lives and
|
partition and the El Torito catalog agree on where the FAT image lives and
|
||||||
@@ -5,7 +5,7 @@ isolation; fault → kill the process → keep the core (`onException`; the
|
|||||||
`fault-recovery` test); the supervisor notification **with exit reasons**
|
`fault-recovery` test); the supervisor notification **with exit reasons**
|
||||||
([process-lifecycle.md](process-lifecycle.md) — clean exit, fault class, or
|
([process-lifecycle.md](process-lifecycle.md) — clean exit, fault class, or
|
||||||
killed, recorded before the notice posts); and the **restart policy itself**
|
killed, recorded before the notice posts); and the **restart policy itself**
|
||||||
([device-manager.md](device-manager.md)): the device manager supervises every
|
([device-manager.md](../device-driver-development-guide/device-manager.md)): the device manager supervises every
|
||||||
driver, restarts crashes with backoff, caps crash loops, and re-claims work
|
driver, restarts crashes with backoff, caps crash loops, and re-claims work
|
||||||
because the kernel releases a dead process's claims. The `driver-restart` and
|
because the kernel releases a dead process's claims. The `driver-restart` and
|
||||||
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
|
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
|
||||||
@@ -14,9 +14,9 @@ more of the system moved into restartable processes (the discovery migration,
|
|||||||
[discovery.md](discovery.md), is the next rung). This is the property danos is really chasing:
|
[discovery.md](discovery.md), is the next rung). This is the property danos is really chasing:
|
||||||
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
|
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
|
||||||
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
|
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
|
||||||
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
|
the reason the [microkernel](../vision.md) shape was chosen, and it's a *separate* goal
|
||||||
from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) —
|
from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) —
|
||||||
one that's less pervasive to build (see [vision.md](vision.md)).
|
one that's less pervasive to build (see [vision.md](../vision.md)).
|
||||||
|
|
||||||
## The idea: "let it crash" + supervision
|
## The idea: "let it crash" + supervision
|
||||||
|
|
||||||
@@ -44,7 +44,7 @@ down. **Keeping the kernel minimal is a resilience strategy, not just an aesthet
|
|||||||
## The building blocks
|
## The building blocks
|
||||||
|
|
||||||
1. **Address-space isolation.** A fault in one component can't corrupt another or the
|
1. **Address-space isolation.** A fault in one component can't corrupt another or the
|
||||||
kernel. This is the [user-mode milestone](vision.md) (ring 3, per-process page
|
kernel. This is the [user-mode milestone](../vision.md) (ring 3, per-process page
|
||||||
tables) — the shared prerequisite for *any* of this, and it's needed regardless.
|
tables) — the shared prerequisite for *any* of this, and it's needed regardless.
|
||||||
2. **Fault detection** — how the system notices a component is dead or sick:
|
2. **Fault detection** — how the system notices a component is dead or sick:
|
||||||
- **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps
|
- **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps
|
||||||
@@ -81,7 +81,7 @@ Detecting and killing is the easy half. The genuinely tricky questions are about
|
|||||||
- **In-flight IPC**: messages sent to the dead component, or replies its clients are
|
- **In-flight IPC**: messages sent to the dead component, or replies its clients are
|
||||||
blocked waiting for. The channel has to break cleanly and unblock the waiters with
|
blocked waiting for. The channel has to break cleanly and unblock the waiters with
|
||||||
an error rather than hang them forever (a design constraint that reaches back into
|
an error rather than hang them forever (a design constraint that reaches back into
|
||||||
[ipc.md](ipc.md) — channels need a "peer died" outcome).
|
[ipc.md](../device-driver-development-guide/ipc.md) — channels need a "peer died" outcome).
|
||||||
- **Clients**: how does a client discover the service it was talking to is gone and
|
- **Clients**: how does a client discover the service it was talking to is gone and
|
||||||
has been replaced? Options: capability revocation makes stale handles fail; or a
|
has been replaced? Options: capability revocation makes stale handles fail; or a
|
||||||
**name server** re-binds clients to the new instance; or clients retry through a
|
**name server** re-binds clients to the new instance; or clients retry through a
|
||||||
@@ -137,7 +137,7 @@ Honest boundaries:
|
|||||||
|
|
||||||
Resilience needs **structural** features (isolation + supervision + a resource
|
Resilience needs **structural** features (isolation + supervision + a resource
|
||||||
model); real-time needs a **pervasive** timing invariant. They're separable, and
|
model); real-time needs a **pervasive** timing invariant. They're separable, and
|
||||||
resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](vision.md)).
|
resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](../vision.md)).
|
||||||
Note the overlap, though: **preemptive scheduling** and **priorities** — already
|
Note the overlap, though: **preemptive scheduling** and **priorities** — already
|
||||||
built — serve resilience too (you can preempt and kill a misbehaving component, and
|
built — serve resilience too (you can preempt and kill a misbehaving component, and
|
||||||
run the supervisor at high priority). So danos keeps the useful *mechanisms* of the
|
run the supervisor at high priority). So danos keeps the useful *mechanisms* of the
|
||||||
@@ -156,10 +156,10 @@ real-time work without owing anyone a timing *guarantee*.
|
|||||||
|
|
||||||
## Related
|
## Related
|
||||||
|
|
||||||
- [vision.md](vision.md) — the goals this serves (learning by doing; resilience over
|
- [vision.md](../vision.md) — the goals this serves (learning by doing; resilience over
|
||||||
hard real-time).
|
hard real-time).
|
||||||
- [scheduling.md](scheduling.md) — preemption, which makes runaway components killable.
|
- [scheduling.md](scheduling.md) — preemption, which makes runaway components killable.
|
||||||
- [ipc.md](ipc.md) — channels that need a "peer died" outcome for clean restart.
|
- [ipc.md](../device-driver-development-guide/ipc.md) — channels that need a "peer died" outcome for clean restart.
|
||||||
- [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and
|
- [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and
|
||||||
restart" instead of "halt".
|
restart" instead of "halt".
|
||||||
- [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context.
|
- [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context.
|
||||||
@@ -3,7 +3,7 @@
|
|||||||
The scheduler turns danos from a linear "boot then halt" kernel into a **running
|
The scheduler turns danos from a linear "boot then halt" kernel into a **running
|
||||||
multitasking system**. It's **fixed-priority preemptive**: the highest-priority
|
multitasking system**. It's **fixed-priority preemptive**: the highest-priority
|
||||||
ready task always runs, and tasks at the same priority take turns. That model is
|
ready task always runs, and tasks at the same priority take turns. That model is
|
||||||
chosen for [real-time](vision.md) — it's predictable (you can reason about which
|
chosen for [real-time](../vision.md) — it's predictable (you can reason about which
|
||||||
task runs when) and its decisions are O(1), unlike a fair-share scheduler.
|
task runs when) and its decisions are O(1), unlike a fair-share scheduler.
|
||||||
|
|
||||||
The scheduler proper (`system/kernel/scheduler.zig`) is generic; the context switch and new-task
|
The scheduler proper (`system/kernel/scheduler.zig`) is generic; the context switch and new-task
|
||||||
@@ -37,7 +37,7 @@ down a return address pointing at `task_trampoline` and zeroed callee-saved slot
|
|||||||
`schedule()` — pick the best task and switch — runs from two places:
|
`schedule()` — pick the best task and switch — runs from two places:
|
||||||
|
|
||||||
- **`yield()`** — a task voluntarily gives up the CPU.
|
- **`yield()`** — a task voluntarily gives up the CPU.
|
||||||
- **`tick()`** — the 1000 Hz [timer](device-interrupts.md) preempts the running
|
- **`tick()`** — the 1000 Hz [timer](../device-driver-development-guide/device-interrupts.md) preempts the running
|
||||||
task. This is what lets a task that never yields still share the CPU.
|
task. This is what lets a task that never yields still share the CPU.
|
||||||
|
|
||||||
The subtlety in mixing them is the **interrupt flag (IF)**. The rule: `switch_context`
|
The subtlety in mixing them is the **interrupt flag (IF)**. The rule: `switch_context`
|
||||||
@@ -97,7 +97,7 @@ marks the task blocked with a wake deadline and switches away. On every tick the
|
|||||||
timer wakes any task whose deadline has passed (a bounded scan, so it stays
|
timer wakes any task whose deadline has passed (a bounded scan, so it stays
|
||||||
deterministic), which makes it ready again; the scheduler then runs it when its
|
deterministic), which makes it ready again; the scheduler then runs it when its
|
||||||
priority comes up. `sleep` measures its deadline on the [calibrated
|
priority comes up. `sleep` measures its deadline on the [calibrated
|
||||||
clock](device-interrupts.md), so it's real time.
|
clock](../device-driver-development-guide/device-interrupts.md), so it's real time.
|
||||||
|
|
||||||
When *every* task is blocked, something still has to run — so there's an **idle
|
When *every* task is blocked, something still has to run — so there's an **idle
|
||||||
task** at the lowest priority that just `hlt`s until the next interrupt (see
|
task** at the lowest priority that just `hlt`s until the next interrupt (see
|
||||||
@@ -111,7 +111,7 @@ The other form of blocking is waiting for an **event** rather than a duration. A
|
|||||||
the caller on it, `wake(wq)` moves the highest-priority waiter back to ready
|
the caller on it, `wake(wq)` moves the highest-priority waiter back to ready
|
||||||
(preempting if it now outranks the running task). A task links into a wait queue
|
(preempting if it now outranks the running task). A task links into a wait queue
|
||||||
through the same field the ready queues use — it's in exactly one queue at a time.
|
through the same field the ready queues use — it's in exactly one queue at a time.
|
||||||
These are the primitives locks, semaphores and [IPC](ipc.md) are built on.
|
These are the primitives locks, semaphores and [IPC](../device-driver-development-guide/ipc.md) are built on.
|
||||||
|
|
||||||
Blocking safely needs **composable critical sections**. A blanket `cli`/`sti` pair
|
Blocking safely needs **composable critical sections**. A blanket `cli`/`sti` pair
|
||||||
doesn't nest: an IPC channel that `cli`s and then calls `wait` would have `wait`'s
|
doesn't nest: an IPC channel that `cli`s and then calls `wait` would have `wait`'s
|
||||||
@@ -124,7 +124,7 @@ caller's state.
|
|||||||
|
|
||||||
## Verifying it
|
## Verifying it
|
||||||
|
|
||||||
Three tests (see [testing.md](testing.md)) prove the guarantees:
|
Three tests (see [testing.md](../testing.md)) prove the guarantees:
|
||||||
|
|
||||||
- **`sched`** spawns three tasks that busy-loop *without ever yielding*. They all
|
- **`sched`** spawns three tasks that busy-loop *without ever yielding*. They all
|
||||||
make progress — which can only happen if the timer is **preempting** between them
|
make progress — which can only happen if the timer is **preempting** between them
|
||||||
@@ -140,7 +140,7 @@ Three tests (see [testing.md](testing.md)) prove the guarantees:
|
|||||||
|
|
||||||
- **Priority inheritance** — still open. Tasks now do block on shared resources
|
- **Priority inheritance** — still open. Tasks now do block on shared resources
|
||||||
(IPC rendezvous, the big kernel lock), and nothing yet bounds priority
|
(IPC rendezvous, the big kernel lock), and nothing yet bounds priority
|
||||||
inversion — a [real-time](vision.md) requirement.
|
inversion — a [real-time](../vision.md) requirement.
|
||||||
- **Task exit / a reaper** — done. A dying task goes on its core's reap list in a
|
- **Task exit / a reaper** — done. A dying task goes on its core's reap list in a
|
||||||
`.reaping` state; the timer tick drains the list, frees the stack back to the
|
`.reaping` state; the timer tick drains the list, frees the stack back to the
|
||||||
heap, and recycles the task-table slot.
|
heap, and recycles the task-table slot.
|
||||||
@@ -307,4 +307,4 @@ refcount, and no group-kill special case is needed at all.
|
|||||||
Then update [threading.md](threading.md) (the shared-fate gap note),
|
Then update [threading.md](threading.md) (the shared-fate gap note),
|
||||||
[process-lifecycle.md](process-lifecycle.md),
|
[process-lifecycle.md](process-lifecycle.md),
|
||||||
[process-management.md](process-management.md), and
|
[process-management.md](process-management.md), and
|
||||||
[ipc.md](ipc.md)/[drivers.md](drivers.md) mentions.
|
[ipc.md](../device-driver-development-guide/ipc.md)/[drivers.md](../device-driver-development-guide/drivers.md) mentions.
|
||||||
@@ -112,7 +112,7 @@ Yes — and this is the branch that matters for danos right now.
|
|||||||
|
|
||||||
These pull in different directions, so **picking the primary goal comes before
|
These pull in different directions, so **picking the primary goal comes before
|
||||||
picking the SMP design.** (danos's founding assumption was real-time; that's under
|
picking the SMP design.** (danos's founding assumption was real-time; that's under
|
||||||
active reconsideration in favour of resilience — see [vision.md](vision.md).)
|
active reconsideration in favour of resilience — see [vision.md](../vision.md).)
|
||||||
|
|
||||||
## What this would mean for danos
|
## What this would mean for danos
|
||||||
|
|
||||||
@@ -169,7 +169,7 @@ next lands.
|
|||||||
the highest-priority ready task; per-core queues are a later optimisation.
|
the highest-priority ready task; per-core queues are a later optimisation.
|
||||||
- **AP wake to long mode** — `architecture.startSecondary` drives INIT–SIPI–SIPI (via the
|
- **AP wake to long mode** — `architecture.startSecondary` drives INIT–SIPI–SIPI (via the
|
||||||
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
|
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
|
||||||
real mode at a low page and runs the [trampoline](../system/kernel/architecture/x86_64/trampoline.s)
|
real mode at a low page and runs the [trampoline](../../system/kernel/architecture/x86_64/trampoline.s)
|
||||||
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
|
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
|
||||||
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
|
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
|
||||||
all four cores report `online`.
|
all four cores report `online`.
|
||||||
@@ -270,6 +270,6 @@ next lands.
|
|||||||
|
|
||||||
- [scheduling.md](scheduling.md) — the single-core scheduler SMP would extend.
|
- [scheduling.md](scheduling.md) — the single-core scheduler SMP would extend.
|
||||||
- [discovery.md](discovery.md) — enumerating cores is a device-discovery problem.
|
- [discovery.md](discovery.md) — enumerating cores is a device-discovery problem.
|
||||||
- [ipc.md](ipc.md) — the message passing cross-core coordination rides on.
|
- [ipc.md](../device-driver-development-guide/ipc.md) — the message passing cross-core coordination rides on.
|
||||||
- [vision.md](vision.md) — the goals question (real-time vs resilience) this note
|
- [vision.md](../vision.md) — the goals question (real-time vs resilience) this note
|
||||||
keeps bumping into.
|
keeps bumping into.
|
||||||
@@ -51,7 +51,7 @@ Everything else---including`read()`,`write()`,`malloc()`, and`fork()`---will run
|
|||||||
3. **`Yield()`/`Thread_Ctrl()`**
|
3. **`Yield()`/`Thread_Ctrl()`**
|
||||||
- **What it does:**Allows a thread to voluntarily give up its CPU time slice, or allows a root task to spawn/kill threads.
|
- **What it does:**Allows a thread to voluntarily give up its CPU time slice, or allows a root task to spawn/kill threads.
|
||||||
4. **`ipc_send(endpoint, message_buffer)`(Asynchronous Send)**
|
4. **`ipc_send(endpoint, message_buffer)`(Asynchronous Send)**
|
||||||
- **What it does:**Posts a small payload to an endpoint's bounded queue and returns *without* blocking — no rendezvous, no reply. The receiver picks it up through the same `IPC_ReplyWait`, as a buffered message. It is the async counterpart of `IPC_Call`, for one-to-many broadcasts where a synchronous rendezvous would let one dead or slow receiver hang the sender. The [input service](input.md) — keyboard-event fan-out — is its first user. A full queue drops the oldest message (a buffered message is discrete data, unlike a coalescing interrupt notification).
|
- **What it does:**Posts a small payload to an endpoint's bounded queue and returns *without* blocking — no rendezvous, no reply. The receiver picks it up through the same `IPC_ReplyWait`, as a buffered message. It is the async counterpart of `IPC_Call`, for one-to-many broadcasts where a synchronous rendezvous would let one dead or slow receiver hang the sender. The [input service](../device-driver-development-guide/input.md) — keyboard-event fan-out — is its first user. A full queue drops the oldest message (a buffered message is discrete data, unlike a coalescing interrupt notification).
|
||||||
|
|
||||||
* * * * *
|
* * * * *
|
||||||
|
|
||||||
@@ -92,4 +92,4 @@ Managing the Payload Challenge
|
|||||||
Because it is a microkernel, performance lives or dies by how fast your`IPC_Call`can move data from Client to Server. You have two minimal choices for handling the`message_buffer`pointer:[[1](https://anazimzada2020.medium.com/microkernel-architectural-pattern-5e4e9184170e)]
|
Because it is a microkernel, performance lives or dies by how fast your`IPC_Call`can move data from Client to Server. You have two minimal choices for handling the`message_buffer`pointer:[[1](https://anazimzada2020.medium.com/microkernel-architectural-pattern-5e4e9184170e)]
|
||||||
|
|
||||||
- **The Copy Method (Simplest to start):**Your kernel pauses the client, reads the data from the client's memory space, switches page tables to the server, and copies the data into the server's buffer.
|
- **The Copy Method (Simplest to start):**Your kernel pauses the client, reads the data from the client's memory space, switches page tables to the server, and copies the data into the server's buffer.
|
||||||
- **The Shared Memory Method (Fastest):**The kernel sets up a temporary, shared virtual memory page between the client and server. The client writes to it, calls`syscall`/`svc`, and the server reads it instantly without the kernel copying any bytes
|
- **The Shared Memory Method (Fastest):**The kernel sets up a temporary, shared virtual memory page between the client and server. The client writes to it, calls`syscall`/`svc`, and the server reads it instantly without the kernel copying any bytes
|
||||||
@@ -11,7 +11,7 @@ sequential pass and hands the bytes to the kernel unmodified.
|
|||||||
|
|
||||||
The capsule is a *performance artifact*, not a source of truth. The boot
|
The capsule is a *performance artifact*, not a source of truth. The boot
|
||||||
volume's `/system` and `/test` file trees remain the canonical layout (see
|
volume's `/system` and `/test` file trees remain the canonical layout (see
|
||||||
[danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md));
|
[danos-file-system-hierarchy-FSH.md](../file-system-development/danos-file-system-hierarchy-FSH.md));
|
||||||
the capsule is a pre-baked snapshot of the same binaries, derived from the same
|
the capsule is a pre-baked snapshot of the same binaries, derived from the same
|
||||||
build graph, so the running system is identical whether the loader read the
|
build graph, so the running system is identical whether the loader read the
|
||||||
capsule or walked the tree.
|
capsule or walked the tree.
|
||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone
|
The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone
|
||||||
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
|
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
|
||||||
[display-v2-plan.md](display-v2-plan.md). Read threading.md first for the *why*.
|
[display-v2-plan.md](../device-driver-development-guide/display-v2-plan.md). Read threading.md first for the *why*.
|
||||||
|
|
||||||
## Locked decisions (do not relitigate)
|
## Locked decisions (do not relitigate)
|
||||||
|
|
||||||
@@ -13,7 +13,7 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run,
|
|||||||
`single_threaded = false`.
|
`single_threaded = false`.
|
||||||
- **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an
|
- **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an
|
||||||
idle core still halts ([halting.md](halting.md)).
|
idle core still halts ([halting.md](halting.md)).
|
||||||
- **New syscalls are private**: extend [abi.zig](../system/abi.zig) `SystemCall` after
|
- **New syscalls are private**: extend [abi.zig](../../system/abi.zig) `SystemCall` after
|
||||||
`shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`,
|
`shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`,
|
||||||
`futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number.
|
`futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number.
|
||||||
- **Restart granularity stays the process** — a faulting thread kills its process; the
|
- **Restart granularity stays the process** — a faulting thread kills its process; the
|
||||||
@@ -21,10 +21,10 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run,
|
|||||||
|
|
||||||
## Conventions
|
## Conventions
|
||||||
|
|
||||||
Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations,
|
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations,
|
||||||
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
|
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
|
||||||
`addUserBinary` (with the new `threaded` flag where a binary spawns threads) and get
|
`addUserBinary` (with the new `threaded` flag where a binary spawns threads) and get
|
||||||
packed into the initial-ramdisk; new syscalls extend [abi.zig](../system/abi.zig)
|
packed into the initial-ramdisk; new syscalls extend [abi.zig](../../system/abi.zig)
|
||||||
`SystemCall` + a `library/runtime` wrapper; test services live beside the code they
|
`SystemCall` + a `library/runtime` wrapper; test services live beside the code they
|
||||||
exercise and register a `ServiceId` if they must be looked up.
|
exercise and register a `ServiceId` if they must be looked up.
|
||||||
|
|
||||||
@@ -58,7 +58,7 @@ fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration
|
|||||||
5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into
|
5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into
|
||||||
`main` stays a human step.
|
`main` stays a human step.
|
||||||
3. **Implement** every unchecked item in that milestone, including adding its
|
3. **Implement** every unchecked item in that milestone, including adding its
|
||||||
`-Dtest-case` to `CASES` in [test/qemu_test.py](../test/qemu_test.py) (with
|
`-Dtest-case` to `CASES` in [test/qemu_test.py](../../test/qemu_test.py) (with
|
||||||
`smp: true` / a `mem` bump where noted) so the gate is runnable.
|
`smp: true` / a `mem` bump where noted) so the gate is runnable.
|
||||||
4. **Run the gate**: `python3 test/qemu_test.py <case>`, then the full **guardrail
|
4. **Run the gate**: `python3 test/qemu_test.py <case>`, then the full **guardrail
|
||||||
set**, then `zig build` (clean) and `zig build test` (green).
|
set**, then `zig build` (clean) and `zig build test` (green).
|
||||||
@@ -67,7 +67,7 @@ fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration
|
|||||||
the whole guardrail set passes, `zig build` is clean, and host tests are green.
|
the whole guardrail set passes, `zig build` is clean, and host tests are green.
|
||||||
→ tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit`
|
→ tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit`
|
||||||
(`threads(M<n>): <summary>`, no `Co-Authored-By` trailer per
|
(`threads(M<n>): <summary>`, no `Co-Authored-By` trailer per
|
||||||
[coding-standards.md](coding-standards.md)), then **`git push` the working branch to
|
[coding-standards.md](../coding-standards.md)), then **`git push` the working branch to
|
||||||
`origin`** (use `-u` on the first push to set upstream). Continue to the next
|
`origin`** (use `-u` on the first push to set upstream). Continue to the next
|
||||||
milestone in the same iteration if budget remains; otherwise let the loop re-fire.
|
milestone in the same iteration if budget remains; otherwise let the loop re-fire.
|
||||||
- **Red** = anything above fails. Diagnose from the captured serial log
|
- **Red** = anything above fails. Diagnose from the captured serial log
|
||||||
@@ -111,7 +111,7 @@ task's exit; make destruction happen on the **last** exit.
|
|||||||
- [x] A refcount keyed by the address-space root, held in `scheduler.zig`
|
- [x] A refcount keyed by the address-space root, held in `scheduler.zig`
|
||||||
(`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the
|
(`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the
|
||||||
success path, after the slot + stack are secured), all under the big kernel lock.
|
success path, after the slot + stack are secured), all under the big kernel lock.
|
||||||
- [x] Both task-teardown paths ([scheduler.zig](../system/kernel/scheduler.zig):
|
- [x] Both task-teardown paths ([scheduler.zig](../../system/kernel/scheduler.zig):
|
||||||
`exitUserLocked` and `destroyTaskLocked`) call `releaseAddressSpace`, which decrements
|
`exitUserLocked` and `destroyTaskLocked`) call `releaseAddressSpace`, which decrements
|
||||||
and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test
|
and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test
|
||||||
spaces) is destroyed directly, preserving prior behaviour.
|
spaces) is destroyed directly, preserving prior behaviour.
|
||||||
@@ -133,7 +133,7 @@ full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `sm
|
|||||||
Spawn only — no join yet. Prove a second task executes in the **caller's** address
|
Spawn only — no join yet. Prove a second task executes in the **caller's** address
|
||||||
space and exits cleanly.
|
space and exits cleanly.
|
||||||
|
|
||||||
- [x] [abi.zig](../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in
|
- [x] [abi.zig](../../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in
|
||||||
process.zig; `thread_spawn` calls `scheduler.spawnThread` (today, after M3, the
|
process.zig; `thread_spawn` calls `scheduler.spawnThread` (today, after M3, the
|
||||||
handler goes `spawnThreadSupervised` → `scheduler.spawnUserLocked`; shares the caller's
|
handler goes `spawnThreadSupervised` → `scheduler.spawnUserLocked`; shares the caller's
|
||||||
address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)`
|
address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)`
|
||||||
@@ -195,7 +195,7 @@ plus `address-space-refcount`/`thread-spawn`; `zig build` clean, `zig build test
|
|||||||
|
|
||||||
## M4 — Futex: the one blocking primitive ✅
|
## M4 — Futex: the one blocking primitive ✅
|
||||||
|
|
||||||
- [x] [abi.zig](../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a
|
- [x] [abi.zig](../../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a
|
||||||
`.blocked` task tagged with `Task.futex_addr` (no queue linkage);
|
`.blocked` task tagged with `Task.futex_addr` (no queue linkage);
|
||||||
`futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock,
|
`futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock,
|
||||||
parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr,
|
parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr,
|
||||||
@@ -261,7 +261,7 @@ green.
|
|||||||
restores it.
|
restores it.
|
||||||
- [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same
|
- [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same
|
||||||
`Futex`/`Mutex`/`Condition` when wanted.
|
`Futex`/`Mutex`/`Condition` when wanted.
|
||||||
- [x] All `thread-*` cases wired into [test/qemu_test.py](../test/qemu_test.py)
|
- [x] All `thread-*` cases wired into [test/qemu_test.py](../../test/qemu_test.py)
|
||||||
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md
|
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md
|
||||||
status updated to **built**; the worked example is threading.md's win-condition.
|
status updated to **built**; the worked example is threading.md's win-condition.
|
||||||
- [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main
|
- [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main
|
||||||
@@ -405,7 +405,7 @@ guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-re
|
|||||||
### M10 — Per-thread TLS: the thread-pointer mechanism ✅
|
### M10 — Per-thread TLS: the thread-pointer mechanism ✅
|
||||||
|
|
||||||
Give each thread its own thread pointer and private TLS storage — the foundation
|
Give each thread its own thread pointer and private TLS storage — the foundation
|
||||||
self-hosting Zig ([zig-self-hosting.md](zig-self-hosting.md)) will build `threadlocal` on.
|
self-hosting Zig ([zig-self-hosting.md](../zig-self-hosting.md)) will build `threadlocal` on.
|
||||||
|
|
||||||
- [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch
|
- [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch
|
||||||
**only when it changes** (the same conditional-load discipline as CR3;
|
**only when it changes** (the same conditional-load discipline as CR3;
|
||||||
@@ -459,7 +459,7 @@ clean.
|
|||||||
## Deferred (explicitly not in this plan)
|
## Deferred (explicitly not in this plan)
|
||||||
|
|
||||||
- **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a
|
- **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a
|
||||||
physical-address key so two processes share a futex through a [shared-memory](display-v2.md)
|
physical-address key so two processes share a futex through a [shared-memory](../device-driver-development-guide/display-v2.md)
|
||||||
region. Not needed for intra-process threads.
|
region. Not needed for intra-process threads.
|
||||||
- **Per-thread priorities / affinity distinct from the process** — threads inherit the
|
- **Per-thread priorities / affinity distinct from the process** — threads inherit the
|
||||||
process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep.
|
process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep.
|
||||||
@@ -467,5 +467,5 @@ clean.
|
|||||||
([process-lifecycle.md](process-lifecycle.md)).
|
([process-lifecycle.md](process-lifecycle.md)).
|
||||||
- **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more.
|
- **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more.
|
||||||
- **A real `std.Thread` backend** — arrives with self-hosting
|
- **A real `std.Thread` backend** — arrives with self-hosting
|
||||||
([zig-self-hosting.md](zig-self-hosting.md)); it sits on these same primitives, so it
|
([zig-self-hosting.md](../zig-self-hosting.md)); it sits on these same primitives, so it
|
||||||
swaps the impl under `runtime.Thread`, not the call sites.
|
swaps the impl under `runtime.Thread`, not the call sites.
|
||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
A note on danos **threads** — several tasks sharing one address space — provided by a
|
A note on danos **threads** — several tasks sharing one address space — provided by a
|
||||||
`Thread` type that mirrors the shape of Zig's `std.Thread` while keeping every
|
`Thread` type that mirrors the shape of Zig's `std.Thread` while keeping every
|
||||||
kernel entry behind the [runtime](../library/kernel). **Built** (M1–M11, see
|
kernel entry behind the [runtime](../../library/kernel). **Built** (M1–M11, see
|
||||||
[threading-plan.md](threading-plan.md)): `spawn`/`join`/`detach`, cross-core parallelism,
|
[threading-plan.md](threading-plan.md)): `spawn`/`join`/`detach`, cross-core parallelism,
|
||||||
a futex, `Mutex`/`Condition`/`Semaphore`/`RwLock`/`WaitGroup`, `getCurrentId`/`currentCore`,
|
a futex, `Mutex`/`Condition`/`Semaphore`/`RwLock`/`WaitGroup`, `getCurrentId`/`currentCore`,
|
||||||
per-thread thread-pointer TLS, thread-safe allocation, and a task reaper that reclaims dead
|
per-thread thread-pointer TLS, thread-safe allocation, and a task reaper that reclaims dead
|
||||||
@@ -36,12 +36,12 @@ implementation underneath, not the API above.
|
|||||||
[Why not literal std.Thread](#why-not-literal-stdthread).
|
[Why not literal std.Thread](#why-not-literal-stdthread).
|
||||||
- **Threads are a narrow, opt-in capability — not the default concurrency tool.** The
|
- **Threads are a narrow, opt-in capability — not the default concurrency tool.** The
|
||||||
default for resilience stays **process + IPC** ([resilience.md](resilience.md),
|
default for resilience stays **process + IPC** ([resilience.md](resilience.md),
|
||||||
[ipc.md](ipc.md)). See [Where threads fit](#where-threads-fit-the-resilience-tension).
|
[ipc.md](../device-driver-development-guide/ipc.md)). See [Where threads fit](#where-threads-fit-the-resilience-tension).
|
||||||
- **Blocking synchronization is futex-backed, never spin-backed.** Waiters sleep in
|
- **Blocking synchronization is futex-backed, never spin-backed.** Waiters sleep in
|
||||||
the kernel so an idle core still halts ([halting.md](halting.md)).
|
the kernel so an idle core still halts ([halting.md](halting.md)).
|
||||||
- **Per-binary opt-in to multi-threaded codegen.** Only a service that asks for
|
- **Per-binary opt-in to multi-threaded codegen.** Only a service that asks for
|
||||||
threads is built `single_threaded = false`; the rest stay lean and single-threaded.
|
threads is built `single_threaded = false`; the rest stay lean and single-threaded.
|
||||||
- **The thread ABI is private.** New syscalls extend [abi.zig](../system/abi.zig)
|
- **The thread ABI is private.** New syscalls extend [abi.zig](../../system/abi.zig)
|
||||||
`SystemCall` and are reached only through `library/kernel` wrappers, exactly like
|
`SystemCall` and are reached only through `library/kernel` wrappers, exactly like
|
||||||
every other danos syscall ([syscall.md](syscall.md)) — numbers stay renumberable.
|
every other danos syscall ([syscall.md](syscall.md)) — numbers stay renumberable.
|
||||||
|
|
||||||
@@ -56,12 +56,12 @@ runtime — rebuilt in lockstep — knows the mapping.
|
|||||||
`std.Thread` is incompatible with that invariant on two counts:
|
`std.Thread` is incompatible with that invariant on two counts:
|
||||||
|
|
||||||
1. **It selects its backend from `builtin.os.tag`, and issues syscalls directly.**
|
1. **It selects its backend from `builtin.os.tag`, and issues syscalls directly.**
|
||||||
danos targets `.os_tag = .freestanding` ([build.zig](../build.zig)), for which
|
danos targets `.os_tag = .freestanding` ([build.zig](../../build.zig)), for which
|
||||||
`std.Thread` resolves to an unsupported stub that `@compileError`s. Adding a real
|
`std.Thread` resolves to an unsupported stub that `@compileError`s. Adding a real
|
||||||
backend would either bake danos syscall numbers into std (breaking ABI privacy and
|
backend would either bake danos syscall numbers into std (breaking ABI privacy and
|
||||||
renumbering) or fork std to route back through the runtime — a permanent rebase
|
renumbering) or fork std to route back through the runtime — a permanent rebase
|
||||||
cost that buys nothing the native type doesn't.
|
cost that buys nothing the native type doesn't.
|
||||||
2. **Our user binaries are built `single_threaded = true`** ([build.zig](../build.zig)
|
2. **Our user binaries are built `single_threaded = true`** ([build.zig](../../build.zig)
|
||||||
`addUserBinary`), which compiles threading out entirely and makes atomics and TLS
|
`addUserBinary`), which compiles threading out entirely and makes atomics and TLS
|
||||||
single-threaded. Threads need this flipped per binary regardless.
|
single-threaded. Threads need this flipped per binary regardless.
|
||||||
|
|
||||||
@@ -134,7 +134,7 @@ Deviations from `std.Thread`, called out honestly:
|
|||||||
|
|
||||||
## Kernel primitives (new private syscalls)
|
## Kernel primitives (new private syscalls)
|
||||||
|
|
||||||
Five core entries extend [abi.zig](../system/abi.zig) `SystemCall` after
|
Five core entries extend [abi.zig](../../system/abi.zig) `SystemCall` after
|
||||||
`shared_memory_physical = 36` (plus small helpers `current_core`, `thread_self`, and
|
`shared_memory_physical = 36` (plus small helpers `current_core`, `thread_self`, and
|
||||||
`set_thread_pointer`), each with a `library/kernel` wrapper:
|
`set_thread_pointer`), each with a `library/kernel` wrapper:
|
||||||
|
|
||||||
@@ -155,12 +155,12 @@ Plus one invariant change with no new syscall: **address-space reference countin
|
|||||||
Before this work an address space was 1:1 with a task: `spawnUserLocked` records
|
Before this work an address space was 1:1 with a task: `spawnUserLocked` records
|
||||||
`address_space` on the Task (as it still does), and teardown did
|
`address_space` on the Task (as it still does), and teardown did
|
||||||
`destroyAddressSpace(t.address_space)` when **any** user task exited
|
`destroyAddressSpace(t.address_space)` when **any** user task exited
|
||||||
([scheduler.zig](../system/kernel/scheduler.zig)). With threads, several tasks share
|
([scheduler.zig](../../system/kernel/scheduler.zig)). With threads, several tasks share
|
||||||
one `address_space`, so the first to exit would rip the address space out from under its
|
one `address_space`, so the first to exit would rip the address space out from under its
|
||||||
siblings.
|
siblings.
|
||||||
|
|
||||||
Fix: a small refcount keyed by the address-space root, kept in
|
Fix: a small refcount keyed by the address-space root, kept in
|
||||||
[scheduler.zig](../system/kernel/scheduler.zig): `retainAddressSpace` takes a
|
[scheduler.zig](../../system/kernel/scheduler.zig): `retainAddressSpace` takes a
|
||||||
reference for every user task `spawnUserLocked` starts (count 1 on the first take, so
|
reference for every user task `spawnUserLocked` starts (count 1 on the first take, so
|
||||||
a thread sharing the caller's space increments it); task teardown calls
|
a thread sharing the caller's space increments it); task teardown calls
|
||||||
`releaseAddressSpace`, which only calls `destroyAddressSpace` at **zero**. All of
|
`releaseAddressSpace`, which only calls `destroyAddressSpace` at **zero**. All of
|
||||||
@@ -172,7 +172,7 @@ that must land and be proven before anything shares an address space.
|
|||||||
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle
|
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle
|
||||||
values through scratch registers — `startUserTask` reads the entry/stack (and the
|
values through scratch registers — `startUserTask` reads the entry/stack (and the
|
||||||
thread's closure arg, delivered in `rdi` via `jumpToUserArg`) from the Task
|
thread's closure arg, delivered in `rdi` via `jumpToUserArg`) from the Task
|
||||||
([scheduler.zig](../system/kernel/scheduler.zig)). That makes the thread path clean:
|
([scheduler.zig](../../system/kernel/scheduler.zig)). That makes the thread path clean:
|
||||||
|
|
||||||
1. The runtime's `spawn` `mmap`s a stack (syscall `4`) and writes the closure —
|
1. The runtime's `spawn` `mmap`s a stack (syscall `4`) and writes the closure —
|
||||||
`{ tls_base, args }`, the std "Instance" pattern — at the **top of the new stack
|
`{ tls_base, args }`, the std "Instance" pattern — at the **top of the new stack
|
||||||
@@ -214,7 +214,7 @@ Keying: threads share an address space, so a **virtual address within that addre
|
|||||||
identifies a futex uniquely; the kernel keys its wait queue by `(address_space_root, virtual_address)`.
|
identifies a futex uniquely; the kernel keys its wait queue by `(address_space_root, virtual_address)`.
|
||||||
Keying by the **physical** address instead (translate `virtual_address -> physical_address` on entry) is a
|
Keying by the **physical** address instead (translate `virtual_address -> physical_address` on entry) is a
|
||||||
deliberate forward door: it lets two *processes* share a futex through an
|
deliberate forward door: it lets two *processes* share a futex through an
|
||||||
[shared-memory](display-v2.md) region later, without changing the API. We start with the
|
[shared-memory](../device-driver-development-guide/display-v2.md) region later, without changing the API. We start with the
|
||||||
private-per-address-space key and note the physical-key upgrade.
|
private-per-address-space key and note the physical-key upgrade.
|
||||||
|
|
||||||
No spinning: a contended lock parks the task in the kernel and the core is free to run
|
No spinning: a contended lock parks the task in the kernel and the core is free to run
|
||||||
@@ -264,9 +264,9 @@ stays single-threaded and lean.
|
|||||||
**process**, which respawns its threads from a known-good state — restart
|
**process**, which respawns its threads from a known-good state — restart
|
||||||
granularity stays the process. The leader's recorded exit reason carries the fault
|
granularity stays the process. The leader's recorded exit reason carries the fault
|
||||||
class even when a worker faulted, so restart policy is unchanged.
|
class even when a worker faulted, so restart policy is unchanged.
|
||||||
- **IPC — two consequences threads forced ([ipc.md](ipc.md)):**
|
- **IPC — two consequences threads forced ([ipc.md](../device-driver-development-guide/ipc.md)):**
|
||||||
- *Handles do not cross threads.* The handle table lives on the `Task`
|
- *Handles do not cross threads.* The handle table lives on the `Task`
|
||||||
([scheduler.zig](../system/kernel/scheduler.zig)), so a handle number is meaningful
|
([scheduler.zig](../../system/kernel/scheduler.zig)), so a handle number is meaningful
|
||||||
only to the thread that created it — thread A's endpoint handle `3` is not thread B's.
|
only to the thread that created it — thread A's endpoint handle `3` is not thread B's.
|
||||||
A thread that needs to reach an endpoint another thread owns looks it up
|
A thread that needs to reach an endpoint another thread owns looks it up
|
||||||
(`ipc.lookup(service)`) to install its **own** handle to the same underlying endpoint.
|
(`ipc.lookup(service)`) to install its **own** handle to the same underlying endpoint.
|
||||||
@@ -284,7 +284,7 @@ stays single-threaded and lean.
|
|||||||
|
|
||||||
The ordered, `/loop`-runnable milestones live in
|
The ordered, `/loop`-runnable milestones live in
|
||||||
**[threading-plan.md](threading-plan.md)** (shaped like
|
**[threading-plan.md](threading-plan.md)** (shaped like
|
||||||
[display-v2-plan.md](display-v2-plan.md)): every milestone lands on its own and ends in
|
[display-v2-plan.md](../device-driver-development-guide/display-v2-plan.md)): every milestone lands on its own and ends in
|
||||||
a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
|
a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
|
||||||
`zig build test` for host unit tests). The stages below are the shape it expands.
|
`zig build test` for host unit tests). The stages below are the shape it expands.
|
||||||
|
|
||||||
@@ -307,13 +307,13 @@ a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
|
|||||||
the consumer blocked, e.g. via a low idle tick count).
|
the consumer blocked, e.g. via a low idle tick count).
|
||||||
- **Stage 3 — polish.** Per-thread TLS / thread pointer and `threadlocal` (only if a
|
- **Stage 3 — polish.** Per-thread TLS / thread pointer and `threadlocal` (only if a
|
||||||
consumer needs it), `RwLock`/`WaitGroup` as demanded, and this doc's cases wired
|
consumer needs it), `RwLock`/`WaitGroup` as demanded, and this doc's cases wired
|
||||||
into [test/qemu_test.py](../test/qemu_test.py).
|
into [test/qemu_test.py](../../test/qemu_test.py).
|
||||||
|
|
||||||
## Conventions
|
## Conventions
|
||||||
|
|
||||||
Follow [coding-standards.md](coding-standards.md): spell out non-acronym
|
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym
|
||||||
abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New syscalls
|
abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New syscalls
|
||||||
extend [abi.zig](../system/abi.zig) `SystemCall` + a `library/kernel` wrapper
|
extend [abi.zig](../../system/abi.zig) `SystemCall` + a `library/kernel` wrapper
|
||||||
([syscall.md](syscall.md)). `Thread` is a first-class runtime module, the same
|
([syscall.md](syscall.md)). `Thread` is a first-class runtime module, the same
|
||||||
way `process` ([process-lifecycle.md](process-lifecycle.md)) and `ipc`
|
way `process` ([process-lifecycle.md](process-lifecycle.md)) and `ipc`
|
||||||
are — user code never names a syscall.
|
are — user code never names a syscall.
|
||||||
@@ -331,7 +331,7 @@ are — user code never names a syscall.
|
|||||||
## The self-hosting endgame
|
## The self-hosting endgame
|
||||||
|
|
||||||
When danos becomes a real Zig target and we (eventually) add a danos backend to std
|
When danos becomes a real Zig target and we (eventually) add a danos backend to std
|
||||||
([zig-self-hosting.md](zig-self-hosting.md)), `std.Thread` can sit *on top of* these
|
([zig-self-hosting.md](../zig-self-hosting.md)), `std.Thread` can sit *on top of* these
|
||||||
same kernel primitives — the danos `std.Thread.Impl` would call the very
|
same kernel primitives — the danos `std.Thread.Impl` would call the very
|
||||||
`thread_spawn`/`futex_*` wrappers `Thread` already uses. Because
|
`thread_spawn`/`futex_*` wrappers `Thread` already uses. Because
|
||||||
`Thread` was built API-compatible from day one, that transition swaps the
|
`Thread` was built API-compatible from day one, that transition swaps the
|
||||||
@@ -341,9 +341,9 @@ the later self-hosting lift cheap.
|
|||||||
## Further reading
|
## Further reading
|
||||||
|
|
||||||
- [scheduling.md](scheduling.md), [smp.md](smp.md) — the task model these threads join.
|
- [scheduling.md](scheduling.md), [smp.md](smp.md) — the task model these threads join.
|
||||||
- [resilience.md](resilience.md), [vision.md](vision.md) — why isolation is the default
|
- [resilience.md](resilience.md), [vision.md](../vision.md) — why isolation is the default
|
||||||
and threads are the exception.
|
and threads are the exception.
|
||||||
- [syscall.md](syscall.md), [ipc.md](ipc.md) — the private ABI and the messaging model
|
- [syscall.md](syscall.md), [ipc.md](../device-driver-development-guide/ipc.md) — the private ABI and the messaging model
|
||||||
threads sit beside.
|
threads sit beside.
|
||||||
- [halting.md](halting.md) — the idle/halt property futex-backed blocking preserves.
|
- [halting.md](halting.md) — the idle/halt property futex-backed blocking preserves.
|
||||||
- [zig-self-hosting.md](zig-self-hosting.md) — the target this bends toward.
|
- [zig-self-hosting.md](../zig-self-hosting.md) — the target this bends toward.
|
||||||
@@ -7,7 +7,7 @@ Two different needs hide under the word "timer", and danos keeps them apart:
|
|||||||
|
|
||||||
Both are answered by the **kernel**, because the kernel already owns a timer: it has
|
Both are answered by the **kernel**, because the kernel already owns a timer: it has
|
||||||
to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this
|
to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this
|
||||||
are built in [device-interrupts.md](device-interrupts.md); the scheduler's blocking and
|
are built in [device-interrupts.md](../device-driver-development-guide/device-interrupts.md); the scheduler's blocking and
|
||||||
wait queues are in [scheduling.md](scheduling.md). This page is about the surface a
|
wait queues are in [scheduling.md](scheduling.md). This page is about the surface a
|
||||||
ring-3 program actually uses, and one deliberate absence: **there is no user-space time
|
ring-3 program actually uses, and one deliberate absence: **there is no user-space time
|
||||||
service.**
|
service.**
|
||||||
@@ -32,11 +32,11 @@ danos checks both — the invariant-TSC CPUID bit (`0x80000007` EDX[8], set on I
|
|||||||
AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET
|
AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET
|
||||||
counter when either fails. So `now()` stays accurate on a real Intel box, a real AMD box,
|
counter when either fails. So `now()` stays accurate on a real Intel box, a real AMD box,
|
||||||
and inside a VM alike; only the source behind it differs. The mechanism is in
|
and inside a VM alike; only the source behind it differs. The mechanism is in
|
||||||
[device-interrupts.md](device-interrupts.md).
|
[device-interrupts.md](../device-driver-development-guide/device-interrupts.md).
|
||||||
|
|
||||||
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
|
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
|
||||||
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
|
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
|
||||||
model; that role now lives in [drivers.md](drivers.md), as documentation.) The one place
|
model; that role now lives in [drivers.md](../device-driver-development-guide/drivers.md), as documentation.) The one place
|
||||||
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
|
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
|
||||||
at the end; it is deliberately not built yet.
|
at the end; it is deliberately not built yet.
|
||||||
|
|
||||||
@@ -55,7 +55,7 @@ Time and waiting are three entries in the small syscall table ([syscall.md](sysc
|
|||||||
service can keep answering messages on the same endpoint while a deadline is pending.
|
service can keep answering messages on the same endpoint while a deadline is pending.
|
||||||
This is the timed wait that stop-sequence escalation, hello deadlines, and restart
|
This is the timed wait that stop-sequence escalation, hello deadlines, and restart
|
||||||
backoff are built from ([process-lifecycle.md](process-lifecycle.md),
|
backoff are built from ([process-lifecycle.md](process-lifecycle.md),
|
||||||
[device-manager.md](device-manager.md)).
|
[device-manager.md](../device-driver-development-guide/device-manager.md)).
|
||||||
|
|
||||||
The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space;
|
The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space;
|
||||||
programs read the TSC through `clock` and get timed wakeups through `sleep`/`timer_bind`,
|
programs read the TSC through `clock` and get timed wakeups through `sleep`/`timer_bind`,
|
||||||
@@ -50,7 +50,7 @@ The public danos ABI then has exactly two layers, neither of which is
|
|||||||
| Layer | Contract | Spoken by |
|
| Layer | Contract | Spoken by |
|
||||||
|-------|----------|-----------|
|
|-------|----------|-----------|
|
||||||
| **vDSO** | C-ABI functions, this note | every language's thin shim (the `system-call` module for Zig, a `-sys` crate for Rust, a header for C) |
|
| **vDSO** | C-ABI functions, this note | every language's thin shim (the `system-call` module for Zig, a `-sys` crate for Rust, a header for C) |
|
||||||
| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](vfs-protocol.md) is the first one documented) | any client that can lay out bytes |
|
| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](../file-system-development/vfs-protocol.md) is the first one documented) | any client that can lay out bytes |
|
||||||
|
|
||||||
Everything above those — the heap, `file_system`, the service harness — is
|
Everything above those — the heap, `file_system`, the service harness — is
|
||||||
per-language convenience, compiled into each binary from source, exactly as
|
per-language convenience, compiled into each binary from source, exactly as
|
||||||
@@ -118,7 +118,7 @@ hypervisor configured for UEFI firmware and an xHCI USB controller.
|
|||||||
(`system/kernel/acpi.zig:3`)
|
(`system/kernel/acpi.zig:3`)
|
||||||
- The loader reads `/system/kernel` off the FAT boot volume, then loads user
|
- The loader reads `/system/kernel` off the FAT boot volume, then loads user
|
||||||
space: a prebuilt `boot\system.img` capsule
|
space: a prebuilt `boot\system.img` capsule
|
||||||
([system-image.md](system-image.md)) when present, otherwise it walks
|
([system-image.md](os-development-guide/system-image.md)) when present, otherwise it walks
|
||||||
the volume's `/system` and optional `/test` trees (init included) into the
|
the volume's `/system` and optional `/test` trees (init included) into the
|
||||||
initial ramdisk. The kernel can boot "kernel-only" without either.
|
initial ramdisk. The kernel can boot "kernel-only" without either.
|
||||||
(`efi.zig:16`, `efi.zig:68`)
|
(`efi.zig:16`, `efi.zig:68`)
|
||||||
|
|||||||
+2
-2
@@ -27,7 +27,7 @@ boot log, memory summary, exception reports — appears on serial as plain text.
|
|||||||
|
|
||||||
QEMU captures that with `-serial file:serial.log`, giving a machine-readable
|
QEMU captures that with `-serial file:serial.log`, giving a machine-readable
|
||||||
transcript. Serial is per-architecture (x86 uses port I/O; an ARM board uses a
|
transcript. Serial is per-architecture (x86 uses port I/O; an ARM board uses a
|
||||||
memory-mapped UART), so it lives behind the [architecture](architecture.md) boundary — and adding
|
memory-mapped UART), so it lives behind the [architecture](os-development-guide/architecture.md) boundary — and adding
|
||||||
a new architecture's UART is what makes the same tests run there.
|
a new architecture's UART is what makes the same tests run there.
|
||||||
|
|
||||||
The serial log sink is **compiled in only under `-Dserial`** (off by default).
|
The serial log sink is **compiled in only under `-Dserial`** (off by default).
|
||||||
@@ -88,7 +88,7 @@ table in `test/qemu_test.py`):
|
|||||||
| `fault-recovery` | a ring-3 process that faults is killed and reaped while init keeps heartbeating — the OS survives | `DANOS-TEST-RESULT: PASS` |
|
| `fault-recovery` | a ring-3 process that faults is killed and reaped while init keeps heartbeating — the OS survives | `DANOS-TEST-RESULT: PASS` |
|
||||||
|
|
||||||
The faulting cases don't print a result line — they deliberately raise a CPU
|
The faulting cases don't print a result line — they deliberately raise a CPU
|
||||||
exception, and the harness asserts on the [exception report](interrupts.md) the
|
exception, and the harness asserts on the [exception report](os-development-guide/interrupts.md) the
|
||||||
handler prints (which also reaches serial). This reuses the real fault path as the
|
handler prints (which also reaches serial). This reuses the real fault path as the
|
||||||
test oracle: if the IDT/TSS weren't wired up, `fault-df` would triple-fault and the
|
test oracle: if the IDT/TSS weren't wired up, `fault-df` would triple-fault and the
|
||||||
marker would never appear.
|
marker would never appear.
|
||||||
|
|||||||
-118
@@ -1,118 +0,0 @@
|
|||||||
# Vision: a microkernel, built to learn
|
|
||||||
|
|
||||||
danos exists first and foremost as a **learning-by-doing project**: the point is to
|
|
||||||
build a real operating system, bump into the hard constraints for real, and research
|
|
||||||
them from a position of having actually hit them. The docs in this folder are part of
|
|
||||||
that — they're where a constraint gets understood once it's been met.
|
|
||||||
|
|
||||||
That framing sets the priorities. danos is not chasing a spec or a product; it's
|
|
||||||
chasing understanding, with a concrete, motivating **win condition** to aim at.
|
|
||||||
|
|
||||||
## The win condition
|
|
||||||
|
|
||||||
danos is a "win" when it:
|
|
||||||
|
|
||||||
- **boots and runs on real hardware** — the author's **PC** (x86-64) and **both
|
|
||||||
Raspberry Pis**: the **Zero 2 W** and the **Pi 5** (both `aarch64`, one backend —
|
|
||||||
see [arm.md](arm.md)),
|
|
||||||
- **has a graphical user interface**, ideally — building on the framebuffer it
|
|
||||||
already draws to.
|
|
||||||
|
|
||||||
Everything below serves that, or serves the curiosity that the project runs on.
|
|
||||||
|
|
||||||
## Why a microkernel: resilience
|
|
||||||
|
|
||||||
The kernel stays **minimal** — only what genuinely must run privileged:
|
|
||||||
|
|
||||||
- scheduling,
|
|
||||||
- inter-process communication (IPC),
|
|
||||||
- memory management (address spaces, page tables),
|
|
||||||
- low-level interrupt dispatch.
|
|
||||||
|
|
||||||
Everything else — device drivers, filesystems, the GUI, the network stack — runs as
|
|
||||||
an **isolated user-space server**, each in its own address space with only the
|
|
||||||
privileges it needs.
|
|
||||||
|
|
||||||
The reason for this shape is **resilience**: the ability to **re-initialise parts of
|
|
||||||
the OS while it runs**. A driver bug can't corrupt the kernel or another driver; a
|
|
||||||
crashed or wedged component is contained, killed, and **restarted** — "if I break
|
|
||||||
something, I can just fix it," without rebooting. Keeping the kernel tiny is part of
|
|
||||||
that strategy: the one thing that *can't* be restarted is the trusted base, so the
|
|
||||||
less code in it, the less that can take the whole system down. This is the project's
|
|
||||||
real motivation, and it has its own design note: [resilience.md](resilience.md).
|
|
||||||
|
|
||||||
The cost is that **IPC becomes the backbone**: what used to be a function call inside
|
|
||||||
a monolithic kernel is now a message between address spaces. In a microkernel, IPC
|
|
||||||
performance essentially *is* system performance (the lesson of L4), so it's a
|
|
||||||
first-class concern. Hardware interrupts become IPC too: the kernel turns an IRQ into
|
|
||||||
a message to the driver that owns the device.
|
|
||||||
|
|
||||||
## On real-time: an option, not a commitment
|
|
||||||
|
|
||||||
danos was originally framed as a hard **real-time** OS. That's now held as **one
|
|
||||||
interesting constraint to explore, not a requirement** — because real-time is a
|
|
||||||
*pervasive* invariant (every operation must be provably time-bounded, everywhere)
|
|
||||||
that would slow every milestone, whereas resilience is a set of *structural* features
|
|
||||||
that's lighter to build and is what the project actually wants. The trade-off is
|
|
||||||
written up in [smp.md](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience).
|
|
||||||
|
|
||||||
What danos keeps from the real-time direction, because it's cheap and useful anyway:
|
|
||||||
|
|
||||||
- **Fixed-priority preemptive scheduling** — the highest-priority ready task runs, and
|
|
||||||
preemption lets a runaway component be interrupted and killed (which *serves
|
|
||||||
resilience*). Already built ([scheduling.md](scheduling.md)).
|
|
||||||
- **A calibrated, deterministic clock** — already built ([device-interrupts.md](device-interrupts.md)).
|
|
||||||
|
|
||||||
What danos does *not* owe anyone unless it deliberately chooses real-time later:
|
|
||||||
timing *guarantees*, priority inheritance, bounded allocators, tickless timers, MCS
|
|
||||||
scheduling contexts. Concretely, the current [heap](heap.md) is a first-fit free list
|
|
||||||
with unbounded allocation time — fine here, and only a problem *if* a hard-real-time
|
|
||||||
path is ever added. Note that **QNX is both** a real-time and a restartable
|
|
||||||
microkernel, so choosing resilience now doesn't close the real-time door — it just
|
|
||||||
doesn't pay the tax yet.
|
|
||||||
|
|
||||||
## The roadmap — tracks, not a strict line
|
|
||||||
|
|
||||||
Because the driver is curiosity plus the win condition, the roadmap is a set of
|
|
||||||
**tracks** with dependencies, not a rigid sequence. Pick by interest; mind the
|
|
||||||
prerequisites.
|
|
||||||
|
|
||||||
**Done:** UEFI boot, framebuffer + [serial](testing.md), [physical frames](frame-allocator.md)
|
|
||||||
(with boot-services memory reclaimed), [paging](paging.md) with W^X, [exceptions and
|
|
||||||
interrupts](interrupts.md), a [calibrated timer + ns clock](device-interrupts.md), a
|
|
||||||
[heap](heap.md), a [fixed-priority preemptive scheduler](scheduling.md) with blocking,
|
|
||||||
in-kernel [IPC channels](ipc.md), SMP (all cores scheduling, with affinity), a
|
|
||||||
**higher-half kernel** with a physmap, and **user space**: per-process address
|
|
||||||
spaces, `syscall`/`sysret` with the `swapgs` discipline, a user-ELF loader, and
|
|
||||||
`/system/services/init` — a real user ELF built from `system/services/init/`, running at CPL 3 as PID 1 on its
|
|
||||||
own page tables — plus a [test harness](testing.md).
|
|
||||||
|
|
||||||
- **Isolation track** — **user mode + address-space isolation**. *Done: a
|
|
||||||
higher-half kernel with a physmap (the low half is user space), per-process
|
|
||||||
address spaces with CR3 switched on context switch, the `swapgs` discipline,
|
|
||||||
`syscall`/`sysret`, a user-ELF loader, an address-space/stack reaper for exited
|
|
||||||
tasks, and `/system/services/init` running as a real preemptive ring-3 process
|
|
||||||
(PID 1). Remaining polish: SMAP + fault-recovering copy-in/out, and TLB shootdown
|
|
||||||
once a process has more than one thread. (The real IPC syscalls —
|
|
||||||
`ipc_call`/`ipc_reply_wait` — have since been built and are the backbone every
|
|
||||||
driver and service speaks; see [ipc.md](ipc.md).)*
|
|
||||||
- **Resilience track** — fault → kill → notify, a supervisor/reincarnation server,
|
|
||||||
resource cleanup on death, then a restartable driver as proof. Needs isolation.
|
|
||||||
See [resilience.md](resilience.md).
|
|
||||||
- **ARM track** — the `aarch64` port so danos runs on the Zero 2 W and Pi 5. Largely
|
|
||||||
independent of the others (it's the [architecture layer](architecture.md)); directly serves the win
|
|
||||||
condition. Likely via aarch64-UEFI first (QEMU `virt` + AAVMF), then real boards.
|
|
||||||
See [arm.md](arm.md), and [discovery.md](discovery.md) for the device tree it needs.
|
|
||||||
- **GUI track** — a framebuffer-based windowing/compositor, and the input + display
|
|
||||||
drivers under it. Builds on the neutral framebuffer (so it's arch-independent), and
|
|
||||||
on the driver model from the isolation/resilience tracks. The visible payoff.
|
|
||||||
|
|
||||||
The natural spine is **isolation → (resilience + drivers) → GUI**, with the **ARM
|
|
||||||
track** pursued alongside whenever the itch to see it boot on a Pi wins out.
|
|
||||||
|
|
||||||
## How to use this page
|
|
||||||
|
|
||||||
Read it before adding anything structural. When a design decision comes up, the
|
|
||||||
question is: does it serve the **win condition** (runs on the three machines, with a
|
|
||||||
GUI), or the **learning** (a constraint worth meeting)? If it serves neither — e.g.
|
|
||||||
paying the full real-time tax with no payoff in sight — it can wait.
|
|
||||||
@@ -108,7 +108,7 @@ localised (below).
|
|||||||
## The architecture decision: `runtime.os` + `runtime.fs`, and retire `posix`
|
## The architecture decision: `runtime.os` + `runtime.fs`, and retire `posix`
|
||||||
|
|
||||||
danos already has the right split ([the private-ABI boundary](../README.md)): the
|
danos already has the right split ([the private-ABI boundary](../README.md)): the
|
||||||
kernel exposes a minimal syscall ABI ([syscall.md](syscall.md)); the **`runtime`**
|
kernel exposes a minimal syscall ABI ([syscall.md](os-development-guide/syscall.md)); the **`runtime`**
|
||||||
library is the stable, danos-native application ABI. What this roadmap adds:
|
library is the stable, danos-native application ABI. What this roadmap adds:
|
||||||
|
|
||||||
- **`runtime.os` — the seam.** A C-ABI-shaped module of the ~30 operations
|
- **`runtime.os` — the seam.** A C-ABI-shaped module of the ~30 operations
|
||||||
@@ -165,7 +165,7 @@ What the seam needs, and what danos already provides:
|
|||||||
| mmap / munmap | native syscalls ([abi.zig](../system/abi.zig)) | none |
|
| mmap / munmap | native syscalls ([abi.zig](../system/abi.zig)) | none |
|
||||||
| page allocator | over `mmap`, via `root.os.heap.page_allocator` override | ~30-line hook |
|
| page allocator | over `mmap`, via `root.os.heap.page_allocator` override | ~30-line hook |
|
||||||
| monotonic clock | `clock` syscall | none |
|
| monotonic clock | `clock` syscall | none |
|
||||||
| args / argv | SysV entry stack ([sysv.md](sysv.md)), `runtime.process.Init` | none |
|
| args / argv | SysV entry stack ([sysv.md](os-development-guide/sysv.md)), `runtime.process.Init` | none |
|
||||||
| stdout / stderr | `debug_write` today | wire fd 1/2 to a console **byte** stream |
|
| stdout / stderr | `debug_write` today | wire fd 1/2 to a console **byte** stream |
|
||||||
| mkdir / unlink / rename / truncate | done — engine + VFS + `runtime.fs` (Phase 2) | — |
|
| mkdir / unlink / rename / truncate | done — engine + VFS + `runtime.fs` (Phase 2) | — |
|
||||||
| stat fields | `{size, kind, mtime}` | **mode / inode** still missing (cache validity) |
|
| stat fields | `{size, kind, mtime}` | **mode / inode** still missing (cache validity) |
|
||||||
@@ -209,7 +209,7 @@ build); point danos's `build.zig`/CI at the resulting binary. Four localised pat
|
|||||||
plan9/serenity;
|
plan9/serenity;
|
||||||
- add `danos` to the freestanding/other **no-op `_start` list** in `std`'s `start.zig`,
|
- add `danos` to the freestanding/other **no-op `_start` list** in `std`'s `start.zig`,
|
||||||
so std does *not* emit its own System-V `_start` — danos keeps owning the entry shim
|
so std does *not* emit its own System-V `_start` — danos keeps owning the entry shim
|
||||||
and `Init`/argv construction it already builds ([sysv.md](sysv.md));
|
and `Init`/argv construction it already builds ([sysv.md](os-development-guide/sysv.md));
|
||||||
- wire the `system` selector `.danos => std.os.danos` in `std.posix`;
|
- wire the `system` selector `.danos => std.os.danos` in `std.posix`;
|
||||||
- add `std/os/danos.zig` — **the seam itself**, promoted near-verbatim from the
|
- add `std/os/danos.zig` — **the seam itself**, promoted near-verbatim from the
|
||||||
`runtime.os` developed first in Phase 1 (against the stock toolchain, so the fork is
|
`runtime.os` developed first in Phase 1 (against the stock toolchain, so the fork is
|
||||||
@@ -344,10 +344,10 @@ Two current decisions fall out of this roadmap:
|
|||||||
## Related
|
## Related
|
||||||
|
|
||||||
- [vision.md](vision.md) — the north star this serves.
|
- [vision.md](vision.md) — the north star this serves.
|
||||||
- [syscall.md](syscall.md) — the kernel↔runtime ABI `runtime.os` is built on.
|
- [syscall.md](os-development-guide/syscall.md) — the kernel↔runtime ABI `runtime.os` is built on.
|
||||||
- [sysv.md](sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs.
|
- [sysv.md](os-development-guide/sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs.
|
||||||
- [ipc.md](ipc.md) — the IPC the VFS/FAT operations travel over.
|
- [ipc.md](device-driver-development-guide/ipc.md) — the IPC the VFS/FAT operations travel over.
|
||||||
- [danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md) — the
|
- [danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md) — the
|
||||||
filesystem layout the file surface serves.
|
filesystem layout the file surface serves.
|
||||||
- [coding-standards.md](coding-standards.md) — danos naming (why the compat spellings
|
- [coding-standards.md](coding-standards.md) — danos naming (why the compat spellings
|
||||||
are confined, and now retired).
|
are confined, and now retired).
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# xkeyboard-config — X11 keyboard layouts, compiled to Zig
|
# xkeyboard-config — X11 keyboard layouts, compiled to Zig
|
||||||
|
|
||||||
This module turns a physical key (a **USB HID usage**, as the [input module](../../docs/input.md)
|
This module turns a physical key (a **USB HID usage**, as the [input module](../../docs/device-driver-development-guide/input.md)
|
||||||
delivers in `KeyEvent.keycode`) plus a modifier state into a **keysym** and, when the key
|
delivers in `KeyEvent.keycode`) plus a modifier state into a **keysym** and, when the key
|
||||||
produces one, a **character** (a Unicode scalar). It is what lets a `keycode` become a
|
produces one, a **character** (a Unicode scalar). It is what lets a `keycode` become a
|
||||||
`character` — a keymap — without danos shipping an X11 runtime.
|
`character` — a keymap — without danos shipping an X11 runtime.
|
||||||
|
|||||||
Reference in New Issue
Block a user