From 52d6e372fd0b715428a57da7e197c5ef117d3b2d Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Thu, 23 Jul 2026 00:24:01 +0100 Subject: [PATCH] re-org docs --- README.md | 2 +- docs/README.md | 132 +++++++++--------- .../amd-gpus.md | 0 .../device-interrupts.md | 20 +-- .../device-manager.md | 18 +-- .../display-plan.md | 28 ++-- .../display-v2-plan.md | 6 +- .../display-v2.md | 2 +- .../display.md | 66 ++++----- .../driver-model.md | 12 +- .../drivers.md | 12 +- .../input.md | 26 ++-- .../intel-igpu.md | 0 .../ipc.md | 10 +- .../nvidia-gpus.md | 0 .../usb-hub.md | 0 .../danos-file-system-hierarchy-FSH.md | 6 +- .../vfs-protocol.md | 2 +- .../README.MD | 0 docs/{ => os-development-guide}/acpi.md | 0 .../architecture.md | 4 +- docs/{ => os-development-guide}/arm.md | 4 +- docs/{ => os-development-guide}/discovery.md | 18 +-- docs/{ => os-development-guide}/efi.md | 2 +- .../frame-allocator.md | 0 .../{ => os-development-guide}/framebuffer.md | 0 docs/{ => os-development-guide}/gop.md | 0 docs/{ => os-development-guide}/halting.md | 0 docs/{ => os-development-guide}/heap.md | 2 +- docs/{ => os-development-guide}/interrupts.md | 4 +- docs/{ => os-development-guide}/logging.md | 0 docs/{ => os-development-guide}/memory-map.md | 0 docs/{ => os-development-guide}/paging.md | 6 +- docs/{ => os-development-guide}/power.md | 6 +- .../process-lifecycle.md | 8 +- .../process-management.md | 0 .../{ => os-development-guide}/release-iso.md | 4 +- docs/{ => os-development-guide}/resilience.md | 16 +-- docs/{ => os-development-guide}/scheduling.md | 12 +- .../shared-fate-plan.md | 2 +- docs/{ => os-development-guide}/smp.md | 8 +- docs/{ => os-development-guide}/syscall.md | 4 +- .../system-image.md | 2 +- docs/{ => os-development-guide}/sysv.md | 0 .../threading-plan.md | 26 ++-- docs/{ => os-development-guide}/threading.md | 40 +++--- docs/{ => os-development-guide}/timers.md | 8 +- docs/{ => os-development-guide}/vdso.md | 2 +- docs/system-requirements.md | 2 +- docs/testing.md | 4 +- docs/vision.md | 118 ---------------- docs/zig-self-hosting.md | 14 +- library/xkeyboard-config/README.md | 2 +- 53 files changed, 271 insertions(+), 389 deletions(-) rename docs/{ => device-driver-development-guide}/amd-gpus.md (100%) rename docs/{ => device-driver-development-guide}/device-interrupts.md (92%) rename docs/{ => device-driver-development-guide}/device-manager.md (91%) rename docs/{ => device-driver-development-guide}/display-plan.md (89%) rename docs/{ => device-driver-development-guide}/display-v2-plan.md (97%) rename docs/{ => device-driver-development-guide}/display-v2.md (98%) rename docs/{ => device-driver-development-guide}/display.md (87%) rename docs/{ => device-driver-development-guide}/driver-model.md (97%) rename docs/{ => device-driver-development-guide}/drivers.md (97%) rename docs/{ => device-driver-development-guide}/input.md (88%) rename docs/{ => device-driver-development-guide}/intel-igpu.md (100%) rename docs/{ => device-driver-development-guide}/ipc.md (93%) rename docs/{ => device-driver-development-guide}/nvidia-gpus.md (100%) rename docs/{ => device-driver-development-guide}/usb-hub.md (100%) rename docs/{ => file-system-development}/danos-file-system-hierarchy-FSH.md (97%) rename docs/{ => file-system-development}/vfs-protocol.md (99%) rename docs/{os-developer-guide => os-development-guide}/README.MD (100%) rename docs/{ => os-development-guide}/acpi.md (100%) rename docs/{ => os-development-guide}/architecture.md (97%) rename docs/{ => os-development-guide}/arm.md (97%) rename docs/{ => os-development-guide}/discovery.md (93%) rename docs/{ => os-development-guide}/efi.md (99%) rename docs/{ => os-development-guide}/frame-allocator.md (100%) rename docs/{ => os-development-guide}/framebuffer.md (100%) rename docs/{ => os-development-guide}/gop.md (100%) rename docs/{ => os-development-guide}/halting.md (100%) rename docs/{ => os-development-guide}/heap.md (97%) rename docs/{ => os-development-guide}/interrupts.md (97%) rename docs/{ => os-development-guide}/logging.md (100%) rename docs/{ => os-development-guide}/memory-map.md (100%) rename docs/{ => os-development-guide}/paging.md (96%) rename docs/{ => os-development-guide}/power.md (96%) rename docs/{ => os-development-guide}/process-lifecycle.md (98%) rename docs/{ => os-development-guide}/process-management.md (100%) rename docs/{ => os-development-guide}/release-iso.md (95%) rename docs/{ => os-development-guide}/resilience.md (93%) rename docs/{ => os-development-guide}/scheduling.md (93%) rename docs/{ => os-development-guide}/shared-fate-plan.md (99%) rename docs/{ => os-development-guide}/smp.md (98%) rename docs/{ => os-development-guide}/syscall.md (97%) rename docs/{ => os-development-guide}/system-image.md (98%) rename docs/{ => os-development-guide}/sysv.md (100%) rename docs/{ => os-development-guide}/threading-plan.md (96%) rename docs/{ => os-development-guide}/threading.md (91%) rename docs/{ => os-development-guide}/timers.md (93%) rename docs/{ => os-development-guide}/vdso.md (99%) delete mode 100644 docs/vision.md diff --git a/README.md b/README.md index e385e0a..d4d261e 100644 --- a/README.md +++ b/README.md @@ -63,7 +63,7 @@ zig build release-x86-64 Produces `zig-out/danos-x86-64.iso`, a hybrid ISO that boots flashed raw to a USB stick (balenaEtcher, dd) or burned to optical media — see -[docs/release-iso.md](docs/release-iso.md). `zig build check-iso-image` +[docs/release-iso.md](docs/os-development-guide/release-iso.md). `zig build check-iso-image` validates it without booting. ## Run diff --git a/docs/README.md b/docs/README.md index 347187f..45970e7 100644 --- a/docs/README.md +++ b/docs/README.md @@ -3,102 +3,102 @@ Notes on how danos boots and draws, written to explain the *why* behind the code rather than restate it. Roughly in the order things happen at runtime: -1. **[efi.md](efi.md) — EFI / the boot process.** How UEFI firmware finds and +1. **[efi.md](os-development-guide/efi.md) — EFI / the boot process.** How UEFI firmware finds and runs the bootloader, what the loader gathers before `ExitBootServices`, how it loads the kernel ELF, and the ABI contract for the jump into the kernel. Start here. -2. **[system-image.md](system-image.md) — system.img, the boot capsule.** The +2. **[system-image.md](os-development-guide/system-image.md) — system.img, the boot capsule.** The bundled user binaries packed into one file in the initial-ramdisk wire format, because one open + one sequential read is the only file I/O shape firmware is fast at. The trivial container format, the three artifacts one build list derives (tree, manifest, capsule), the loader's three-strategy fallback chain, and the capsule's kernel-side life as both the spawn table and the read-only `/system` mount. -3. **[gop.md](gop.md) — the Graphics Output Protocol.** How UEFI exposes graphics +3. **[gop.md](os-development-guide/gop.md) — the Graphics Output Protocol.** How UEFI exposes graphics modes (unlike fixed VGA modes), how we detect the monitor's native resolution from EDID and switch to it, and the pixel formats we accept or reject. -4. **[framebuffer.md](framebuffer.md) — the framebuffer.** What the linear +4. **[framebuffer.md](os-development-guide/framebuffer.md) — the framebuffer.** What the linear framebuffer the loader hands over actually is, and what **pitch** (stride) means versus width — the detail you have to get right to avoid a skewed image. -5. **[memory-map.md](memory-map.md) — the memory map.** How the loader learns what +5. **[memory-map.md](os-development-guide/memory-map.md) — the memory map.** How the loader learns what physical RAM exists and hands it to the kernel in danos's own neutral format, rather than leaking UEFI's memory descriptors across the boundary. -6. **[frame-allocator.md](frame-allocator.md) — the physical frame allocator.** The +6. **[frame-allocator.md](os-development-guide/frame-allocator.md) — the physical frame allocator.** The bitmap allocator that hands out and reclaims 4 KiB physical frames from that map — the primitive page tables and the heap are built on. -7. **[interrupts.md](interrupts.md) — interrupts and exceptions.** The GDT, IDT and +7. **[interrupts.md](os-development-guide/interrupts.md) — interrupts and exceptions.** The GDT, IDT and TSS, the exception stubs, and the handler that reports a CPU fault in red instead of letting it triple-fault into a silent reset. -8. **[paging.md](paging.md) — the kernel's page tables.** Building our own 4-level +8. **[paging.md](os-development-guide/paging.md) — the kernel's page tables.** Building our own 4-level page tables, identity-mapping the low 4 GiB, and switching CR3 off the firmware's tables onto ours. -9. **[device-interrupts.md](device-interrupts.md) — device interrupts.** The Local +9. **[device-interrupts.md](device-driver-development-guide/device-interrupts.md) — device interrupts.** The Local APIC and its timer — the kernel's first interrupt that is *handled and returned from*, giving it a heartbeat. -10. **[heap.md](heap.md) — the kernel heap.** A growable free-list allocator built on +10. **[heap.md](os-development-guide/heap.md) — the kernel heap.** A growable free-list allocator built on the VMM, exposed as a `std.mem.Allocator` so std containers work — dynamic allocation for the kernel. -11. **[scheduling.md](scheduling.md) — the scheduler.** Fixed-priority preemptive +11. **[scheduling.md](os-development-guide/scheduling.md) — the scheduler.** Fixed-priority preemptive multitasking: kernel threads, the context switch, O(1) priority selection, and blocking (sleep, wait queues) — the leap to a running system. -12. **[ipc.md](ipc.md) — inter-process communication.** Bounded blocking +12. **[ipc.md](device-driver-development-guide/ipc.md) — inter-process communication.** Bounded blocking message-passing channels, then synchronous call/reply between *processes* over endpoints — the backbone the microkernel's isolated servers talk over. -13. **[syscall.md](syscall.md) — system calls.** How ring 3 asks the kernel for +13. **[syscall.md](os-development-guide/syscall.md) — system calls.** How ring 3 asks the kernel for something: the `syscall`/`sysret` fast path, the trap frame, and why the table is - deliberately tiny. The numbers are a **private** ABI — [vdso.md](vdso.md) designs + deliberately tiny. The numbers are a **private** ABI — [vdso.md](os-development-guide/vdso.md) designs the public boundary that will hide them. -14. **[vfs-protocol.md](vfs-protocol.md) — the VFS wire protocol.** The language-neutral +14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral byte-level spec of the file protocol spoken over IPC: request/reply headers, the operation table, mount routing, and the append-only evolution rules — the first IPC protocol documented as public ABI. -15. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an +15. **[drivers.md](device-driver-development-guide/drivers.md) — writing a driver.** The payoff: a driver is an ordinary ring-3 process that claims a device, maps its registers, and **sleeps until its hardware interrupts it**. The claim is the capability; `irq_ack` is the unmask. -16. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How +16. **[driver-model.md](device-driver-development-guide/driver-model.md) — buses, classes and host controllers.** How real driver stacks factor into three shapes and how families share code. The three primitives it proposed are long since built (M13 capability passing, M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them — hello, supervision, restart — is built too (device-manager.md, M18). -17. **[usb-hub.md](usb-hub.md) — USB hubs.** Built (M22): why hub topology is handled +17. **[usb-hub.md](device-driver-development-guide/usb-hub.md) — USB hubs.** Built (M22): why hub topology is handled *inside* the `usb-xhci-bus` driver rather than a separate hub class driver — a device behind a hub is reached by the **controller**, programmed with a route string in its slot context — plus the compound-hub reality (a USB 3.0 hub is physically two hubs) and detection via the hub's status-change interrupt endpoint. -18. **[process-management.md](process-management.md) — process management.** The +18. **[process-management.md](os-development-guide/process-management.md) — process management.** The microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the supervision link as the kill authority, and child-exit notifications over the same endpoints IRQs arrive on. -19. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built +19. **[process-lifecycle.md](os-development-guide/process-lifecycle.md) — the process lifecycle.** Built (M17): signals over IPC as the one lifecycle vocabulary every process speaks — the POSIX.1-1990 words with message delivery instead of stack hijack, the stable `process` module interface, exit reasons, published exit events any stateful service can subscribe to (the VFS releasing dead clients' handles), and the two iron rules (cleanup is the kernel's job; kill is not a signal). -20. **[device-manager.md](device-manager.md) — the device manager.** Built (M18, +20. **[device-manager.md](device-driver-development-guide/device-manager.md) — the device manager.** Built (M18, through the app surface): the tree, the matcher, and the supervisor. Tree structure lives in the manager, authority stays in the kernel; bus drivers report what they see; drivers are restarted through the lifecycle vocabulary — the plan that turns - [resilience.md](resilience.md)'s restart goal into increments. -21. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard, + [resilience.md](os-development-guide/resilience.md)'s restart goal into increments. +21. **[input.md](device-driver-development-guide/input.md) — the input module.** Broadcasting input events (keyboard, mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish service layered on top. -22. **[display.md](display.md) — the display service.** The display half of the GUI +22. **[display.md](device-driver-development-guide/display.md) — the display service.** The display half of the GUI track: a user-space compositor that owns the framebuffer, composes a layer stack into a double buffer, and presents it. Why GOP and the PCI display device are two views of one controller, the device-node + write-combining handoff, and what flicker-free buys - that tear-free doesn't. Plan: [display-plan.md](display-plan.md). **v2** (complete) makes + that tear-free doesn't. Plan: [display-plan.md](device-driver-development-guide/display-plan.md). **v2** (complete) makes scanout a pluggable backend — GOP floor + a native virtio-gpu driver, hot-attached, with runtime mode-set, EDID, fenced vsync presents, and restart re-attach: - [display-v2.md](display-v2.md), plan [display-v2-plan.md](display-v2-plan.md). Looking + [display-v2.md](device-driver-development-guide/display-v2.md), plan [display-v2-plan.md](device-driver-development-guide/display-v2-plan.md). Looking further out, three research snapshots survey what a *native* driver for real GPU silicon - would take as another `.scanout` backend: [nvidia-gpus.md](nvidia-gpus.md) (RTX 3060 / - Ampere), [amd-gpus.md](amd-gpus.md) (RX 6600 / RDNA2), and [intel-igpu.md](intel-igpu.md) + would take as another `.scanout` backend: [nvidia-gpus.md](device-driver-development-guide/nvidia-gpus.md) (RTX 3060 / + Ampere), [amd-gpus.md](device-driver-development-guide/amd-gpus.md) (RX 6600 / RDNA2), and [intel-igpu.md](device-driver-development-guide/intel-igpu.md) (Intel iGPU). -23. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and +23. **[halting.md](os-development-guide/halting.md) — halting.** Why a kernel can't just "exit", and how `while (true) hlt` parks the CPU safely once there's nothing left to do. Start with the north star: @@ -108,7 +108,7 @@ Start with the north star: **resilience** (restartable components). Win condition: runs on the author's PC and both Raspberry Pis, ideally with a GUI. Real-time is an option to explore, not a requirement. The *why* that shapes everything below. -- **[resilience.md](resilience.md) — resilience.** A design note (not built yet) on +- **[resilience.md](os-development-guide/resilience.md) — resilience.** A design note (not built yet) on fault isolation + live restart — the reincarnation-server + capability model that makes "if I break it, I can restart it" real. danos's core motivation. - **[zig-self-hosting.md](zig-self-hosting.md) — running Zig on danos.** A design note @@ -117,19 +117,19 @@ Start with the north star: port to **one seam** (`std.os.danos`), so we build an `os` seam module (→ that seam) plus the thin `file-system` module, retire the `posix` shim, and follow a phased path to `zig build-exe hello.zig` running on danos — **not** Linux-ABI emulation. -- **[threading.md](threading.md) — threads, the std-shaped way.** **Built** (M1–M6): +- **[threading.md](os-development-guide/threading.md) — threads, the std-shaped way.** **Built** (M1–M6): the `thread` module's `Thread` mirrors `std.Thread`'s API (spawn/join/detach, Mutex/Condition/ Semaphore) over a **private** thread ABI — several tasks sharing one address space via a `thread_spawn` syscall, futex-backed blocking, address-space refcounting. Why it's the - native type and not literal `std.Thread` (the [private ABI](syscall.md)), and why - threads stay a narrow opt-in against the [resilience](resilience.md) default. Build - plan + gates: [threading-plan.md](threading-plan.md). -- **[vdso.md](vdso.md) — the vDSO, the public system-call boundary.** A design note + native type and not literal `std.Thread` (the [private ABI](os-development-guide/syscall.md)), and why + threads stay a narrow opt-in against the [resilience](os-development-guide/resilience.md) default. Build + plan + gates: [threading-plan.md](os-development-guide/threading-plan.md). +- **[vdso.md](os-development-guide/vdso.md) — the vDSO, the public system-call boundary.** A design note (not built yet) on keeping `abi.zig` genuinely private: a kernel-supplied, C-ABI entry blob mapped into every process as the *only* way into the kernel — so the syscall numbers can be renumbered or randomised at will, and Rust/C binaries get a stable boundary without danos growing a dynamic linker. danos's public ABI = the - vDSO + the documented IPC wire protocols ([vfs-protocol.md](vfs-protocol.md) first). + vDSO + the documented IPC wire protocols ([vfs-protocol.md](file-system-development/vfs-protocol.md) first). Cutting across all of these: @@ -137,69 +137,69 @@ Cutting across all of these: hardware needed to run danos: minimum specs (UEFI x86-64, ACPI, PCIe ECAM, xHCI, ~128 MiB RAM) grounded in what the boot path actually assumes, plus a plain-language guide matching Intel/AMD CPU generations by name. -- **[release-iso.md](release-iso.md) — the release ISO.** The flashable boot +- **[release-iso.md](os-development-guide/release-iso.md) — the release ISO.** The flashable boot media: `zig build release-x86-64` wraps the FAT32 boot volume in a hybrid ISO (MBR ESP partition + El Torito EFI entry, one embedded image) that Etcher/dd flash to USB or a burner writes to disc — built by an in-repo pure-Python tool, like the FAT image itself. -- **[architecture.md](architecture.md) — the architecture split.** How CPU-specific code is kept +- **[architecture.md](os-development-guide/architecture.md) — the architecture split.** How CPU-specific code is kept behind a build-time `arch` module so the generic kernel never names x86_64, leaving room for other systems (e.g. an AArch64 Raspberry Pi) later. -- **[arm.md](arm.md) — ARM targets.** The Raspberry Pi landscape the arch split is +- **[arm.md](os-development-guide/arm.md) — ARM targets.** The Raspberry Pi landscape the arch split is aiming at: `arm` (32-bit, Pi Zero W) vs `aarch64` (64-bit, Pi 3-5), UEFI vs device-tree boot, and what each layer needs. -- **[discovery.md](discovery.md) — device discovery.** A design note on learning what +- **[discovery.md](os-development-guide/discovery.md) — device discovery.** A design note on learning what hardware exists via ACPI (x86) or device tree (ARM) behind one neutral device model — when to build it, and how to keep it architecture-agnostic. -- **[acpi.md](acpi.md) — finding the ACPI tables.** The concrete x86 locator chain: +- **[acpi.md](os-development-guide/acpi.md) — finding the ACPI tables.** The concrete x86 locator chain: how the loader captures the **RSDP**, hands its physical address across in `BootInformation`, and how the platform derives the **RSDT/XSDT** from it and walks the SDTs — plus the live event side (the SCI, the power button, GPE/Notify) the ring-3 acpi service runs. -- **[power.md](power.md) — the power service.** System power as a domain-named +- **[power.md](os-development-guide/power.md) — the power service.** System power as a domain-named service: button/lid/battery events published to subscribers, and init's orderly - shutdown composing the [lifecycle](process-lifecycle.md) stop sequence with an ACPI + shutdown composing the [lifecycle](os-development-guide/process-lifecycle.md) stop sequence with an ACPI S5 write. Firmware-neutral — a PSCI backend drops in on ARM. -- **[timers.md](timers.md) — timers and time.** The ring-3 surface for reading the +- **[timers.md](os-development-guide/timers.md) — timers and time.** The ring-3 surface for reading the clock and waiting: why `now()` is a syscall rather than a service, and the one-shot timer notification (`timer_bind`) that gives supervisors a timed wait — built on the - LAPIC heartbeat and calibrated TSC of [device-interrupts.md](device-interrupts.md). -- **[smp.md](smp.md) — multiple cores.** A design/research note on how microkernels + LAPIC heartbeat and calibrated TSC of [device-interrupts.md](device-driver-development-guide/device-interrupts.md). +- **[smp.md](os-development-guide/smp.md) — multiple cores.** A design/research note on how microkernels (L4, seL4) handle SMP — big kernel lock vs per-CPU vs multikernel — and how the right choice depends on whether danos is chasing real-time or resilience. - **[coding-standards.md](coding-standards.md) — coding standards.** The naming rule the tree follows: non-acronyms are spelled out in full (`message`, not `msg`), files are `kebab-case`, code follows Zig's case conventions, and the handful of exceptions (POSIX/C ABI names, `init`/`len`/`ptr`, acronyms). -- **[sysv.md](sysv.md) — the calling convention.** What "the kernel is SysV" means, +- **[sysv.md](os-development-guide/sysv.md) — the calling convention.** What "the kernel is SysV" means, and why the loader→kernel boundary has to pin it (the RDI-vs-RCX handoff). - **[testing.md](testing.md) — testing.** How the kernel is tested by booting it in QEMU and asserting on its serial output — reproducibly, and structured so the same tests run across architectures. -- **[logging.md](logging.md) — logging.** The multi-sink diagnostic log (serial, +- **[logging.md](os-development-guide/logging.md) — logging.** The multi-sink diagnostic log (serial, 0xE9 debugcon, file later) kept separate from the framebuffer display, plus the robustness path: optional framebuffer, POST-code checkpoints, and a persistent panic breadcrumb so the kernel survives — and can be diagnosed — with no output. ## How the pieces relate -The boot flow ties them together: UEFI runs the loader ([efi.md](efi.md)), which -queries the **GOP** to pick a graphics mode ([gop.md](gop.md)), hands the kernel a -**framebuffer** to draw into ([framebuffer.md](framebuffer.md)) and a **memory -map** of physical RAM ([memory-map.md](memory-map.md)); the kernel turns that map -into a **frame allocator** ([frame-allocator.md](frame-allocator.md)), installs -its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.md)), -builds its own **page tables** and switches onto them ([paging.md](paging.md)), -brings up the **heap** for dynamic allocation ([heap.md](heap.md)), starts the -**scheduler** ([scheduling.md](scheduling.md)) and the **timer** that preempts it -([device-interrupts.md](device-interrupts.md)) — with tasks blocking, sleeping and -passing messages over **[IPC](ipc.md)** channels — runs, its CPU-specific bits -behind the [architecture](architecture.md) boundary, and when idle, or on a panic, it **halts** -([halting.md](halting.md)). +The boot flow ties them together: UEFI runs the loader ([efi.md](os-development-guide/efi.md)), which +queries the **GOP** to pick a graphics mode ([gop.md](os-development-guide/gop.md)), hands the kernel a +**framebuffer** to draw into ([framebuffer.md](os-development-guide/framebuffer.md)) and a **memory +map** of physical RAM ([memory-map.md](os-development-guide/memory-map.md)); the kernel turns that map +into a **frame allocator** ([frame-allocator.md](os-development-guide/frame-allocator.md)), installs +its **descriptor tables** so CPU faults are caught ([interrupts.md](os-development-guide/interrupts.md)), +builds its own **page tables** and switches onto them ([paging.md](os-development-guide/paging.md)), +brings up the **heap** for dynamic allocation ([heap.md](os-development-guide/heap.md)), starts the +**scheduler** ([scheduling.md](os-development-guide/scheduling.md)) and the **timer** that preempts it +([device-interrupts.md](device-driver-development-guide/device-interrupts.md)) — with tasks blocking, sleeping and +passing messages over **[IPC](device-driver-development-guide/ipc.md)** channels — runs, its CPU-specific bits +behind the [architecture](os-development-guide/architecture.md) boundary, and when idle, or on a panic, it **halts** +([halting.md](os-development-guide/halting.md)). -Above that line the microkernel proper begins: **discovery** ([discovery.md](discovery.md), -[acpi.md](acpi.md)) learns what hardware exists, ring-3 processes ask the kernel for -things through the small **[syscall](syscall.md)** table, isolated servers reach each -other over IPC **endpoints** ([ipc.md](ipc.md)), and a **[driver](drivers.md)** claims +Above that line the microkernel proper begins: **discovery** ([discovery.md](os-development-guide/discovery.md), +[acpi.md](os-development-guide/acpi.md)) learns what hardware exists, ring-3 processes ask the kernel for +things through the small **[syscall](os-development-guide/syscall.md)** table, isolated servers reach each +other over IPC **endpoints** ([ipc.md](device-driver-development-guide/ipc.md)), and a **[driver](device-driver-development-guide/drivers.md)** claims a device, maps its registers, and sleeps until the hardware interrupts it — which is the whole reason for the arrangement ([vision.md](vision.md)). @@ -208,7 +208,7 @@ the whole reason for the arrangement ([vision.md](vision.md)). danos is a **monorepo of sub-projects**. Each service or driver is a directory that is its own Zig module — it can hold as many files as it needs, and other sub-projects reach it *by module name*, never by a path into its files. The source tree deliberately -**mirrors the runtime FHS** ([danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md)): +**mirrors the runtime FHS** ([danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md)): what you see under `system/` in the source is what a running danos represents under `/system`. diff --git a/docs/amd-gpus.md b/docs/device-driver-development-guide/amd-gpus.md similarity index 100% rename from docs/amd-gpus.md rename to docs/device-driver-development-guide/amd-gpus.md diff --git a/docs/device-interrupts.md b/docs/device-driver-development-guide/device-interrupts.md similarity index 92% rename from docs/device-interrupts.md rename to docs/device-driver-development-guide/device-interrupts.md index 716eb81..790a7f5 100644 --- a/docs/device-interrupts.md +++ b/docs/device-driver-development-guide/device-interrupts.md @@ -1,6 +1,6 @@ # Device interrupts -CPU exceptions ([interrupts.md](interrupts.md)) are the kernel reacting to its own +CPU exceptions ([interrupts.md](../os-development-guide/interrupts.md)) are the kernel reacting to its own mistakes. **Device interrupts** are the opposite: hardware asking for attention — a timer firing, a key pressed, a packet arriving. They share the IDT, but differ in one fundamental way: an exception here is terminal (we report and halt), while a @@ -10,7 +10,7 @@ back — the same mechanism a scheduler will later use to preempt tasks. The first device we bring up is the **timer**, because it's the simplest: it lives entirely on the CPU's local interrupt controller, needing no external routing. -It's all x86_64-specific, behind the [architecture](architecture.md) boundary. +It's all x86_64-specific, behind the [architecture](../os-development-guide/architecture.md) boundary. ## The APIC, not the PIC @@ -40,7 +40,7 @@ count that becomes the reload value. From then on it fires vector 32 repeatedly, its own, forever. The reload count isn't picked arbitrarily — it's **calibrated to real time**, -which the [real-time](vision.md) scheduling guarantees depend on. Since the LAPIC +which the [real-time](../vision.md) scheduling guarantees depend on. Since the LAPIC timer's raw rate is bus-clock dependent and unknown up front, `calibrate` runs the LAPIC timer one-shot from its maximum count while a **reference clock** counts out a known 10 ms, then sees how far the LAPIC got — its counts-per-millisecond, from which @@ -53,7 +53,7 @@ a missing PIT would hang the boot): 1. **CPUID leaf 0x15** — the CPU's TSC frequency directly, needing no external timer at all (the LAPIC is then measured against the TSC). -2. The **HPET**, discovered via ACPI (see [discovery](discovery.md) / [acpi](acpi.md)). +2. The **HPET**, discovered via ACPI (see [discovery](../os-development-guide/discovery.md) / [acpi](../os-development-guide/acpi.md)). 3. The **ACPI PM timer** (a fixed 3.579545 MHz counter from the FADT). 4. The **PIT** (legacy 8254, 1.193182 MHz) — last resort, and bounded so it can't hang. @@ -100,7 +100,7 @@ values (a second socket, some firmware), so a thread migrating from a core readi check** as each application processor comes online (`checkWarpSource`, adapted from Linux's): the waking core and the BSP hammer a shared "highest seen" TSC under a lock, and if either ever reads below it, the cores' TSCs are skewed. It's pairwise because APs -come up one at a time ([smp.md](smp.md)). +come up one at a time ([smp.md](../os-development-guide/smp.md)). **The fallback.** When the TSC fails either test — non-invariant (a bare VM such as the default qemu64), or warped between cores — danos moves the monotonic clock onto the @@ -160,7 +160,7 @@ A device handler is a plain `fn () void` — a timer or keyboard handler doesn't the interrupted registers. (The stubs originally didn't save the SSE/vector registers, so a handler couldn't use them; `isr_common` now does an `fxsave`/`fxrstor` of the full SSE/x87 state around dispatch — see -[interrupts.md](interrupts.md).) +[interrupts.md](../os-development-guide/interrupts.md).) ## Turning them on @@ -168,11 +168,11 @@ Exceptions can't be masked, which is why they worked all along. Maskable device interrupts don't fire until the CPU's interrupt flag is set — so the final step is `sti` (`arch.enableInterrupts()`), after the APIC and timer are configured. From that instant the kernel has a heartbeat, and its idle `hlt` loop -([halting.md](halting.md)) wakes on every tick and dozes off again. +([halting.md](../os-development-guide/halting.md)) wakes on every tick and dozes off again. ## Verifying it -The `timer` test (see [testing.md](testing.md)) is the proof that an interrupt both +The `timer` test (see [testing.md](../testing.md)) is the proof that an interrupt both *fires* and *returns*: it records the tick count, busy-waits, and checks the count advanced on its own. @@ -188,13 +188,13 @@ spinning in unrelated code — is the whole mechanism working end to end. ## Since (done elsewhere) - **Preemption**: the timer handler is where the scheduler decides to switch — the - reason a *returning* interrupt matters. See [scheduling.md](scheduling.md). + reason a *returning* interrupt matters. See [scheduling.md](../os-development-guide/scheduling.md). - **`sleep()` / timeouts** built on the calibrated clock. - **The I/O APIC, routed**: external device lines now reach a vector, and the interrupt is delivered onward to a *user-space* driver as an IPC message. See [drivers.md](drivers.md). - **Uncacheable MMIO**: device grants are mapped `PCD|PWT` (strong-uncacheable) for - user drivers — see [paging.md](paging.md). + user drivers — see [paging.md](../os-development-guide/paging.md). ## What's next (partly done since) diff --git a/docs/device-manager.md b/docs/device-driver-development-guide/device-manager.md similarity index 91% rename from docs/device-manager.md rename to docs/device-driver-development-guide/device-manager.md index 26473e8..1c635f8 100644 --- a/docs/device-manager.md +++ b/docs/device-driver-development-guide/device-manager.md @@ -10,18 +10,18 @@ mirrors them and prunes a dead reporter's children, and the `usb-report` scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13): `enumerate` and `subscribe` over IPC, with `device-list` as the first client — the manager is now the one answer to "what devices exist" for applications. -The primitives underneath are real ([process-management.md](process-management.md): +The primitives underneath are real ([process-management.md](../os-development-guide/process-management.md): spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first per-device driver spawn works (the device manager matches the xHCI controller by PCI class and spawns `usb-xhci-bus` with the device id as argv[1]). This document designs the rest: the device manager as **the tree, the matcher, and the supervisor** — the -policy process that turns [resilience.md](resilience.md)'s restart goal into practice +policy process that turns [resilience.md](../os-development-guide/resilience.md)'s restart goal into practice for drivers. How processes stop, reload, and report their deaths is deliberately **not** in this document: that is the universal lifecycle every danos process speaks — -[process-lifecycle.md](process-lifecycle.md), signals over IPC and the stable +[process-lifecycle.md](../os-development-guide/process-lifecycle.md), signals over IPC and the stable `process` interface. The device manager is that design's first serious customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a driver is stopped, health-checked, and buried exactly like any other process. @@ -35,7 +35,7 @@ The device tree is two things fused: *information* (what exists, how it nests) a claims, resource containment on `device_register`, the `mmio_map`/`irq_bind`/`msi_bind` gates — and **cleans all of it up when a process dies** (settled; it is increment 1 of - [process-lifecycle.md](process-lifecycle.md)). The three invariants in + [process-lifecycle.md](../os-development-guide/process-lifecycle.md)). The three invariants in [driver-model.md](driver-model.md) stay exactly where they are. A device manager that could mint MMIO mappings by its own say-so would be a second kernel, and a buggy one would un-earn everything the microkernel bought. @@ -51,7 +51,7 @@ enumeration is a **pci-bus driver**: the manager spawns it against the host brid like any bus reports children. ACPI becomes an **acpi service** that interprets the tables and reports the namespace. The manager only orchestrates and merges. Moving AML interpretation out of ring 0 is its own project on its own track; nothing here -depends on when it lands. (It landed: [discovery.md](discovery.md), M19–M20.) +depends on when it lands. (It landed: [discovery.md](../os-development-guide/discovery.md), M19–M20.) `device_register` is **idempotent on exact match**: a re-registration with an identical (parent, class, identity, resources) tuple returns the existing id @@ -82,7 +82,7 @@ one world. deadline means wrong binary, wrong protocol version, or wedged before main — apply the stop sequence and the restart policy. Everything else lifecycle-shaped (terminate, the common `ping` liveness call, exit reasons) arrives through -[process-lifecycle.md](process-lifecycle.md)'s vocabulary, not this protocol. +[process-lifecycle.md](../os-development-guide/process-lifecycle.md)'s vocabulary, not this protocol. Assignment stays argv (`usb-xhci-bus `) for now — simple, and it works. The step after `hello` exists is delegation: the manager claims (or is granted) the @@ -97,7 +97,7 @@ from usb-ids.zig — each bus's native language, decoded by the shared ids modul Every driver is spawned with the manager's exit endpoint (`spawnSupervised` — built). On a death notification: -1. **Read the reason** ([process-lifecycle.md](process-lifecycle.md) increment 2). +1. **Read the reason** ([process-lifecycle.md](../os-development-guide/process-lifecycle.md) increment 2). Clean exit → it meant to; don't restart. Fault or missed `hello` deadline → restart with **backoff**, and a crash-loop cap (three fast deaths → mark failed, stop respawning, log loudly; a later `reload` to the manager can retry). @@ -136,7 +136,7 @@ way. ## Increments Increments 1–4 are the lifecycle prerequisites and live in -[process-lifecycle.md](process-lifecycle.md) (claim cleanup on death, exit reasons, +[process-lifecycle.md](../os-development-guide/process-lifecycle.md) (claim cleanup on death, exit reasons, published exit events, signals + `process`). On top of those: 5. **device-manager-protocol**: `hello`, supervised spawn with restart policy; @@ -147,7 +147,7 @@ published exit events, signals + `process`). On top of those: to a manager-internal seam. 8. **Discovery migration** — DONE (M19–M20, 2026-07-13): enumeration moved to ring 3 as swappable per-firmware discoverers — the pci-bus driver (M19) then - the acpi service (M20), see [discovery.md](discovery.md); of the enumerable + the acpi service (M20), see [discovery.md](../os-development-guide/discovery.md); of the enumerable devices, the kernel seeds only the host bridge and the acpi-tables node (the non-enumerable platform nodes — processors, interrupt controllers, the HPET, the loader's framebuffer — stay kernel-seeded too). Matching moved with it: diff --git a/docs/display-plan.md b/docs/device-driver-development-guide/display-plan.md similarity index 89% rename from docs/display-plan.md rename to docs/device-driver-development-guide/display-plan.md index 71d75b7..d77fbe7 100644 --- a/docs/display-plan.md +++ b/docs/device-driver-development-guide/display-plan.md @@ -18,9 +18,9 @@ Read [display.md](display.md) first for the *why*; this is the *what* and the *o ## Conventions -Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations in +Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations in full, kebab-case file names, no `Co-Authored-By` trailers on commits. New user binaries -go through `addUserBinary` in [build.zig](../build.zig) and get packed into the +go through `addUserBinary` in [build.zig](../../build.zig) and get packed into the initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported into the `runtime` module. @@ -29,7 +29,7 @@ initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported i - `zig build test` — host unit tests (compositor math: layer clipping, damage merge, pitch/format blits are all host-testable with a fake framebuffer). - `python3 test/qemu_test.py ` — boots the real kernel in QEMU; assert on the - serial log ([tests.zig](../system/kernel/tests.zig) is the registry). + serial log ([tests.zig](../../system/kernel/tests.zig) is the registry). - The `run-efi` target renders to QEMU's display (`-device VGA,edid=on,xres=1280,yres=720`) — a screenshot confirms pixels for the milestones whose gate is visual. @@ -39,22 +39,22 @@ initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported i Make the boot framebuffer reachable and mappable **write-combining** from user space. -- [x] [device-abi.zig](../library/device/model/device-abi.zig): added `DeviceClass.display`; a +- [x] [device-abi.zig](../../library/device/model/device-abi.zig): added `DeviceClass.display`; a `DisplayInfo{ width, height, pitch, format }` carried on the descriptor; a `flags` field on `ResourceDescriptor` + `resource_flag_write_combining`. -- [x] [devices-broker.zig](../system/kernel/devices-broker.zig): `seedDisplay(base, w, h, +- [x] [devices-broker.zig](../../system/kernel/devices-broker.zig): `seedDisplay(base, w, h, pitch, format)` publishes a root-level `display` node with one WC-flagged `memory` resource `[base, height*pitch]` + the `DisplayInfo`; `displayDevice()` / `displayClaimed()`. Seeded from `kmain` after `devices_broker.init`. -- [x] [process.zig](../system/kernel/process.zig) `systemMmioMap` + paging +- [x] [process.zig](../../system/kernel/process.zig) `systemMmioMap` + paging (`mapUserDeviceInto` gains a `write_combining` bool): a resource's WC flag maps it through the WC PAT slot (`setupPat`) instead of strong-uncacheable. -- [x] [console.zig](../system/kernel/console.zig): `setSuppressed` quiesces `write` while +- [x] [console.zig](../../system/kernel/console.zig): `setSuppressed` quiesces `write` while the display device is claimed (driven from `systemDeviceClaim` / release); the terminal panic + exception paths clear it first so a dying machine still draws. **Gate (met, automated):** the `display` kernel test (`python3 test/qemu_test.py display`, -`displayTest` in [tests.zig](../system/kernel/tests.zig)) asserts the seeded node's shape +`displayTest` in [tests.zig](../../system/kernel/tests.zig)) asserts the seeded node's shape and geometry, then walks the real claim + `mmio_map` path into a throwaway address space and verifies the leaf is **write-combining** (PAT entry 4: PAT bit set, PCD/PWT clear) — with an uncacheable-still-uncacheable regression guard. Chosen over the original @@ -70,7 +70,7 @@ Stand up the named service and the double-buffer, no layers yet. - [x] `library/protocol/display/display-protocol.zig`: `Operation{ info, create_layer, configure_layer, destroy_layer, fill_rect, blit_tile, damage, present }`; `extern` `Request`/`Reply`; size + `maximum_payload` consts. (Model: block/protocol.zig.) -- [x] [abi.zig](../system/abi.zig): `ServiceId.display = 9`. +- [x] [abi.zig](../../system/abi.zig): `ServiceId.display = 9`. - [x] `system/services/display/display.zig`: `main` → enumerate + claim + WC-map the LFB (front) → `mmap` a cacheable back buffer of `height*pitch` → `runtime.service.run`. `info` and a whole-screen `present` (back → front) are live; layer ops fail-stub @@ -78,8 +78,8 @@ Stand up the named service and the double-buffer, no layers yet. - [x] [library/runtime/display.zig](../library/runtime/runtime.zig) (+ barrel export of `display` and `display_protocol`): `info()` and `present()`, cached `.display` lookup with retry (model: block.zig). -- [x] [init.zig](../system/services/init/init.zig): `"display"` added to `boot_services`. -- [x] [build.zig](../build.zig): `display-protocol` module on the runtime; `display` exe +- [x] [init.zig](../../system/services/init/init.zig): `"display"` added to `boot_services`. +- [x] [build.zig](../../build.zig): `display-protocol` module on the runtime; `display` exe via `addUserBinary`; packed into the initial-ramdisk; installed to `/system/services/display`. - [x] **Kernel fix the back buffer surfaced:** `mmap` was capped at 256 pages (1 MiB) by @@ -105,7 +105,7 @@ The heart: composite an ordered layer stack, present only what changed. - [x] `create_layer` / `configure_layer` (damages old + new footprints) / `destroy_layer`, `fill_rect`, `blit_tile` (reads the inline tile from the IPC payload, unaligned-safe), `damage`, `present`. -- [x] Pure, host-tested [compositor.zig](../system/services/display/compositor.zig): `Rect` +- [x] Pure, host-tested [compositor.zig](../../system/services/display/compositor.zig): `Rect` (intersect/unite), `Surface`, `fillRect`, `composite` (opaque, clipped to a damage rect), `blitTile`. `present` clears the damaged region to the wallpaper, paints the visible layers bottom-to-top (z-sorted), and flushes just that rect back → front (WC). @@ -148,10 +148,10 @@ still pass, and the default `zig build` is clean. - [x] The three integration cases exist and pass: `display` (D1 handoff, kernel), `display-service` (D2/D3 compositor + self-check), and `display-demo` (D4 full pipeline: spawn `display` + `display-demo`, match `display-demo: ok`) — - [tests.zig](../system/kernel/tests.zig) + [qemu_test.py](../test/qemu_test.py). Plus + [tests.zig](../../system/kernel/tests.zig) + [qemu_test.py](../../test/qemu_test.py). Plus the pure host tests (`zig build test`). - [x] [display.md](display.md) updated to the built state (the "Verifying it" section names - the real cases); [README index](README.md) entry present (#19); the `display-track` + the real cases); [README index](../README.md) entry present (#19); the `display-track` memory marked DONE with the commits. **Gate (met):** `python3 test/qemu_test.py display display-service display-demo` all pass, diff --git a/docs/display-v2-plan.md b/docs/device-driver-development-guide/display-v2-plan.md similarity index 97% rename from docs/display-v2-plan.md rename to docs/device-driver-development-guide/display-v2-plan.md index 7da5123..179a8cf 100644 --- a/docs/display-v2-plan.md +++ b/docs/device-driver-development-guide/display-v2-plan.md @@ -16,11 +16,11 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, ## Conventions -Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations, +Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through `addUserBinary` and get packed into the initial-ramdisk; protocols are `b.addModule("…-protocol", …)` imported into `runtime`; new syscalls extend -[abi.zig](../system/abi.zig) `SystemCall` + a `library/runtime` wrapper. +[abi.zig](../../system/abi.zig) `SystemCall` + a `library/runtime` wrapper. ## How to verify along the way @@ -59,7 +59,7 @@ is the only backend), and `zig build test` stays green. ## V2 — The shared-memory cross-process capability (kernel) ✅ -- [x] [abi.zig](../system/abi.zig): `shared_memory_create` (34) / `shared_memory_map` (35) syscalls + a +- [x] [abi.zig](../../system/abi.zig): `shared_memory_create` (34) / `shared_memory_map` (35) syscalls + a `shared_memory_test` service id. Handlers in process.zig: `shared_memory_create(len)` allocates contiguous, zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability handle, maps them into the caller's shared-memory arena → returns virtual_address + handle; `shared_memory_map(cap)` diff --git a/docs/display-v2.md b/docs/device-driver-development-guide/display-v2.md similarity index 98% rename from docs/display-v2.md rename to docs/device-driver-development-guide/display-v2.md index 3a226e7..c2d7d58 100644 --- a/docs/display-v2.md +++ b/docs/device-driver-development-guide/display-v2.md @@ -144,4 +144,4 @@ path in VMs**, where danos development happens. The framebuffer floor never goes - [display.md](display.md) — v1: the compositor, the GOP-vs-device split, the WC discipline. - [display-v2-plan.md](display-v2-plan.md) — the ordered build-out. - [driver-model.md](driver-model.md) — claim / `mmio_map` / MSI / capability passing (M13). -- [resilience.md](resilience.md) — the restart machinery the hot-attach leans on. +- [resilience.md](../os-development-guide/resilience.md) — the restart machinery the hot-attach leans on. diff --git a/docs/display.md b/docs/device-driver-development-guide/display.md similarity index 87% rename from docs/display.md rename to docs/device-driver-development-guide/display.md index 5cf2484..c25b6f2 100644 --- a/docs/display.md +++ b/docs/device-driver-development-guide/display.md @@ -1,12 +1,12 @@ # The display service: a framebuffer compositor -The [framebuffer](framebuffer.md) the loader hands over is a flat block of pixel -memory, and the kernel's [bootstrap console](../system/kernel/console.zig) draws text +The [framebuffer](../os-development-guide/framebuffer.md) the loader hands over is a flat block of pixel +memory, and the kernel's [bootstrap console](../../system/kernel/console.zig) draws text into it directly. That console is a stop-gap. The **display service** (`system/services/display/`) is the real thing: an ordinary ring-3 process that *owns* the framebuffer, composes a stack of **layers** into an off-screen back buffer, and **presents** finished frames to the screen — the display half of the GUI track -([vision.md](vision.md)), the sibling of the [input service](input.md). +([vision.md](../vision.md)), the sibling of the [input service](input.md). This note is the architecture and the reasoning behind it. The concrete build order lives in [display-plan.md](display-plan.md). @@ -20,20 +20,20 @@ which one you're holding decides what you can do. - **GOP is firmware's *temporary* driver** for the display controller. It gives you a linear framebuffer pointer and can set video modes — but only until - `ExitBootServices`. The loader already leans on this: [`queryFramebuffer`](../boot/efi.zig) + `ExitBootServices`. The loader already leans on this: [`queryFramebuffer`](../../boot/efi.zig) reads the monitor's EDID, picks the native mode, and calls `set_mode` **before** - exiting ([gop.md](gop.md)). Once the kernel runs, GOP is **gone** — no `set_mode`, no + exiting ([gop.md](../os-development-guide/gop.md)). Once the kernel runs, GOP is **gone** — no `set_mode`, no mode list, no EDID. What survives is the frozen snapshot in - [`BootInformation.framebuffer`](../system/boot-handoff.zig): `{base, width, height, + [`BootInformation.framebuffer`](../../system/boot-handoff.zig): `{base, width, height, pitch, format, refresh_hz}`, and nothing more. - **The PCI class-0x03 device is the raw controller** — BARs, config space, registers, IO ports. It is what you actually *own* after boot. On QEMU's emulated adapter - ([`-device VGA,edid=on`](../build.zig), the Bochs VBE/DISPI model) the `base` GOP handed + ([`-device VGA,edid=on`](../../build.zig), the Bochs VBE/DISPI model) the `base` GOP handed you *is* that device's linear-framebuffer BAR — the same physical memory, seen through a different door. On a real discrete GPU, GOP's `base` is an aperture inside the GPU's VRAM BAR. danos already decodes this device - ([pci-class.zig](../library/device/pci/pci-class.zig) has the full `display` namespace, and + ([pci-class.zig](../../library/device/pci/pci-class.zig) has the full `display` namespace, and `pci-bus` already reports it to the [device manager](device-manager.md) with its class triple) — but nothing binds it yet. @@ -63,16 +63,16 @@ rest of the system hasn't had to face: 1. **The framebuffer is kernel-only today.** It arrives through the boot handoff, is mapped into the kernel's physmap, and is touched only by - [`console.zig`](../system/kernel/console.zig). It is *not* a - [devices-broker](../system/kernel/devices-broker.zig) node, so + [`console.zig`](../../system/kernel/console.zig). It is *not* a + [devices-broker](../../system/kernel/devices-broker.zig) node, so `device.claim`/`mmio_map` cannot reach it, and there is no framebuffer - [syscall](syscall.md). A user-space display service needs a **new mechanism just to + [syscall](../os-development-guide/syscall.md). A user-space display service needs a **new mechanism just to touch the pixels**. (See "The handoff" below — this is built.) 2. **danos had no cross-process shared memory.** At v1 the memory syscalls were `mmap` (private, zeroed), `mmio_map` (a *claimed device's* MMIO), and `dma_alloc` (new pinned physical). The block driver's "pass a buffer by physical address" trick - ([block/protocol.zig](../library/protocol/block/block-protocol.zig)) works *only because its + ([block/protocol.zig](../../library/protocol/block/block-protocol.zig)) works *only because its consumer is DMA hardware*. A compositor that CPU-reads and blends client layers can't use it — it would have to *map* another process's memory, which nothing allowed. v1 sidesteps it entirely (see "What v1 does not do"); v2 has since built the primitive @@ -103,10 +103,10 @@ rest of the system hasn't had to face: ``` The bring-up sequence mirrors a hardware driver's — it is the -[`usb-xhci-bus` `initialise`](../system/drivers/usb-xhci-bus/usb-xhci-bus.zig) shape +[`usb-xhci-bus` `initialise`](../../system/drivers/usb-xhci-bus/usb-xhci-bus.zig) shape (claim → `mmio_map` → run loop) — and the request/reply service shell is the -[FAT](../system/services/fat/fat.zig) / [input](../system/services/input/input.zig) shape -([`service.run`](../library/kernel/service.zig) with a `protocol.zig` of +[FAT](../../system/services/fat/fat.zig) / [input](../../system/services/input/input.zig) shape +([`service.run`](../../library/kernel/service.zig) with a `protocol.zig` of `extern struct` messages and an `Operation` tag). **One process, for now.** v1 is a *single* service that both owns the framebuffer and @@ -119,12 +119,12 @@ second backend or a second monitor appears; until then it is complexity with no The framebuffer crosses into user space through the machinery that already exists for every other device, rather than a bespoke syscall — so it inherits ownership, -release-on-death, and re-claim-on-restart for free (the [resilience](resilience.md) +release-on-death, and re-claim-on-restart for free (the [resilience](../os-development-guide/resilience.md) story: a crashed display service returns the LFB to the kernel, and its restart re-claims it). - The kernel seeds a synthetic **display-class** node into the - [devices-broker](../system/kernel/devices-broker.zig) at init (`seedDisplay`), from + [devices-broker](../../system/kernel/devices-broker.zig) at init (`seedDisplay`), from `BootInformation.framebuffer`: one `ResourceKind.memory` resource spanning `[base, height*pitch]`, tagged **write-combining**, plus a small `DisplayInfo{width, height, pitch, format, refresh_hz}` (the memory resource says *where* @@ -135,7 +135,7 @@ re-claims it). - The service `device.claim`s it and `mmio_map`s the resource. The map is **write-combining**, not the strong-uncacheable that `mmio_map` uses for register MMIO. The kernel already programs a WC PAT slot for its own console - ([`setupPat`](../system/kernel/architecture/x86_64/paging.zig)); this reaches it from + ([`setupPat`](../../system/kernel/architecture/x86_64/paging.zig)); this reaches it from the user mapping path. **This matters:** an uncacheable framebuffer makes the back→front blit unusably slow. - On `claim`, the kernel's bootstrap console goes quiet, so the two never fight over the @@ -143,7 +143,7 @@ re-claims it). panic on screen wins. The display service is a **named boot service**: `init` spawns it by name alongside -`input`/`device-manager`/`fat` ([init.zig](../system/services/init/init.zig)), and it +`input`/`device-manager`/`fat` ([init.zig](../../system/services/init/init.zig)), and it self-discovers the display node with `device.enumerate` (matching on `DeviceClass.display`). The [device manager](device-manager.md) matching path (PCI class 0x03 → a driver) is reserved for the future *native* backend, not this singleton synthetic node. @@ -160,8 +160,8 @@ Two buffers, with deliberately different memory types: So a frame is: compose every dirty layer into the cacheable back buffer, then **present** — copy the changed regions back→front in sequential, WC-friendly writes. Two details the -[framebuffer](framebuffer.md) note already establishes carry over: step rows by `pitch`, -not `width*4`; and handle both `rgbx` and `bgrx` [pixel formats](gop.md). +[framebuffer](../os-development-guide/framebuffer.md) note already establishes carry over: step rows by `pitch`, +not `width*4`; and handle both `rgbx` and `bgrx` [pixel formats](../os-development-guide/gop.md). ## Flicker vs. tearing — what double buffering does and doesn't buy @@ -188,10 +188,10 @@ The compositor holds an **ordered stack of layers**. Each layer has a rectangle, z-order, a visibility flag, and a surface. Presenting walks the stack bottom-to-top, painting each dirty layer into the back buffer, then flushes the damage to the front. Damage is tracked by one of two interchangeable trackers behind a compile-time -`damage_mode` A/B switch ([display.zig](../system/services/display/display.zig)): a +`damage_mode` A/B switch ([display.zig](../../system/services/display/display.zig)): a free-form dirty-rectangle **list** (tight bounds, heuristic merging) or a fixed 64-px **tile grid** (exact O(1) merging, tile-quantized repaints) — the grid is the default; -[compositor.zig](../system/services/display/compositor.zig) has both, with the trade-off +[compositor.zig](../../system/services/display/compositor.zig) has both, with the trade-off discussion. In v1 the surfaces are **server-owned**, and clients draw into them with a small @@ -210,7 +210,7 @@ shell, a terminal, a cursor, and a wallpaper: | `present` | request a repaint: composited at the next frame-clock tick | Text is intentionally *not* an operation — a client renders glyphs by blitting tiles -(the [PSF font](../system/kernel/font.psf) path the console already uses can move into a +(the [PSF font](../../system/kernel/font.psf) path the console already uses can move into a client). Keeping the protocol to rectangles and tiles keeps the compositor small and the policy in the client. @@ -224,8 +224,8 @@ pixels on screen synchronously (initialisation, the self-checks) bypass the cloc ## `display` -Clients speak the protocol through a new [`library/client/display/display.zig`](../library/client/display/display.zig), -the [`block`](../library/device/block/block.zig) shape (a cached `.display` lookup +Clients speak the protocol through a new [`library/client/display/display.zig`](../../library/client/display/display.zig), +the [`block`](../../library/device/block/block.zig) shape (a cached `.display` lookup with a boot-race retry): `display.info()`, a `Layer` handle with `fill` / `blitTile` / `damage`, and `present()`. Application code never issues the raw syscalls — it calls the client module, as with every other danos service. @@ -234,7 +234,7 @@ client module, as with every other danos service. The compositor is the single owner of the framebuffer — only the main `service.run` loop touches the backend and the layer stack. Tracking the mouse without breaking that -ownership is the display's first use of [threads](threading.md): the service is built +ownership is the display's first use of [threads](../os-development-guide/threading.md): the service is built multi-threaded (`addThreadedUserBinary`) and, at startup, spawns a **mouse-listener thread** beside the compositor loop. @@ -242,7 +242,7 @@ thread** beside the compositor loop. (`input.subscribeMouse()`), accumulates the relative `dx`/`dy` motion into an absolute cursor position clamped to the screen, and hands it to the compositor. It never touches the compositor — so no lock guards the framebuffer. A parked `next()` leaves its core - free to halt ([halting.md](halting.md)). + free to halt ([halting.md](../os-development-guide/halting.md)). - **The channel.** A single-slot *latest-value* cell (`CursorChannel`) guarded by a `Thread.Mutex`: the renderer wants where the cursor *is now*, not a replay of every delta, so a new position overwrites the old. The listener also **pokes** the @@ -254,13 +254,13 @@ thread** beside the compositor loop. which is just a top-z compositor layer — with the existing `configure` + `present` path (it damages the old and new footprints, so only those two rectangles repaint). -Two threading facts shape this (both in [threading.md](threading.md)). IPC **handles do +Two threading facts shape this (both in [threading.md](../os-development-guide/threading.md)). IPC **handles do not cross threads**, so the listener can't reuse the main loop's endpoint handle — it `ipc.lookup(.display)`s its *own* handle to the same endpoint to poke through. And a multi-threaded service doing concurrent IPC is why the kernel's endpoint-create / register / lookup syscalls now serialize under the big kernel lock. Shared fate applies: a fault in the listener takes the whole display down, and the supervisor restarts the process -([resilience.md](resilience.md)). +([resilience.md](../os-development-guide/resilience.md)). ## What v1 does not do (and why that's fine) @@ -284,7 +284,7 @@ both are clean additions behind the interfaces v1 establishes. ## Verifying it -Four QEMU test cases ([tests.zig](../system/kernel/tests.zig), `python3 +Four QEMU test cases ([tests.zig](../../system/kernel/tests.zig), `python3 test/qemu_test.py `), each layering on the last: - **`display`** — the kernel handoff: the seeded `display` device is shaped correctly and @@ -317,8 +317,8 @@ packing are additionally covered by pure host unit tests under `zig build test`. ## See also -- [framebuffer.md](framebuffer.md) — the linear framebuffer, pitch vs. width, `volatile`. -- [gop.md](gop.md) — GOP, and why only linear RGBX/BGRX modes are paintable. +- [framebuffer.md](../os-development-guide/framebuffer.md) — the linear framebuffer, pitch vs. width, `volatile`. +- [gop.md](../os-development-guide/gop.md) — GOP, and why only linear RGBX/BGRX modes are paintable. - [input.md](input.md) — the sibling service; the async `ipc_send` fan-out. - [driver-model.md](driver-model.md) — claim / `mmio_map`, capability passing, the trust model. - [device-manager.md](device-manager.md) — matching and supervision (the native backend's route). diff --git a/docs/driver-model.md b/docs/device-driver-development-guide/driver-model.md similarity index 97% rename from docs/driver-model.md rename to docs/device-driver-development-guide/driver-model.md index 7f5adef..f696644 100644 --- a/docs/driver-model.md +++ b/docs/device-driver-development-guide/driver-model.md @@ -34,7 +34,7 @@ plain bus driver with no controller — a USB hub — is also a real thing. danos already has the right central structure. `system/kernel/devices-broker.zig` holds a table of `DeviceDescriptor`, each with a parent, a class, and a set of resources. Firmware discovery -seeds it ([discovery.md](discovery.md)); `device_register` grows it. +seeds it ([discovery.md](../os-development-guide/discovery.md)); `device_register` grows it. Three invariants make it a capability system rather than a directory: @@ -97,7 +97,7 @@ A "family" is two modules, not one: danos already has one of each: `library/device/pci/pci.zig` is a logic module (the `Function` view of a claimed PCI function), -[`library/protocol/vfs/vfs-protocol.zig`](../library/protocol/vfs/vfs-protocol.zig) is a +[`library/protocol/vfs/vfs-protocol.zig`](../../library/protocol/vfs/vfs-protocol.zig) is a protocol module shared by the mount backends (today the fat server) and their clients. (The user-space VFS server it was originally written against has since retired — path routing moved into the kernel, `system/kernel/vfs.zig`'s `fs_resolve` — but the protocol @@ -199,7 +199,7 @@ class driver, the device manager, or the kernel may share them freely. `system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a fresh ring-3 process; `name` becomes the child's argv[0] and the optional NUL-separated `arguments` blob its argv[1..], delivered on a SysV entry stack - ([sysv.md](sysv.md)). This is what + ([sysv.md](../os-development-guide/sysv.md)). This is what turned the device manager from "log the match" into "run the driver": the kernel now spawns only `init`, `init` spawns the services, and the **device-manager** discovers the hardware and spawns each driver ([drivers.md](drivers.md)). Ungated for now — a @@ -321,7 +321,7 @@ sprinkling of `asm volatile`: | DMA cache coherency | coherent; nothing to do | **not guaranteed**; needs non-cacheable buffers or cache maintenance | x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away -with a compiler barrier alone. ARM is not, and [vision.md](vision.md) makes ARM the win +with a compiler barrier alone. ARM is not, and [vision.md](../vision.md) makes ARM the win condition. Build the abstraction while there is one caller to fix. (Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full @@ -398,6 +398,6 @@ from hand-rolling `*volatile` and getting ARM wrong. ## See also - [drivers.md](drivers.md) — how to write one, concretely. -- [discovery.md](discovery.md) / [acpi.md](acpi.md) — where the device table comes from. +- [discovery.md](../os-development-guide/discovery.md) / [acpi.md](../os-development-guide/acpi.md) — where the device table comes from. - [ipc.md](ipc.md) — endpoints, badges, and the notification path an IRQ arrives on. -- [resilience.md](resilience.md) — restart, the reason any of this is worth the trouble. +- [resilience.md](../os-development-guide/resilience.md) — restart, the reason any of this is worth the trouble. diff --git a/docs/drivers.md b/docs/device-driver-development-guide/drivers.md similarity index 97% rename from docs/drivers.md rename to docs/device-driver-development-guide/drivers.md index 922aa87..4b3c22f 100644 --- a/docs/drivers.md +++ b/docs/device-driver-development-guide/drivers.md @@ -4,13 +4,13 @@ In a monolithic kernel a driver is a function call away from everything: it runs ring 0, dereferences any physical address, and its interrupt handler *is* the ISR. In danos a driver is **an ordinary ring-3 process**. It has its own address space, it can crash without taking the kernel with it, and — the point of this document — it -can be restarted ([resilience](resilience.md)). +can be restarted ([resilience](../os-development-guide/resilience.md)). That leaves three questions the kernel has to answer, because a process can't answer them for itself: 1. **What hardware exists?** → `device_enumerate`, over the device table discovery built - ([discovery](discovery.md), [acpi](acpi.md)). + ([discovery](../os-development-guide/discovery.md), [acpi](../os-development-guide/acpi.md)). 2. **How do I touch its registers?** → `device_claim` + `mmio_map`: the kernel maps the device's physical MMIO window into your address space, and from then on it's plain memory. No syscall per register access. @@ -51,7 +51,7 @@ the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Every is the **driver supervisor**. It does the three steps a monolithic kernel would do in its probe path, entirely from ring 3: 1. **Discover** — `device_enumerate` snapshots the device table the kernel built from - ACPI/PCI ([discovery](discovery.md)). + ACPI/PCI ([discovery](../os-development-guide/discovery.md)). 2. **Match** — for each device it looks up a driver. The match policy is code, a few small per-bus tables: from the boot snapshot only the PCI host bridge matches (→ `pci-bus`); everything else arrives later as bus reports and matches on @@ -269,7 +269,7 @@ already owns. A device with **no resources** is legal and common. A USB device is reached through its controller, not by MMIO, so it gets `resource_count = 0`. -See [`system/drivers/pci-bus/pci-bus.zig`](../system/drivers/pci-bus/pci-bus.zig) for a +See [`system/drivers/pci-bus/pci-bus.zig`](../../system/drivers/pci-bus/pci-bus.zig) for a real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function it finds as a child — and [driver-model.md](driver-model.md) for how bus drivers, class drivers and host controller drivers fit together. @@ -293,7 +293,7 @@ the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so uncacheable, physical address exposed), **memory barriers** (`library/device/mmio`'s `memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier`, imported as the `mmio` module), **fault isolation** (a ring-3 fault kills only the faulting process — `killCurrentProcess` — and the machine keeps running, -[resilience](resilience.md)), and **reclaim + restart on death** (every path out of a +[resilience](../os-development-guide/resilience.md)), and **reclaim + restart on death** (every path out of a process releases its claims and IRQ/MSI bindings — `releaseAllOwnedBy`, `irq.releaseOwner` — and the device manager respawns the driver with backoff, [device-manager.md](device-manager.md)). What remains: @@ -394,7 +394,7 @@ Claiming and mapping is half of being a danos driver; the other half is the - Build on `service.run` — one replyWait loop folding protocol requests, signals, and notifications into callbacks. The harness answers the universal zero-length ping and turns `terminate` into a clean exit for you - ([process-lifecycle.md](process-lifecycle.md)). + ([process-lifecycle.md](../os-development-guide/process-lifecycle.md)). - A driver spawned with an assignment (its device id as argv[1]) sends the versioned `hello` to the device manager inside the deadline, and a **bus** driver reports what it discovers with `child_added` diff --git a/docs/input.md b/docs/device-driver-development-guide/input.md similarity index 88% rename from docs/input.md rename to docs/device-driver-development-guide/input.md index 28c3765..7e45b70 100644 --- a/docs/input.md +++ b/docs/device-driver-development-guide/input.md @@ -5,14 +5,14 @@ window server, a logger. None of them owns the hardware, and the driver should n who is listening. So between the drivers and the listeners sits the **input service** (`system/services/input/`): drivers **publish** events to it, programs **subscribe**, and it fans each event out to every interested subscriber. It is an ordinary ring-3 process -reached over IPC, like the [FAT server](../system/services/fat/fat.zig) — no kernel knows +reached over IPC, like the [FAT server](../../system/services/fat/fat.zig) — no kernel knows what a key is. ## One service, several device classes The service carries three device classes today — **keyboard**, **mouse**, and **joystick/gamepad** — and is built to take more -([protocol.zig](../library/protocol/input/input-protocol.zig)). Each class has its own typed +([protocol.zig](../../library/protocol/input/input-protocol.zig)). Each class has its own typed event: - `KeyEvent` — `key_down`/`key_up` (physical make/break) and `key_press` (a character was @@ -45,7 +45,7 @@ consequences decide the whole design: `ipc_call`, it would block until each subscriber replied. `ipc_call` has no timeout, and a subscriber's endpoint is an *unregistered* capability the kernel's death path cannot reach (since display v2's V6, `killOwnedEndpointsLocked` in - [ipc-synchronous.zig](../system/kernel/ipc-synchronous.zig) marks a dead owner's + [ipc-synchronous.zig](../../system/kernel/ipc-synchronous.zig) marks a dead owner's *registered* endpoints dead and wakes parked callers with `-EPEER` — but unregistered ones just drop with the task's handle table). One subscriber that exits mid-delivery would wedge input for everyone. That is the opposite of the resilience the microkernel @@ -66,7 +66,7 @@ the badge (distinguishing it from a bare IRQ/child-exit notification), the sende in the low bits, and the payload in the receive buffer, with no reply owed. The queue holds 16 messages per endpoint; a full queue **drops the oldest**, because a buffered message is discrete data, not a coalescing "level" like an interrupt. See -[ipc-synchronous.zig](../system/kernel/ipc-synchronous.zig) (`sendLocked`, `popPost`, and +[ipc-synchronous.zig](../../system/kernel/ipc-synchronous.zig) (`sendLocked`, `popPost`, and the `replyWait` receive loop). This is the async counterpart of `ipc_call`, and the input service is its first consumer. @@ -89,7 +89,7 @@ This is the async counterpart of `ipc_call`, and the input service is its first - A **subscriber** calls `input.subscribe(mask)` — or a typed helper: `subscribeKeyboard()`, `subscribeMouse()`, `subscribeJoystick()` (one class, `next()` returns the decoded event), or `subscribeAll()` (every class, `next()` returns a tagged `InputEvent`) - ([library/client/input/input.zig](../library/client/input/input.zig)). It creates its own endpoint + ([library/client/input/input.zig](../../library/client/input/input.zig)). It creates its own endpoint and hands it to the service as a **capability** (M13 capability passing — the input service is that feature's first real user), along with its `device_mask`. Then it loops on `next()`, a `replyWait` on that endpoint returning each pushed event. @@ -98,7 +98,7 @@ This is the async counterpart of `ipc_call`, and the input service is its first `publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at once; the service's own fan-out is asynchronous, so publishing never blocks on a slow subscriber. -- The **service** ([input.zig](../system/services/input/input.zig)) keeps a small subscriber +- The **service** ([input.zig](../../system/services/input/input.zig)) keeps a small subscriber table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the event to every subscriber whose mask includes the event's device class. On `subscribe` it stores the passed capability and mask and, as housekeeping, prunes any slot whose owning @@ -113,18 +113,18 @@ the service delivers to its endpoint, which only the same thread could receive). - **The keyboard is real.** The `ps2-bus` driver owns PNP0303, which carries *both* the 0x60/0x64 ports and IRQ1, so reading the hardware lives in the bus, not in - [keyboard.zig](../system/drivers/ps2-bus/keyboard.zig): the bus binds IRQ1 and, on each + [keyboard.zig](../../system/drivers/ps2-bus/keyboard.zig): the bus binds IRQ1 and, on each interrupt, drains port 0x60, routing every byte by the status register's auxiliary-output bit to whichever child driver **attached** for that device (an `AttachRequest` to the well-known `ps2_bus` service, carrying the child's endpoint as a capability; the bytes then arrive as asynchronous `ForwardedByte` messages, so the IRQ path never blocks on a child). The keyboard driver decodes the stream — scancode **set 2**, what the keyboard sends with the 8042's legacy translation off, decoded by - [scancode.zig](../system/drivers/ps2-bus/scancode.zig) into USB HID usage keycodes with + [scancode.zig](../../system/drivers/ps2-bus/scancode.zig) into USB HID usage keycodes with make/break, typematic-repeat, and modifier tracking (host-tested under `zig build test`) — and publishes real `key_down`/`key_press`/`key_up` events. - **Keycode → character** is wired in: the keyboard driver fills a `key_press` event's - `character` through [`library/xkeyboard-config`](../library/xkeyboard-config/README.md) + `character` through [`library/xkeyboard-config`](../../library/xkeyboard-config/README.md) (`xkb.map(layout, keycode, mods)` → keysym + Unicode character), synthesizing the ASCII control characters for Enter/Tab/Backspace/Escape, whose keysyms map to no Unicode. The layout defaults to `us`; the bus can pass another as the driver's argv[2] — the seam for @@ -132,9 +132,9 @@ the service delivers to its endpoint, which only the same thread could receive). - **The mouse is real too.** IRQ12 is enumerated on the auxiliary device's own ACPI node (PNP0F13), so the bus claims that node alongside the controller and routes both IRQs to its one endpoint, acking whichever line the notification's badge names. - [mouse.zig](../system/drivers/ps2-bus/mouse.zig) attaches the way the keyboard does and + [mouse.zig](../../system/drivers/ps2-bus/mouse.zig) attaches the way the keyboard does and assembles the forwarded bytes with - [mouse-packet.zig](../system/drivers/ps2-bus/mouse-packet.zig) (three-byte stream-mode + [mouse-packet.zig](../../system/drivers/ps2-bus/mouse-packet.zig) (three-byte stream-mode packets: sync/overflow handling, nine-bit movement, screen-convention `dy` — host-tested under `zig build test`) into `button_down`/`button_up` transitions and `motion` events. **Follow-up:** the IntelliMouse magic-knock for a scroll wheel (four-byte packets) and @@ -150,7 +150,7 @@ the service delivers to its endpoint, which only the same thread could receive). ## Verifying it The `input` case (`python3 test/qemu_test.py input`, in -[tests.zig](../system/kernel/tests.zig) `inputTest`) boots the real kernel and spawns the +[tests.zig](../../system/kernel/tests.zig) `inputTest`) boots the real kernel and spawns the service, the synthetic source (which cycles keyboard, mouse, and joystick events), and a subscriber that took all three classes. It passes only when the subscriber heartbeats `input-test: ok` — proof that an event travelled source → service → subscriber over IPC, @@ -160,5 +160,5 @@ serial line names the class received, so the log shows all three arriving on one ## See also - [ipc.md](ipc.md) — the synchronous rendezvous and the notification path `ipc_send` extends. -- [syscall.md](syscall.md) — the system-call surface, including `ipc_send`. +- [syscall.md](../os-development-guide/syscall.md) — the system-call surface, including `ipc_send`. - [driver-model.md](driver-model.md) — class drivers, capability passing (M13), the trust model. diff --git a/docs/intel-igpu.md b/docs/device-driver-development-guide/intel-igpu.md similarity index 100% rename from docs/intel-igpu.md rename to docs/device-driver-development-guide/intel-igpu.md diff --git a/docs/ipc.md b/docs/device-driver-development-guide/ipc.md similarity index 93% rename from docs/ipc.md rename to docs/device-driver-development-guide/ipc.md index 117f974..fb48144 100644 --- a/docs/ipc.md +++ b/docs/device-driver-development-guide/ipc.md @@ -1,7 +1,7 @@ # IPC: message-passing channels Inter-process communication is the **backbone of a microkernel**. Once drivers and -services run isolated in their own address spaces ([vision](vision.md)), they can't +services run isolated in their own address spaces ([vision](../vision.md)), they can't just call each other — a request becomes a **message**. In a microkernel, whatever was a function call across a monolithic kernel is IPC, so it's a first-class concern, not an afterthought. @@ -18,7 +18,7 @@ There are two layers, built a milestone apart: The first form is a **bounded blocking channel** (`system/kernel/ipc.zig`): a fixed-size ring buffer of messages with a producer/consumer rendezvous, built on the -scheduler's [wait queues](scheduling.md). +scheduler's [wait queues](../os-development-guide/scheduling.md). `Channel(T, capacity)` is generic over the message type and buffer size. It holds a ring buffer, a count, and two wait queues: @@ -37,7 +37,7 @@ Two details make it correct: rather than assuming the slot is still available — another waiter may have taken it first. This is the standard guard against spurious or racing wakeups. - **One critical section.** `send`/`receive` run under the [big kernel - lock](smp.md) (`sync.enter` / `sync.leave`), which disables interrupts on this + lock](../os-development-guide/smp.md) (`sync.enter` / `sync.leave`), which disables interrupts on this core *and* takes the kernel's one spinlock — since SMP, the interrupt flag alone is not atomicity, because `cli` on one core does nothing to another. So checking the condition and committing the block/enqueue happen atomically both with respect @@ -47,7 +47,7 @@ Two details make it correct: ## Verifying it -The `ipc` test (see [testing.md](testing.md)) runs a producer and a consumer passing +The `ipc` test (see [testing.md](../testing.md)) runs a producer and a consumer passing **100 messages through a 4-slot channel**. The small buffer means the channel goes full and empty over and over, so both the blocking-send and blocking-receive paths are exercised heavily. The messages arrive intact and in order (their sum is the @@ -114,7 +114,7 @@ This is what makes a user-space driver possible at all, and it's the subject of ## Lifecycle conventions over IPC (M17) -Three conventions from [process-lifecycle.md](process-lifecycle.md) ride the +Three conventions from [process-lifecycle.md](../os-development-guide/process-lifecycle.md) ride the notification mechanism: - **Signals** arrive as notifications on the endpoint a process nominated with diff --git a/docs/nvidia-gpus.md b/docs/device-driver-development-guide/nvidia-gpus.md similarity index 100% rename from docs/nvidia-gpus.md rename to docs/device-driver-development-guide/nvidia-gpus.md diff --git a/docs/usb-hub.md b/docs/device-driver-development-guide/usb-hub.md similarity index 100% rename from docs/usb-hub.md rename to docs/device-driver-development-guide/usb-hub.md diff --git a/docs/danos-file-system-hierarchy-FSH.md b/docs/file-system-development/danos-file-system-hierarchy-FSH.md similarity index 97% rename from docs/danos-file-system-hierarchy-FSH.md rename to docs/file-system-development/danos-file-system-hierarchy-FSH.md index 2038736..a1ddaed 100644 --- a/docs/danos-file-system-hierarchy-FSH.md +++ b/docs/file-system-development/danos-file-system-hierarchy-FSH.md @@ -49,7 +49,7 @@ addressed by device id. `/dev` is the much smaller set of devices that have a dr willing to serve them, addressed by name. A device node is not a file the VFS can read. The bytes live in a driver process -([drivers.md](drivers.md)), so opening a `/dev` name has to resolve to that driver's +([drivers.md](../device-driver-development-guide/drivers.md)), so opening a `/dev` name has to resolve to that driver's IPC endpoint, and subsequent reads and writes are calls against it. Resolve-to-endpoint is exactly what the kernel's `fs_resolve` already does for any mounted backend, and `FileStatus.kind` is the field that marks a device node; **what is not implemented today @@ -92,7 +92,7 @@ descriptor ring and left to read and write memory on its own. That ring is exact **`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its physical address disclosed — and **`/lib/device/mmio`**'s barriers order the descriptor writes against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver -can be written today (the M14/M15 work in [driver-model.md](driver-model.md); the earlier +can be written today (the M14/M15 work in [driver-model.md](../device-driver-development-guide/driver-model.md); the earlier "cannot host a block driver at all" is no longer true). What is *not* yet true is that it is safe. A device programmed with an arbitrary physical @@ -125,4 +125,4 @@ caller can tell a character device from a regular file. options on this kernel are `RDRAND`/`RDSEED` where CPUID advertises them, and the HPET counter's low bits as a poor fallback. Neither is a seeded CSPRNG, and a `/dev/random` that is merely unpredictable-looking is worse than none — nothing should be keyed from -it until it is a real one. \ No newline at end of file +it until it is a real one. diff --git a/docs/vfs-protocol.md b/docs/file-system-development/vfs-protocol.md similarity index 99% rename from docs/vfs-protocol.md rename to docs/file-system-development/vfs-protocol.md index a5ae41f..be958eb 100644 --- a/docs/vfs-protocol.md +++ b/docs/file-system-development/vfs-protocol.md @@ -9,7 +9,7 @@ > backend, unchanged. The Zig source of truth is `library/protocol/vfs/vfs-protocol.zig` > (the `vfs-protocol` module), whose unit test pins a sample of the sizes > and values below. This page is the **language-neutral wire specification** -> of that contract — what a Rust or C client implements ([vdso.md](vdso.md) +> of that contract — what a Rust or C client implements ([vdso.md](../os-development-guide/vdso.md) > explains why the IPC protocols, not the syscall numbers, are danos's > public ABI). diff --git a/docs/os-developer-guide/README.MD b/docs/os-development-guide/README.MD similarity index 100% rename from docs/os-developer-guide/README.MD rename to docs/os-development-guide/README.MD diff --git a/docs/acpi.md b/docs/os-development-guide/acpi.md similarity index 100% rename from docs/acpi.md rename to docs/os-development-guide/acpi.md diff --git a/docs/architecture.md b/docs/os-development-guide/architecture.md similarity index 97% rename from docs/architecture.md rename to docs/os-development-guide/architecture.md index 1398f5e..9e217d3 100644 --- a/docs/architecture.md +++ b/docs/os-development-guide/architecture.md @@ -75,9 +75,9 @@ There are really two independent questions, and it's worth not conflating them: - **`system/kernel/architecture/x86_64/paging.zig`** — the kernel's page tables and address-space management (see [paging.md](paging.md)). - **`system/kernel/architecture/x86_64/apic.zig`** / **`ioapic.zig`** — the Local APIC, its timer, - and the I/O APIC for device interrupts (see [device-interrupts.md](device-interrupts.md)). + and the I/O APIC for device interrupts (see [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). - **`system/kernel/architecture/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's - machine-readable log channel, see [testing.md](testing.md)) and the shared port-I/O + MSR primitives. + machine-readable log channel, see [testing.md](../testing.md)) and the shared port-I/O + MSR primitives. - **`system/kernel/architecture/x86_64/smp.zig`** / **`per-cpu.zig`** — application-processor bring-up and per-CPU state (GS base, system-call entry point, see [scheduling.md](scheduling.md)). - **`system/kernel/architecture/x86_64/isr.s`** — the exception stubs, the `lgdt`/`lidt`/`ltr` load diff --git a/docs/arm.md b/docs/os-development-guide/arm.md similarity index 97% rename from docs/arm.md rename to docs/os-development-guide/arm.md index 797e7ca..80432a4 100644 --- a/docs/arm.md +++ b/docs/os-development-guide/arm.md @@ -87,7 +87,7 @@ The Pi is not a "standard" ARM platform — expect Broadcom-specific peripherals Pi 5. Everything below is an offset from it. - **UART**: a **PL011** (at base + `0x20_1000`) plus a mini-UART; on some boards the PL011 is wired to Bluetooth, so which one is the console varies. This is the - `aarch64`/`arm` equivalent of our x86 [COM1 serial](testing.md). + `aarch64`/`arm` equivalent of our x86 [COM1 serial](../testing.md). - **Interrupt controller**: *not* a standard ARM GIC on the older parts — the Zero W and Pi 3 use Broadcom's own ARMCTRL controller (Pi 3 adds a per-core "local" controller for timers/mailboxes). The **Pi 4 and 5 do have a GIC-400**. So the @@ -122,4 +122,4 @@ Two routes, mirroring how we test x86-64 with OVMF: - [architecture.md](architecture.md) — the arch-module boundary these targets plug into, and the CPU-arch vs boot-protocol "two axes". - [efi.md](efi.md) — the UEFI loader that carries over to aarch64-UEFI. -- [vision.md](vision.md) — why isolated, portable-across-architectures is the goal. +- [vision.md](../vision.md) — why isolated, portable-across-architectures is the goal. diff --git a/docs/discovery.md b/docs/os-development-guide/discovery.md similarity index 93% rename from docs/discovery.md rename to docs/os-development-guide/discovery.md index 6c612fc..312142a 100644 --- a/docs/discovery.md +++ b/docs/os-development-guide/discovery.md @@ -36,7 +36,7 @@ Two things drive the need, and they set the timing: **aarch64 it's required to boot at all**. The [aarch64 port](arm.md) is what forces the issue. -2. **Isolated user-space drivers need it.** In the [microkernel vision](vision.md), +2. **Isolated user-space drivers need it.** In the [microkernel vision](../vision.md), drivers live in user space — but something has to enumerate the hardware and hand each driver its MMIO regions and IRQs. That enumeration *is* device discovery. So discovery is a prerequisite for real drivers, **not** for user mode itself. @@ -132,7 +132,7 @@ when*: - **User-space enumeration: a device-manager server.** Everything else — PCI devices, peripherals — is parsed (or queried from the kernel's parse) by a privileged user-space server that hands each driver process its MMIO regions and IRQ rights - over [IPC](ipc.md). Combined with **interrupts-as-messages** (an IRQ delivered to a + over [IPC](../device-driver-development-guide/ipc.md). Combined with **interrupts-as-messages** (an IRQ delivered to a driver as a message on a channel — a natural extension of the wait queues and channels already built), that's what makes drivers genuinely isolated. @@ -144,7 +144,7 @@ slice is unavoidably in-kernel. On ARMv8 the generic timer exposes its frequency directly via the `CNTFRQ` register — no calibration needed. That's cleaner than the x86 side, where we measure the LAPIC and TSC against the PIT because nothing tells us their frequency (see -[device-interrupts.md](device-interrupts.md)). Discovery on ARM hands you more for +[device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). Discovery on ARM hands you more for free; discovery on x86 is partly about *finding* what ARM just tells you. ## Suggested ordering @@ -166,18 +166,18 @@ free; discovery on x86 is partly about *finding* what ARM just tells you. - [arm.md](arm.md) — the aarch64 target that forces genuine discovery (DTB, GIC). - [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam, and the note about grabbing the RSDP before `ExitBootServices`. -- [device-interrupts.md](device-interrupts.md) — the LAPIC/timer bring-up that +- [device-interrupts.md](../device-driver-development-guide/device-interrupts.md) — the LAPIC/timer bring-up that discovery will eventually feed (IOAPIC, real IRQ routing). -- [ipc.md](ipc.md) — the channels that interrupts-as-messages and the device manager +- [ipc.md](../device-driver-development-guide/ipc.md) — the channels that interrupts-as-messages and the device manager will ride on. -- [vision.md](vision.md) — why drivers belong in isolated user space at all. +- [vision.md](../vision.md) — why drivers belong in isolated user space at all. ## Update (M19.3, 2026-07-13): PCI enumeration left the kernel The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO apertures derived from the memory map's holes, bus range, and the 16-bit I/O window). The per-function walk moved to the ring-3 `pci-bus` driver -([device-manager.md](device-manager.md)): it claims the bridge, repeats the +([device-manager.md](../device-driver-development-guide/device-manager.md)): it claims the bridge, repeats the ECAM scan through its mmio grant, and `device_register`s what it finds, which the device manager mirrors and matches. The ACPI namespace walk follows in M20; the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side. @@ -190,7 +190,7 @@ for the host bridge, FADT); at this point it also still built the AML namespace but only to read the `\\_S5` sleep type for poweroff. (That remnant is gone too: the kernel now runs no AML at all — soft-off belongs to the acpi service, and the kernel keeps only the AML-free reboot path.) Device discovery is the ring-3 **acpi -service** ([device-manager.md](device-manager.md)): it claims the `acpi-tables` +service** ([device-manager.md](../device-driver-development-guide/device-manager.md)): it claims the `acpi-tables` node the kernel publishes (the AML blobs, a broad io_port grant, the SCI), re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`, and registers + reports each `_HID` device — the device manager matches drivers @@ -206,7 +206,7 @@ ring 0.) Moving PCI and ACPI enumeration out of ring 0 was not just a relocation — it made discovery **firmware-neutral by construction**, which is the whole reason to do it before the second architecture rather than after. Everything at and -above the [device-manager](device-manager.md) protocol — descriptors, +above the [device-manager](../device-driver-development-guide/device-manager.md) protocol — descriptors, containment, reports, matching, supervision — is generic and may never become x86-specific. Discovery is the single firmware-specific piece, and it is isolated as **one swappable process per firmware**: diff --git a/docs/efi.md b/docs/os-development-guide/efi.md similarity index 99% rename from docs/efi.md rename to docs/os-development-guide/efi.md index 8b2c91f..41217e2 100644 --- a/docs/efi.md +++ b/docs/os-development-guide/efi.md @@ -24,7 +24,7 @@ EFI/BOOT/BOOTX64.efi <- the "removable media" default for x86-64 ``` The boot volume is **FHS-shaped** (see the repository-layout note in -[README.md](README.md)): `build.zig` installs `boot/efi.zig` (built for the `uefi` +[README.md](../README.md)): `build.zig` installs `boot/efi.zig` (built for the `uefi` target) at `EFI/BOOT/BOOTX64.efi` — the one path UEFI firmware fixes — and lays the rest out by FHS path: the kernel at `system/kernel`, init at `system/services/init`, the pre-packed boot capsule at `boot/system.img` diff --git a/docs/frame-allocator.md b/docs/os-development-guide/frame-allocator.md similarity index 100% rename from docs/frame-allocator.md rename to docs/os-development-guide/frame-allocator.md diff --git a/docs/framebuffer.md b/docs/os-development-guide/framebuffer.md similarity index 100% rename from docs/framebuffer.md rename to docs/os-development-guide/framebuffer.md diff --git a/docs/gop.md b/docs/os-development-guide/gop.md similarity index 100% rename from docs/gop.md rename to docs/os-development-guide/gop.md diff --git a/docs/halting.md b/docs/os-development-guide/halting.md similarity index 100% rename from docs/halting.md rename to docs/os-development-guide/halting.md diff --git a/docs/heap.md b/docs/os-development-guide/heap.md similarity index 97% rename from docs/heap.md rename to docs/os-development-guide/heap.md index 03c541d..747174e 100644 --- a/docs/heap.md +++ b/docs/os-development-guide/heap.md @@ -51,7 +51,7 @@ rest — works directly on the kernel heap, no bespoke containers required. ## Verifying it -The `heap` test (see [testing.md](testing.md)) exercises the allocator end to end: +The `heap` test (see [testing.md](../testing.md)) exercises the allocator end to end: ``` [PASS] alloc 4096 bytes diff --git a/docs/interrupts.md b/docs/os-development-guide/interrupts.md similarity index 97% rename from docs/interrupts.md rename to docs/os-development-guide/interrupts.md index ece391f..4a4fb3e 100644 --- a/docs/interrupts.md +++ b/docs/os-development-guide/interrupts.md @@ -133,10 +133,10 @@ TSS/IST is wired up: the handler survived a completely broken stack. Both items originally deferred here have landed: -- **The IO-APIC**: [ioapic.zig](../system/kernel/architecture/x86_64/ioapic.zig) +- **The IO-APIC**: [ioapic.zig](../../system/kernel/architecture/x86_64/ioapic.zig) routes external device lines onto vectors — discovered via ACPI's MADT, every input masked at init, lines unmasked one at a time as user-space drivers bind - them (see [device-interrupts.md](device-interrupts.md)). The keyboard followed + them (see [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). The keyboard followed exactly as predicted: the PS/2 bus driver (`system/drivers/ps2-bus/`) claims the 8042 controller and binds its IRQ 1 (and the aux mouse's IRQ 12) through this routing. USB HID keyboards arrive over xHCI instead, which interrupts via diff --git a/docs/logging.md b/docs/os-development-guide/logging.md similarity index 100% rename from docs/logging.md rename to docs/os-development-guide/logging.md diff --git a/docs/memory-map.md b/docs/os-development-guide/memory-map.md similarity index 100% rename from docs/memory-map.md rename to docs/os-development-guide/memory-map.md diff --git a/docs/paging.md b/docs/os-development-guide/paging.md similarity index 96% rename from docs/paging.md rename to docs/os-development-guide/paging.md index efc2fe8..601b902 100644 --- a/docs/paging.md +++ b/docs/os-development-guide/paging.md @@ -112,7 +112,7 @@ kernel heap will build on to map pages on demand. ## Verifying it -Four tests (see [testing.md](testing.md)) pin down the guarantees: +Four tests (see [testing.md](../testing.md)) pin down the guarantees: - **`vmm`** — map a fresh frame at an unused virtual address, write and read it back. Proves `map` works end to end. @@ -138,8 +138,8 @@ Four tests (see [testing.md](testing.md)) pin down the guarantees: processes own the low half. - **Per-address-space tables** — done: each user process gets its own root with the kernel half shared, and refcounted shared-memory mappings exist - ([ipc.md](ipc.md)). Copy-on-write remains unbuilt — nothing has needed it yet. + ([ipc.md](../device-driver-development-guide/ipc.md)). Copy-on-write remains unbuilt — nothing has needed it yet. - **Uncacheable MMIO** — half done: user-space device and DMA mappings are strong-uncacheable and the framebuffer is write-combining via the PAT, but the kernel's own `mapMmio` path is still writeback — the LAPIC included (see - [device-interrupts.md](device-interrupts.md)). + [device-interrupts.md](../device-driver-development-guide/device-interrupts.md)). diff --git a/docs/power.md b/docs/os-development-guide/power.md similarity index 96% rename from docs/power.md rename to docs/os-development-guide/power.md index 282f75d..fb94f29 100644 --- a/docs/power.md +++ b/docs/os-development-guide/power.md @@ -7,7 +7,7 @@ them owns the hardware that reported the event, and the reporter should not know who is listening. So system power is a **service**: an event source **publishes** button/lid/battery/AC events, interested processes **subscribe**, and one privileged caller — init — can ask it to power the machine off. It is the same -publish/subscribe shape as the [input service](input.md), applied to power. +publish/subscribe shape as the [input service](../device-driver-development-guide/input.md), applied to power. ## Why a service, and why it is named for the domain, not the firmware @@ -27,7 +27,7 @@ unchanged. ## The protocol -The `power-protocol` module ([library/protocol/power/power-protocol.zig](../library/protocol/power/power-protocol.zig)) +The `power-protocol` module ([library/protocol/power/power-protocol.zig](../../library/protocol/power/power-protocol.zig)) follows the vfs-protocol pattern — extern-struct messages, a version, reserved fields. Three operations: @@ -126,5 +126,5 @@ until laptop sleep), and thermal zones. firmware neutrality that makes a PSCI backend drop-in on ARM. - [process-lifecycle.md](process-lifecycle.md) — the stop sequence (`terminate → deadline → kill`) and signals init composes into shutdown. -- [device-manager.md](device-manager.md) — the supervision model init mirrors for +- [device-manager.md](../device-driver-development-guide/device-manager.md) — the supervision model init mirrors for its own children. diff --git a/docs/process-lifecycle.md b/docs/os-development-guide/process-lifecycle.md similarity index 98% rename from docs/process-lifecycle.md rename to docs/os-development-guide/process-lifecycle.md index 90fa8b9..9d43f45 100644 --- a/docs/process-lifecycle.md +++ b/docs/os-development-guide/process-lifecycle.md @@ -9,7 +9,7 @@ the layer above them — the standard vocabulary a danos process speaks about it life, and the stable `process` interface that carries it. Nothing here is device- or driver-specific: a driver, the VFS, and a user application all stop, reload, and die the same way. The device manager is simply this design's first -serious customer ([device-manager.md](device-manager.md)). +serious customer ([device-manager.md](../device-driver-development-guide/device-manager.md)). **"POSIX" in this document means the concepts, never the letter of the standard.** danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE @@ -19,7 +19,7 @@ rule is danos's own and it is strict: plain words that communicate intent (`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for a concept that already has one. Literal POSIX arrives later and lives elsewhere: the `std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C -layer** on the same native surface (see [zig-self-hosting.md](zig-self-hosting.md)) — +layer** on the same native surface (see [zig-self-hosting.md](../zig-self-hosting.md)) — musl's syscall surface retargeted at danos system calls and IPC protocols (files onto the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever networking becomes). Ported programs see POSIX; the system underneath never does. @@ -172,7 +172,7 @@ zombie state or privileged snooping: the server's reply with `-EPEER`; a server that dies fails its waiting clients the same way. This covers the *synchronous* case only. 3. **The subscribers** — the new piece, and it is the input service's - publish/subscribe shape ([input.md](input.md)) applied to exits. A stateful + publish/subscribe shape ([input.md](../device-driver-development-guide/input.md)) applied to exits. A stateful service accumulates per-client state across many requests: a filesystem server (FAT today) holds a dead client's open file handles, the input service holds its subscriptions, a future network stack holds its sockets. None of these @@ -320,7 +320,7 @@ get POSIX; danos-native programs never pay for it. `process` grows the interface above; the service harness handles `terminate` and answers the common `ping`; `stop()` for supervisors. -[device-manager.md](device-manager.md) builds directly on all four. +[device-manager.md](../device-driver-development-guide/device-manager.md) builds directly on all four. ## Settled questions (2026-07-12) diff --git a/docs/process-management.md b/docs/os-development-guide/process-management.md similarity index 100% rename from docs/process-management.md rename to docs/os-development-guide/process-management.md diff --git a/docs/release-iso.md b/docs/os-development-guide/release-iso.md similarity index 95% rename from docs/release-iso.md rename to docs/os-development-guide/release-iso.md index 8019cd7..76e6742 100644 --- a/docs/release-iso.md +++ b/docs/os-development-guide/release-iso.md @@ -44,7 +44,7 @@ tables of contents, both pointing at the same embedded FAT image: same `BOOTX64.efi` off it. Neither path involves the legacy BIOS boot-sector machinery: danos is -UEFI-only ([system-requirements.md](system-requirements.md)), so the MBR holds +UEFI-only ([system-requirements.md](../system-requirements.md)), so the MBR holds no boot code, just the partition entry, and the El Torito entry is EFI-class, not floppy emulation. @@ -57,7 +57,7 @@ allocated from the front) either way. The USB path has no such cap. ## The builder `tools/make-iso-image.py` follows the house rule of -[make-fat-image.py](../tools/make-fat-image.py): pure Python 3 standard +[make-fat-image.py](../../tools/make-fat-image.py): pure Python 3 standard library, no external tools (no xorriso, mkisofs, or isohybrid), with a `--verify` mode the `check-iso-image` step runs — it checks that the MBR partition and the El Torito catalog agree on where the FAT image lives and diff --git a/docs/resilience.md b/docs/os-development-guide/resilience.md similarity index 93% rename from docs/resilience.md rename to docs/os-development-guide/resilience.md index 2b05e41..383445d 100644 --- a/docs/resilience.md +++ b/docs/os-development-guide/resilience.md @@ -5,7 +5,7 @@ isolation; fault → kill the process → keep the core (`onException`; the `fault-recovery` test); the supervisor notification **with exit reasons** ([process-lifecycle.md](process-lifecycle.md) — clean exit, fault class, or killed, recorded before the notice posts); and the **restart policy itself** -([device-manager.md](device-manager.md)): the device manager supervises every +([device-manager.md](../device-driver-development-guide/device-manager.md)): the device manager supervises every driver, restarts crashes with backoff, caps crash loops, and re-claims work because the kernel releases a dead process's claims. The `driver-restart` and `usb-report` scenarios prove kill → release → respawn → re-claim → re-report @@ -14,9 +14,9 @@ more of the system moved into restartable processes (the discovery migration, [discovery.md](discovery.md), is the next rung). This is the property danos is really chasing: **if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.** A crashed driver gets restarted; a wedged service gets killed and brought back. It's -the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal +the reason the [microkernel](../vision.md) shape was chosen, and it's a *separate* goal from [real-time](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience) — -one that's less pervasive to build (see [vision.md](vision.md)). +one that's less pervasive to build (see [vision.md](../vision.md)). ## The idea: "let it crash" + supervision @@ -44,7 +44,7 @@ down. **Keeping the kernel minimal is a resilience strategy, not just an aesthet ## The building blocks 1. **Address-space isolation.** A fault in one component can't corrupt another or the - kernel. This is the [user-mode milestone](vision.md) (ring 3, per-process page + kernel. This is the [user-mode milestone](../vision.md) (ring 3, per-process page tables) — the shared prerequisite for *any* of this, and it's needed regardless. 2. **Fault detection** — how the system notices a component is dead or sick: - **Crash**: a CPU fault in a user process (page fault, illegal instruction) traps @@ -81,7 +81,7 @@ Detecting and killing is the easy half. The genuinely tricky questions are about - **In-flight IPC**: messages sent to the dead component, or replies its clients are blocked waiting for. The channel has to break cleanly and unblock the waiters with an error rather than hang them forever (a design constraint that reaches back into - [ipc.md](ipc.md) — channels need a "peer died" outcome). + [ipc.md](../device-driver-development-guide/ipc.md) — channels need a "peer died" outcome). - **Clients**: how does a client discover the service it was talking to is gone and has been replaced? Options: capability revocation makes stale handles fail; or a **name server** re-binds clients to the new instance; or clients retry through a @@ -137,7 +137,7 @@ Honest boundaries: Resilience needs **structural** features (isolation + supervision + a resource model); real-time needs a **pervasive** timing invariant. They're separable, and -resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](vision.md)). +resilience is the lighter commitment (see [smp.md](smp.md) and [vision.md](../vision.md)). Note the overlap, though: **preemptive scheduling** and **priorities** — already built — serve resilience too (you can preempt and kill a misbehaving component, and run the supervisor at high priority). So danos keeps the useful *mechanisms* of the @@ -156,10 +156,10 @@ real-time work without owing anyone a timing *guarantee*. ## Related -- [vision.md](vision.md) — the goals this serves (learning by doing; resilience over +- [vision.md](../vision.md) — the goals this serves (learning by doing; resilience over hard real-time). - [scheduling.md](scheduling.md) — preemption, which makes runaway components killable. -- [ipc.md](ipc.md) — channels that need a "peer died" outcome for clean restart. +- [ipc.md](../device-driver-development-guide/ipc.md) — channels that need a "peer died" outcome for clean restart. - [interrupts.md](interrupts.md) — fault reporting that user mode turns into "kill and restart" instead of "halt". - [smp.md](smp.md) — the real-time-vs-resilience fork, in the SMP context. diff --git a/docs/scheduling.md b/docs/os-development-guide/scheduling.md similarity index 93% rename from docs/scheduling.md rename to docs/os-development-guide/scheduling.md index 2551f5d..82d7a1a 100644 --- a/docs/scheduling.md +++ b/docs/os-development-guide/scheduling.md @@ -3,7 +3,7 @@ The scheduler turns danos from a linear "boot then halt" kernel into a **running multitasking system**. It's **fixed-priority preemptive**: the highest-priority ready task always runs, and tasks at the same priority take turns. That model is -chosen for [real-time](vision.md) — it's predictable (you can reason about which +chosen for [real-time](../vision.md) — it's predictable (you can reason about which task runs when) and its decisions are O(1), unlike a fair-share scheduler. The scheduler proper (`system/kernel/scheduler.zig`) is generic; the context switch and new-task @@ -37,7 +37,7 @@ down a return address pointing at `task_trampoline` and zeroed callee-saved slot `schedule()` — pick the best task and switch — runs from two places: - **`yield()`** — a task voluntarily gives up the CPU. -- **`tick()`** — the 1000 Hz [timer](device-interrupts.md) preempts the running +- **`tick()`** — the 1000 Hz [timer](../device-driver-development-guide/device-interrupts.md) preempts the running task. This is what lets a task that never yields still share the CPU. The subtlety in mixing them is the **interrupt flag (IF)**. The rule: `switch_context` @@ -97,7 +97,7 @@ marks the task blocked with a wake deadline and switches away. On every tick the timer wakes any task whose deadline has passed (a bounded scan, so it stays deterministic), which makes it ready again; the scheduler then runs it when its priority comes up. `sleep` measures its deadline on the [calibrated -clock](device-interrupts.md), so it's real time. +clock](../device-driver-development-guide/device-interrupts.md), so it's real time. When *every* task is blocked, something still has to run — so there's an **idle task** at the lowest priority that just `hlt`s until the next interrupt (see @@ -111,7 +111,7 @@ The other form of blocking is waiting for an **event** rather than a duration. A the caller on it, `wake(wq)` moves the highest-priority waiter back to ready (preempting if it now outranks the running task). A task links into a wait queue through the same field the ready queues use — it's in exactly one queue at a time. -These are the primitives locks, semaphores and [IPC](ipc.md) are built on. +These are the primitives locks, semaphores and [IPC](../device-driver-development-guide/ipc.md) are built on. Blocking safely needs **composable critical sections**. A blanket `cli`/`sti` pair doesn't nest: an IPC channel that `cli`s and then calls `wait` would have `wait`'s @@ -124,7 +124,7 @@ caller's state. ## Verifying it -Three tests (see [testing.md](testing.md)) prove the guarantees: +Three tests (see [testing.md](../testing.md)) prove the guarantees: - **`sched`** spawns three tasks that busy-loop *without ever yielding*. They all make progress — which can only happen if the timer is **preempting** between them @@ -140,7 +140,7 @@ Three tests (see [testing.md](testing.md)) prove the guarantees: - **Priority inheritance** — still open. Tasks now do block on shared resources (IPC rendezvous, the big kernel lock), and nothing yet bounds priority - inversion — a [real-time](vision.md) requirement. + inversion — a [real-time](../vision.md) requirement. - **Task exit / a reaper** — done. A dying task goes on its core's reap list in a `.reaping` state; the timer tick drains the list, frees the stack back to the heap, and recycles the task-table slot. diff --git a/docs/shared-fate-plan.md b/docs/os-development-guide/shared-fate-plan.md similarity index 99% rename from docs/shared-fate-plan.md rename to docs/os-development-guide/shared-fate-plan.md index c8631f7..0480604 100644 --- a/docs/shared-fate-plan.md +++ b/docs/os-development-guide/shared-fate-plan.md @@ -307,4 +307,4 @@ refcount, and no group-kill special case is needed at all. Then update [threading.md](threading.md) (the shared-fate gap note), [process-lifecycle.md](process-lifecycle.md), [process-management.md](process-management.md), and - [ipc.md](ipc.md)/[drivers.md](drivers.md) mentions. + [ipc.md](../device-driver-development-guide/ipc.md)/[drivers.md](../device-driver-development-guide/drivers.md) mentions. diff --git a/docs/smp.md b/docs/os-development-guide/smp.md similarity index 98% rename from docs/smp.md rename to docs/os-development-guide/smp.md index ebc420f..0dd0da5 100644 --- a/docs/smp.md +++ b/docs/os-development-guide/smp.md @@ -112,7 +112,7 @@ Yes — and this is the branch that matters for danos right now. These pull in different directions, so **picking the primary goal comes before picking the SMP design.** (danos's founding assumption was real-time; that's under -active reconsideration in favour of resilience — see [vision.md](vision.md).) +active reconsideration in favour of resilience — see [vision.md](../vision.md).) ## What this would mean for danos @@ -169,7 +169,7 @@ next lands. the highest-priority ready task; per-core queues are a later optimisation. - **AP wake to long mode** — `architecture.startSecondary` drives INIT–SIPI–SIPI (via the LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit - real mode at a low page and runs the [trampoline](../system/kernel/architecture/x86_64/trampoline.s) + real mode at a low page and runs the [trampoline](../../system/kernel/architecture/x86_64/trampoline.s) up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`, publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`: all four cores report `online`. @@ -270,6 +270,6 @@ next lands. - [scheduling.md](scheduling.md) — the single-core scheduler SMP would extend. - [discovery.md](discovery.md) — enumerating cores is a device-discovery problem. -- [ipc.md](ipc.md) — the message passing cross-core coordination rides on. -- [vision.md](vision.md) — the goals question (real-time vs resilience) this note +- [ipc.md](../device-driver-development-guide/ipc.md) — the message passing cross-core coordination rides on. +- [vision.md](../vision.md) — the goals question (real-time vs resilience) this note keeps bumping into. diff --git a/docs/syscall.md b/docs/os-development-guide/syscall.md similarity index 97% rename from docs/syscall.md rename to docs/os-development-guide/syscall.md index 03096f8..bfa0f89 100644 --- a/docs/syscall.md +++ b/docs/os-development-guide/syscall.md @@ -51,7 +51,7 @@ Everything else---including`read()`,`write()`,`malloc()`, and`fork()`---will run 3. **`Yield()`/`Thread_Ctrl()`** - **What it does:**Allows a thread to voluntarily give up its CPU time slice, or allows a root task to spawn/kill threads. 4. **`ipc_send(endpoint, message_buffer)`(Asynchronous Send)** - - **What it does:**Posts a small payload to an endpoint's bounded queue and returns *without* blocking — no rendezvous, no reply. The receiver picks it up through the same `IPC_ReplyWait`, as a buffered message. It is the async counterpart of `IPC_Call`, for one-to-many broadcasts where a synchronous rendezvous would let one dead or slow receiver hang the sender. The [input service](input.md) — keyboard-event fan-out — is its first user. A full queue drops the oldest message (a buffered message is discrete data, unlike a coalescing interrupt notification). + - **What it does:**Posts a small payload to an endpoint's bounded queue and returns *without* blocking — no rendezvous, no reply. The receiver picks it up through the same `IPC_ReplyWait`, as a buffered message. It is the async counterpart of `IPC_Call`, for one-to-many broadcasts where a synchronous rendezvous would let one dead or slow receiver hang the sender. The [input service](../device-driver-development-guide/input.md) — keyboard-event fan-out — is its first user. A full queue drops the oldest message (a buffered message is discrete data, unlike a coalescing interrupt notification). * * * * * @@ -92,4 +92,4 @@ Managing the Payload Challenge Because it is a microkernel, performance lives or dies by how fast your`IPC_Call`can move data from Client to Server. You have two minimal choices for handling the`message_buffer`pointer:[[1](https://anazimzada2020.medium.com/microkernel-architectural-pattern-5e4e9184170e)] - **The Copy Method (Simplest to start):**Your kernel pauses the client, reads the data from the client's memory space, switches page tables to the server, and copies the data into the server's buffer. -- **The Shared Memory Method (Fastest):**The kernel sets up a temporary, shared virtual memory page between the client and server. The client writes to it, calls`syscall`/`svc`, and the server reads it instantly without the kernel copying any bytes \ No newline at end of file +- **The Shared Memory Method (Fastest):**The kernel sets up a temporary, shared virtual memory page between the client and server. The client writes to it, calls`syscall`/`svc`, and the server reads it instantly without the kernel copying any bytes diff --git a/docs/system-image.md b/docs/os-development-guide/system-image.md similarity index 98% rename from docs/system-image.md rename to docs/os-development-guide/system-image.md index 5d31faf..a429af9 100644 --- a/docs/system-image.md +++ b/docs/os-development-guide/system-image.md @@ -11,7 +11,7 @@ sequential pass and hands the bytes to the kernel unmodified. The capsule is a *performance artifact*, not a source of truth. The boot volume's `/system` and `/test` file trees remain the canonical layout (see -[danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md)); +[danos-file-system-hierarchy-FSH.md](../file-system-development/danos-file-system-hierarchy-FSH.md)); the capsule is a pre-baked snapshot of the same binaries, derived from the same build graph, so the running system is identical whether the loader read the capsule or walked the tree. diff --git a/docs/sysv.md b/docs/os-development-guide/sysv.md similarity index 100% rename from docs/sysv.md rename to docs/os-development-guide/sysv.md diff --git a/docs/threading-plan.md b/docs/os-development-guide/threading-plan.md similarity index 96% rename from docs/threading-plan.md rename to docs/os-development-guide/threading-plan.md index 572f15d..ca62cd9 100644 --- a/docs/threading-plan.md +++ b/docs/os-development-guide/threading-plan.md @@ -2,7 +2,7 @@ The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like -[display-v2-plan.md](display-v2-plan.md). Read threading.md first for the *why*. +[display-v2-plan.md](../device-driver-development-guide/display-v2-plan.md). Read threading.md first for the *why*. ## Locked decisions (do not relitigate) @@ -13,7 +13,7 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, `single_threaded = false`. - **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an idle core still halts ([halting.md](halting.md)). -- **New syscalls are private**: extend [abi.zig](../system/abi.zig) `SystemCall` after +- **New syscalls are private**: extend [abi.zig](../../system/abi.zig) `SystemCall` after `shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`, `futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number. - **Restart granularity stays the process** — a faulting thread kills its process; the @@ -21,10 +21,10 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, ## Conventions -Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations, +Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through `addUserBinary` (with the new `threaded` flag where a binary spawns threads) and get -packed into the initial-ramdisk; new syscalls extend [abi.zig](../system/abi.zig) +packed into the initial-ramdisk; new syscalls extend [abi.zig](../../system/abi.zig) `SystemCall` + a `library/runtime` wrapper; test services live beside the code they exercise and register a `ServiceId` if they must be looked up. @@ -58,7 +58,7 @@ fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration 5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into `main` stays a human step. 3. **Implement** every unchecked item in that milestone, including adding its - `-Dtest-case` to `CASES` in [test/qemu_test.py](../test/qemu_test.py) (with + `-Dtest-case` to `CASES` in [test/qemu_test.py](../../test/qemu_test.py) (with `smp: true` / a `mem` bump where noted) so the gate is runnable. 4. **Run the gate**: `python3 test/qemu_test.py `, then the full **guardrail set**, then `zig build` (clean) and `zig build test` (green). @@ -67,7 +67,7 @@ fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration the whole guardrail set passes, `zig build` is clean, and host tests are green. → tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit` (`threads(M): `, no `Co-Authored-By` trailer per - [coding-standards.md](coding-standards.md)), then **`git push` the working branch to + [coding-standards.md](../coding-standards.md)), then **`git push` the working branch to `origin`** (use `-u` on the first push to set upstream). Continue to the next milestone in the same iteration if budget remains; otherwise let the loop re-fire. - **Red** = anything above fails. Diagnose from the captured serial log @@ -111,7 +111,7 @@ task's exit; make destruction happen on the **last** exit. - [x] A refcount keyed by the address-space root, held in `scheduler.zig` (`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the success path, after the slot + stack are secured), all under the big kernel lock. -- [x] Both task-teardown paths ([scheduler.zig](../system/kernel/scheduler.zig): +- [x] Both task-teardown paths ([scheduler.zig](../../system/kernel/scheduler.zig): `exitUserLocked` and `destroyTaskLocked`) call `releaseAddressSpace`, which decrements and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test spaces) is destroyed directly, preserving prior behaviour. @@ -133,7 +133,7 @@ full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `sm Spawn only — no join yet. Prove a second task executes in the **caller's** address space and exits cleanly. -- [x] [abi.zig](../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in +- [x] [abi.zig](../../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in process.zig; `thread_spawn` calls `scheduler.spawnThread` (today, after M3, the handler goes `spawnThreadSupervised` → `scheduler.spawnUserLocked`; shares the caller's address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)` @@ -195,7 +195,7 @@ plus `address-space-refcount`/`thread-spawn`; `zig build` clean, `zig build test ## M4 — Futex: the one blocking primitive ✅ -- [x] [abi.zig](../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a +- [x] [abi.zig](../../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a `.blocked` task tagged with `Task.futex_addr` (no queue linkage); `futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock, parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr, @@ -261,7 +261,7 @@ green. restores it. - [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same `Futex`/`Mutex`/`Condition` when wanted. -- [x] All `thread-*` cases wired into [test/qemu_test.py](../test/qemu_test.py) +- [x] All `thread-*` cases wired into [test/qemu_test.py](../../test/qemu_test.py) (`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md status updated to **built**; the worked example is threading.md's win-condition. - [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main @@ -405,7 +405,7 @@ guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-re ### M10 — Per-thread TLS: the thread-pointer mechanism ✅ Give each thread its own thread pointer and private TLS storage — the foundation -self-hosting Zig ([zig-self-hosting.md](zig-self-hosting.md)) will build `threadlocal` on. +self-hosting Zig ([zig-self-hosting.md](../zig-self-hosting.md)) will build `threadlocal` on. - [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch **only when it changes** (the same conditional-load discipline as CR3; @@ -459,7 +459,7 @@ clean. ## Deferred (explicitly not in this plan) - **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a - physical-address key so two processes share a futex through a [shared-memory](display-v2.md) + physical-address key so two processes share a futex through a [shared-memory](../device-driver-development-guide/display-v2.md) region. Not needed for intra-process threads. - **Per-thread priorities / affinity distinct from the process** — threads inherit the process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep. @@ -467,5 +467,5 @@ clean. ([process-lifecycle.md](process-lifecycle.md)). - **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more. - **A real `std.Thread` backend** — arrives with self-hosting - ([zig-self-hosting.md](zig-self-hosting.md)); it sits on these same primitives, so it + ([zig-self-hosting.md](../zig-self-hosting.md)); it sits on these same primitives, so it swaps the impl under `runtime.Thread`, not the call sites. diff --git a/docs/threading.md b/docs/os-development-guide/threading.md similarity index 91% rename from docs/threading.md rename to docs/os-development-guide/threading.md index 0bc69a0..a24db90 100644 --- a/docs/threading.md +++ b/docs/os-development-guide/threading.md @@ -2,7 +2,7 @@ A note on danos **threads** — several tasks sharing one address space — provided by a `Thread` type that mirrors the shape of Zig's `std.Thread` while keeping every -kernel entry behind the [runtime](../library/kernel). **Built** (M1–M11, see +kernel entry behind the [runtime](../../library/kernel). **Built** (M1–M11, see [threading-plan.md](threading-plan.md)): `spawn`/`join`/`detach`, cross-core parallelism, a futex, `Mutex`/`Condition`/`Semaphore`/`RwLock`/`WaitGroup`, `getCurrentId`/`currentCore`, per-thread thread-pointer TLS, thread-safe allocation, and a task reaper that reclaims dead @@ -36,12 +36,12 @@ implementation underneath, not the API above. [Why not literal std.Thread](#why-not-literal-stdthread). - **Threads are a narrow, opt-in capability — not the default concurrency tool.** The default for resilience stays **process + IPC** ([resilience.md](resilience.md), - [ipc.md](ipc.md)). See [Where threads fit](#where-threads-fit-the-resilience-tension). + [ipc.md](../device-driver-development-guide/ipc.md)). See [Where threads fit](#where-threads-fit-the-resilience-tension). - **Blocking synchronization is futex-backed, never spin-backed.** Waiters sleep in the kernel so an idle core still halts ([halting.md](halting.md)). - **Per-binary opt-in to multi-threaded codegen.** Only a service that asks for threads is built `single_threaded = false`; the rest stay lean and single-threaded. -- **The thread ABI is private.** New syscalls extend [abi.zig](../system/abi.zig) +- **The thread ABI is private.** New syscalls extend [abi.zig](../../system/abi.zig) `SystemCall` and are reached only through `library/kernel` wrappers, exactly like every other danos syscall ([syscall.md](syscall.md)) — numbers stay renumberable. @@ -56,12 +56,12 @@ runtime — rebuilt in lockstep — knows the mapping. `std.Thread` is incompatible with that invariant on two counts: 1. **It selects its backend from `builtin.os.tag`, and issues syscalls directly.** - danos targets `.os_tag = .freestanding` ([build.zig](../build.zig)), for which + danos targets `.os_tag = .freestanding` ([build.zig](../../build.zig)), for which `std.Thread` resolves to an unsupported stub that `@compileError`s. Adding a real backend would either bake danos syscall numbers into std (breaking ABI privacy and renumbering) or fork std to route back through the runtime — a permanent rebase cost that buys nothing the native type doesn't. -2. **Our user binaries are built `single_threaded = true`** ([build.zig](../build.zig) +2. **Our user binaries are built `single_threaded = true`** ([build.zig](../../build.zig) `addUserBinary`), which compiles threading out entirely and makes atomics and TLS single-threaded. Threads need this flipped per binary regardless. @@ -134,7 +134,7 @@ Deviations from `std.Thread`, called out honestly: ## Kernel primitives (new private syscalls) -Five core entries extend [abi.zig](../system/abi.zig) `SystemCall` after +Five core entries extend [abi.zig](../../system/abi.zig) `SystemCall` after `shared_memory_physical = 36` (plus small helpers `current_core`, `thread_self`, and `set_thread_pointer`), each with a `library/kernel` wrapper: @@ -155,12 +155,12 @@ Plus one invariant change with no new syscall: **address-space reference countin Before this work an address space was 1:1 with a task: `spawnUserLocked` records `address_space` on the Task (as it still does), and teardown did `destroyAddressSpace(t.address_space)` when **any** user task exited -([scheduler.zig](../system/kernel/scheduler.zig)). With threads, several tasks share +([scheduler.zig](../../system/kernel/scheduler.zig)). With threads, several tasks share one `address_space`, so the first to exit would rip the address space out from under its siblings. Fix: a small refcount keyed by the address-space root, kept in -[scheduler.zig](../system/kernel/scheduler.zig): `retainAddressSpace` takes a +[scheduler.zig](../../system/kernel/scheduler.zig): `retainAddressSpace` takes a reference for every user task `spawnUserLocked` starts (count 1 on the first take, so a thread sharing the caller's space increments it); task teardown calls `releaseAddressSpace`, which only calls `destroyAddressSpace` at **zero**. All of @@ -172,7 +172,7 @@ that must land and be proven before anything shares an address space. The scheduler already accepts an arbitrary `address_space` and does **not** smuggle values through scratch registers — `startUserTask` reads the entry/stack (and the thread's closure arg, delivered in `rdi` via `jumpToUserArg`) from the Task -([scheduler.zig](../system/kernel/scheduler.zig)). That makes the thread path clean: +([scheduler.zig](../../system/kernel/scheduler.zig)). That makes the thread path clean: 1. The runtime's `spawn` `mmap`s a stack (syscall `4`) and writes the closure — `{ tls_base, args }`, the std "Instance" pattern — at the **top of the new stack @@ -214,7 +214,7 @@ Keying: threads share an address space, so a **virtual address within that addre identifies a futex uniquely; the kernel keys its wait queue by `(address_space_root, virtual_address)`. Keying by the **physical** address instead (translate `virtual_address -> physical_address` on entry) is a deliberate forward door: it lets two *processes* share a futex through an -[shared-memory](display-v2.md) region later, without changing the API. We start with the +[shared-memory](../device-driver-development-guide/display-v2.md) region later, without changing the API. We start with the private-per-address-space key and note the physical-key upgrade. No spinning: a contended lock parks the task in the kernel and the core is free to run @@ -264,9 +264,9 @@ stays single-threaded and lean. **process**, which respawns its threads from a known-good state — restart granularity stays the process. The leader's recorded exit reason carries the fault class even when a worker faulted, so restart policy is unchanged. -- **IPC — two consequences threads forced ([ipc.md](ipc.md)):** +- **IPC — two consequences threads forced ([ipc.md](../device-driver-development-guide/ipc.md)):** - *Handles do not cross threads.* The handle table lives on the `Task` - ([scheduler.zig](../system/kernel/scheduler.zig)), so a handle number is meaningful + ([scheduler.zig](../../system/kernel/scheduler.zig)), so a handle number is meaningful only to the thread that created it — thread A's endpoint handle `3` is not thread B's. A thread that needs to reach an endpoint another thread owns looks it up (`ipc.lookup(service)`) to install its **own** handle to the same underlying endpoint. @@ -284,7 +284,7 @@ stays single-threaded and lean. The ordered, `/loop`-runnable milestones live in **[threading-plan.md](threading-plan.md)** (shaped like -[display-v2-plan.md](display-v2-plan.md)): every milestone lands on its own and ends in +[display-v2-plan.md](../device-driver-development-guide/display-v2-plan.md)): every milestone lands on its own and ends in a verifiable gate (`python3 test/qemu_test.py `, asserting serial markers; `zig build test` for host unit tests). The stages below are the shape it expands. @@ -307,13 +307,13 @@ a verifiable gate (`python3 test/qemu_test.py `, asserting serial markers; the consumer blocked, e.g. via a low idle tick count). - **Stage 3 — polish.** Per-thread TLS / thread pointer and `threadlocal` (only if a consumer needs it), `RwLock`/`WaitGroup` as demanded, and this doc's cases wired - into [test/qemu_test.py](../test/qemu_test.py). + into [test/qemu_test.py](../../test/qemu_test.py). ## Conventions -Follow [coding-standards.md](coding-standards.md): spell out non-acronym +Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New syscalls -extend [abi.zig](../system/abi.zig) `SystemCall` + a `library/kernel` wrapper +extend [abi.zig](../../system/abi.zig) `SystemCall` + a `library/kernel` wrapper ([syscall.md](syscall.md)). `Thread` is a first-class runtime module, the same way `process` ([process-lifecycle.md](process-lifecycle.md)) and `ipc` are — user code never names a syscall. @@ -331,7 +331,7 @@ are — user code never names a syscall. ## The self-hosting endgame When danos becomes a real Zig target and we (eventually) add a danos backend to std -([zig-self-hosting.md](zig-self-hosting.md)), `std.Thread` can sit *on top of* these +([zig-self-hosting.md](../zig-self-hosting.md)), `std.Thread` can sit *on top of* these same kernel primitives — the danos `std.Thread.Impl` would call the very `thread_spawn`/`futex_*` wrappers `Thread` already uses. Because `Thread` was built API-compatible from day one, that transition swaps the @@ -341,9 +341,9 @@ the later self-hosting lift cheap. ## Further reading - [scheduling.md](scheduling.md), [smp.md](smp.md) — the task model these threads join. -- [resilience.md](resilience.md), [vision.md](vision.md) — why isolation is the default +- [resilience.md](resilience.md), [vision.md](../vision.md) — why isolation is the default and threads are the exception. -- [syscall.md](syscall.md), [ipc.md](ipc.md) — the private ABI and the messaging model +- [syscall.md](syscall.md), [ipc.md](../device-driver-development-guide/ipc.md) — the private ABI and the messaging model threads sit beside. - [halting.md](halting.md) — the idle/halt property futex-backed blocking preserves. -- [zig-self-hosting.md](zig-self-hosting.md) — the target this bends toward. +- [zig-self-hosting.md](../zig-self-hosting.md) — the target this bends toward. diff --git a/docs/timers.md b/docs/os-development-guide/timers.md similarity index 93% rename from docs/timers.md rename to docs/os-development-guide/timers.md index cfeed01..1693944 100644 --- a/docs/timers.md +++ b/docs/os-development-guide/timers.md @@ -7,7 +7,7 @@ Two different needs hide under the word "timer", and danos keeps them apart: Both are answered by the **kernel**, because the kernel already owns a timer: it has to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this -are built in [device-interrupts.md](device-interrupts.md); the scheduler's blocking and +are built in [device-interrupts.md](../device-driver-development-guide/device-interrupts.md); the scheduler's blocking and wait queues are in [scheduling.md](scheduling.md). This page is about the surface a ring-3 program actually uses, and one deliberate absence: **there is no user-space time service.** @@ -32,11 +32,11 @@ danos checks both — the invariant-TSC CPUID bit (`0x80000007` EDX[8], set on I AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET counter when either fails. So `now()` stays accurate on a real Intel box, a real AMD box, and inside a VM alike; only the source behind it differs. The mechanism is in -[device-interrupts.md](device-interrupts.md). +[device-interrupts.md](../device-driver-development-guide/device-interrupts.md). So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver -model; that role now lives in [drivers.md](drivers.md), as documentation.) The one place +model; that role now lives in [drivers.md](../device-driver-development-guide/drivers.md), as documentation.) The one place a user-space time service *is* justified — **wall-clock / calendar time** — is discussed at the end; it is deliberately not built yet. @@ -55,7 +55,7 @@ Time and waiting are three entries in the small syscall table ([syscall.md](sysc service can keep answering messages on the same endpoint while a deadline is pending. This is the timed wait that stop-sequence escalation, hello deadlines, and restart backoff are built from ([process-lifecycle.md](process-lifecycle.md), - [device-manager.md](device-manager.md)). + [device-manager.md](../device-driver-development-guide/device-manager.md)). The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space; programs read the TSC through `clock` and get timed wakeups through `sleep`/`timer_bind`, diff --git a/docs/vdso.md b/docs/os-development-guide/vdso.md similarity index 99% rename from docs/vdso.md rename to docs/os-development-guide/vdso.md index 3c4241d..21b6cbe 100644 --- a/docs/vdso.md +++ b/docs/os-development-guide/vdso.md @@ -50,7 +50,7 @@ The public danos ABI then has exactly two layers, neither of which is | Layer | Contract | Spoken by | |-------|----------|-----------| | **vDSO** | C-ABI functions, this note | every language's thin shim (the `system-call` module for Zig, a `-sys` crate for Rust, a header for C) | -| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](vfs-protocol.md) is the first one documented) | any client that can lay out bytes | +| **IPC wire protocols** | byte layouts over `ipc_call` ([vfs-protocol.md](../file-system-development/vfs-protocol.md) is the first one documented) | any client that can lay out bytes | Everything above those — the heap, `file_system`, the service harness — is per-language convenience, compiled into each binary from source, exactly as diff --git a/docs/system-requirements.md b/docs/system-requirements.md index a01adf6..b3e3628 100644 --- a/docs/system-requirements.md +++ b/docs/system-requirements.md @@ -118,7 +118,7 @@ hypervisor configured for UEFI firmware and an xHCI USB controller. (`system/kernel/acpi.zig:3`) - The loader reads `/system/kernel` off the FAT boot volume, then loads user space: a prebuilt `boot\system.img` capsule - ([system-image.md](system-image.md)) when present, otherwise it walks + ([system-image.md](os-development-guide/system-image.md)) when present, otherwise it walks the volume's `/system` and optional `/test` trees (init included) into the initial ramdisk. The kernel can boot "kernel-only" without either. (`efi.zig:16`, `efi.zig:68`) diff --git a/docs/testing.md b/docs/testing.md index 134a902..9e0f721 100644 --- a/docs/testing.md +++ b/docs/testing.md @@ -27,7 +27,7 @@ boot log, memory summary, exception reports — appears on serial as plain text. QEMU captures that with `-serial file:serial.log`, giving a machine-readable transcript. Serial is per-architecture (x86 uses port I/O; an ARM board uses a -memory-mapped UART), so it lives behind the [architecture](architecture.md) boundary — and adding +memory-mapped UART), so it lives behind the [architecture](os-development-guide/architecture.md) boundary — and adding a new architecture's UART is what makes the same tests run there. The serial log sink is **compiled in only under `-Dserial`** (off by default). @@ -88,7 +88,7 @@ table in `test/qemu_test.py`): | `fault-recovery` | a ring-3 process that faults is killed and reaped while init keeps heartbeating — the OS survives | `DANOS-TEST-RESULT: PASS` | The faulting cases don't print a result line — they deliberately raise a CPU -exception, and the harness asserts on the [exception report](interrupts.md) the +exception, and the harness asserts on the [exception report](os-development-guide/interrupts.md) the handler prints (which also reaches serial). This reuses the real fault path as the test oracle: if the IDT/TSS weren't wired up, `fault-df` would triple-fault and the marker would never appear. diff --git a/docs/vision.md b/docs/vision.md deleted file mode 100644 index 98ba86b..0000000 --- a/docs/vision.md +++ /dev/null @@ -1,118 +0,0 @@ -# Vision: a microkernel, built to learn - -danos exists first and foremost as a **learning-by-doing project**: the point is to -build a real operating system, bump into the hard constraints for real, and research -them from a position of having actually hit them. The docs in this folder are part of -that — they're where a constraint gets understood once it's been met. - -That framing sets the priorities. danos is not chasing a spec or a product; it's -chasing understanding, with a concrete, motivating **win condition** to aim at. - -## The win condition - -danos is a "win" when it: - -- **boots and runs on real hardware** — the author's **PC** (x86-64) and **both - Raspberry Pis**: the **Zero 2 W** and the **Pi 5** (both `aarch64`, one backend — - see [arm.md](arm.md)), -- **has a graphical user interface**, ideally — building on the framebuffer it - already draws to. - -Everything below serves that, or serves the curiosity that the project runs on. - -## Why a microkernel: resilience - -The kernel stays **minimal** — only what genuinely must run privileged: - -- scheduling, -- inter-process communication (IPC), -- memory management (address spaces, page tables), -- low-level interrupt dispatch. - -Everything else — device drivers, filesystems, the GUI, the network stack — runs as -an **isolated user-space server**, each in its own address space with only the -privileges it needs. - -The reason for this shape is **resilience**: the ability to **re-initialise parts of -the OS while it runs**. A driver bug can't corrupt the kernel or another driver; a -crashed or wedged component is contained, killed, and **restarted** — "if I break -something, I can just fix it," without rebooting. Keeping the kernel tiny is part of -that strategy: the one thing that *can't* be restarted is the trusted base, so the -less code in it, the less that can take the whole system down. This is the project's -real motivation, and it has its own design note: [resilience.md](resilience.md). - -The cost is that **IPC becomes the backbone**: what used to be a function call inside -a monolithic kernel is now a message between address spaces. In a microkernel, IPC -performance essentially *is* system performance (the lesson of L4), so it's a -first-class concern. Hardware interrupts become IPC too: the kernel turns an IRQ into -a message to the driver that owns the device. - -## On real-time: an option, not a commitment - -danos was originally framed as a hard **real-time** OS. That's now held as **one -interesting constraint to explore, not a requirement** — because real-time is a -*pervasive* invariant (every operation must be provably time-bounded, everywhere) -that would slow every milestone, whereas resilience is a set of *structural* features -that's lighter to build and is what the project actually wants. The trade-off is -written up in [smp.md](smp.md#does-the-right-choice-depend-on-real-time-vs-resilience). - -What danos keeps from the real-time direction, because it's cheap and useful anyway: - -- **Fixed-priority preemptive scheduling** — the highest-priority ready task runs, and - preemption lets a runaway component be interrupted and killed (which *serves - resilience*). Already built ([scheduling.md](scheduling.md)). -- **A calibrated, deterministic clock** — already built ([device-interrupts.md](device-interrupts.md)). - -What danos does *not* owe anyone unless it deliberately chooses real-time later: -timing *guarantees*, priority inheritance, bounded allocators, tickless timers, MCS -scheduling contexts. Concretely, the current [heap](heap.md) is a first-fit free list -with unbounded allocation time — fine here, and only a problem *if* a hard-real-time -path is ever added. Note that **QNX is both** a real-time and a restartable -microkernel, so choosing resilience now doesn't close the real-time door — it just -doesn't pay the tax yet. - -## The roadmap — tracks, not a strict line - -Because the driver is curiosity plus the win condition, the roadmap is a set of -**tracks** with dependencies, not a rigid sequence. Pick by interest; mind the -prerequisites. - -**Done:** UEFI boot, framebuffer + [serial](testing.md), [physical frames](frame-allocator.md) -(with boot-services memory reclaimed), [paging](paging.md) with W^X, [exceptions and -interrupts](interrupts.md), a [calibrated timer + ns clock](device-interrupts.md), a -[heap](heap.md), a [fixed-priority preemptive scheduler](scheduling.md) with blocking, -in-kernel [IPC channels](ipc.md), SMP (all cores scheduling, with affinity), a -**higher-half kernel** with a physmap, and **user space**: per-process address -spaces, `syscall`/`sysret` with the `swapgs` discipline, a user-ELF loader, and -`/system/services/init` — a real user ELF built from `system/services/init/`, running at CPL 3 as PID 1 on its -own page tables — plus a [test harness](testing.md). - -- **Isolation track** — **user mode + address-space isolation**. *Done: a - higher-half kernel with a physmap (the low half is user space), per-process - address spaces with CR3 switched on context switch, the `swapgs` discipline, - `syscall`/`sysret`, a user-ELF loader, an address-space/stack reaper for exited - tasks, and `/system/services/init` running as a real preemptive ring-3 process - (PID 1). Remaining polish: SMAP + fault-recovering copy-in/out, and TLB shootdown - once a process has more than one thread. (The real IPC syscalls — - `ipc_call`/`ipc_reply_wait` — have since been built and are the backbone every - driver and service speaks; see [ipc.md](ipc.md).)* -- **Resilience track** — fault → kill → notify, a supervisor/reincarnation server, - resource cleanup on death, then a restartable driver as proof. Needs isolation. - See [resilience.md](resilience.md). -- **ARM track** — the `aarch64` port so danos runs on the Zero 2 W and Pi 5. Largely - independent of the others (it's the [architecture layer](architecture.md)); directly serves the win - condition. Likely via aarch64-UEFI first (QEMU `virt` + AAVMF), then real boards. - See [arm.md](arm.md), and [discovery.md](discovery.md) for the device tree it needs. -- **GUI track** — a framebuffer-based windowing/compositor, and the input + display - drivers under it. Builds on the neutral framebuffer (so it's arch-independent), and - on the driver model from the isolation/resilience tracks. The visible payoff. - -The natural spine is **isolation → (resilience + drivers) → GUI**, with the **ARM -track** pursued alongside whenever the itch to see it boot on a Pi wins out. - -## How to use this page - -Read it before adding anything structural. When a design decision comes up, the -question is: does it serve the **win condition** (runs on the three machines, with a -GUI), or the **learning** (a constraint worth meeting)? If it serves neither — e.g. -paying the full real-time tax with no payoff in sight — it can wait. diff --git a/docs/zig-self-hosting.md b/docs/zig-self-hosting.md index 339759f..0b56604 100644 --- a/docs/zig-self-hosting.md +++ b/docs/zig-self-hosting.md @@ -108,7 +108,7 @@ localised (below). ## The architecture decision: `runtime.os` + `runtime.fs`, and retire `posix` danos already has the right split ([the private-ABI boundary](../README.md)): the -kernel exposes a minimal syscall ABI ([syscall.md](syscall.md)); the **`runtime`** +kernel exposes a minimal syscall ABI ([syscall.md](os-development-guide/syscall.md)); the **`runtime`** library is the stable, danos-native application ABI. What this roadmap adds: - **`runtime.os` — the seam.** A C-ABI-shaped module of the ~30 operations @@ -165,7 +165,7 @@ What the seam needs, and what danos already provides: | mmap / munmap | native syscalls ([abi.zig](../system/abi.zig)) | none | | page allocator | over `mmap`, via `root.os.heap.page_allocator` override | ~30-line hook | | monotonic clock | `clock` syscall | none | -| args / argv | SysV entry stack ([sysv.md](sysv.md)), `runtime.process.Init` | none | +| args / argv | SysV entry stack ([sysv.md](os-development-guide/sysv.md)), `runtime.process.Init` | none | | stdout / stderr | `debug_write` today | wire fd 1/2 to a console **byte** stream | | mkdir / unlink / rename / truncate | done — engine + VFS + `runtime.fs` (Phase 2) | — | | stat fields | `{size, kind, mtime}` | **mode / inode** still missing (cache validity) | @@ -209,7 +209,7 @@ build); point danos's `build.zig`/CI at the resulting binary. Four localised pat plan9/serenity; - add `danos` to the freestanding/other **no-op `_start` list** in `std`'s `start.zig`, so std does *not* emit its own System-V `_start` — danos keeps owning the entry shim - and `Init`/argv construction it already builds ([sysv.md](sysv.md)); + and `Init`/argv construction it already builds ([sysv.md](os-development-guide/sysv.md)); - wire the `system` selector `.danos => std.os.danos` in `std.posix`; - add `std/os/danos.zig` — **the seam itself**, promoted near-verbatim from the `runtime.os` developed first in Phase 1 (against the stock toolchain, so the fork is @@ -344,10 +344,10 @@ Two current decisions fall out of this roadmap: ## Related - [vision.md](vision.md) — the north star this serves. -- [syscall.md](syscall.md) — the kernel↔runtime ABI `runtime.os` is built on. -- [sysv.md](sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs. -- [ipc.md](ipc.md) — the IPC the VFS/FAT operations travel over. -- [danos-file-system-hierarchy-FSH.md](danos-file-system-hierarchy-FSH.md) — the +- [syscall.md](os-development-guide/syscall.md) — the kernel↔runtime ABI `runtime.os` is built on. +- [sysv.md](os-development-guide/sysv.md) — the entry stack (`argc/argv/envp/auxv`) danos already constructs. +- [ipc.md](device-driver-development-guide/ipc.md) — the IPC the VFS/FAT operations travel over. +- [danos-file-system-hierarchy-FSH.md](file-system-development/danos-file-system-hierarchy-FSH.md) — the filesystem layout the file surface serves. - [coding-standards.md](coding-standards.md) — danos naming (why the compat spellings are confined, and now retired). diff --git a/library/xkeyboard-config/README.md b/library/xkeyboard-config/README.md index afc087b..c458f29 100644 --- a/library/xkeyboard-config/README.md +++ b/library/xkeyboard-config/README.md @@ -1,6 +1,6 @@ # xkeyboard-config — X11 keyboard layouts, compiled to Zig -This module turns a physical key (a **USB HID usage**, as the [input module](../../docs/input.md) +This module turns a physical key (a **USB HID usage**, as the [input module](../../docs/device-driver-development-guide/input.md) delivers in `KeyEvent.keycode`) plus a modifier state into a **keysym** and, when the key produces one, a **character** (a Unicode scalar). It is what lets a `keycode` become a `character` — a keymap — without danos shipping an X11 runtime.