Run LLMs larger than your RAM — native GGUF inference engine with SSD streaming, no GPU required
-
Updated
Apr 2, 2026 - C
Run LLMs larger than your RAM — native GGUF inference engine with SSD streaming, no GPU required
Self-contained C11 CPU inference runtime for GGUF language models, with deterministic decoding, hardened loading, and zero third-party dependencies.
To associate your repository with the wayos topic, visit your repo's landing page and select "manage topics."