Adaptive local inference for Hermes — TurboFit Check, evidence-built TurboFit List, repaired discovery/benchmark campaigns, Qwen 3.8 DFlash2, Bonsai DSpark, FreeToken, and Sirvir.
-
Updated
Aug 24, 2026 - Python
Adaptive local inference for Hermes — TurboFit Check, evidence-built TurboFit List, repaired discovery/benchmark campaigns, Qwen 3.8 DFlash2, Bonsai DSpark, FreeToken, and Sirvir.
Local-first adaptive LLM runtime for Hermes Agent. Matches 8–300 GB hardware, acquires pinned GGUFs through Turbohaul, and safely contracts and heals under VRAM pressure—no remote fallback by default.
Adaptive AI inference runtime built on llama.cpp with intelligent scheduling, self-optimizing execution, runtime learning, and advanced memory management.
Add a description, image, and links to the adaptive-runtime topic page so that developers can more easily learn about it.
To associate your repository with the adaptive-runtime topic, visit your repo's landing page and select "manage topics."