quickthink is a local-first CLI and Python library that wraps Ollama-backed LLM calls with a compressed plan-then-answer scaffold and latency-aware routing. It adds a short validated planning step for multi-step prompts and routes simple ones straight through. Local inference control for small models.
python cli latency routing scaffold inference structured-output local-first ai-tools llm prompt-engineering small-models local-llm llm-inference ollama llm-ops small-llm ai-reliability hermes-labs plan-then-answer
-
Updated
Sep 16, 2026 - Python