macOS AI agent resource hog watcher that sacrifices expendable GUI apps before OOM kills agent work
-
Updated
Jul 3, 2026 - Python
macOS AI agent resource hog watcher that sacrifices expendable GUI apps before OOM kills agent work
Comparing KV-cache-aware scheduling policies for LLM serving: greedy preempts 49% of requests under memory pressure while memory-first achieves 43% higher effective throughput with zero preemptions.
Step-level simulation of chunked prefill under KV cache pressure. Block policy wastes 89.6% of prefill chunks while achieving the same throughput as upfront rejection. Chunk waste rate is the metric standard throughput metrics miss.
To associate your repository with the memory-pressure topic, visit your repo's landing page and select "manage topics."