Event-driven benchmark of speculative prefill during user reading time in multi-turn LLM conversations, measuring net TTFT benefit, contention penalty, and the conditions under which speculation becomes net-negative.
pythonperformance-engineeringbenchmarksimulationlatencyschedulerinferencesystemsgatingservingmulti-turnprefillevent-driven-simulationcontentionkv-cachettftllmspeculative-prefill
-
Updated
Jul 24, 2026 - Python