#
mooncake
Here are 4 public repositories matching this topic...
Discrete-event simulation of LLM request routing across multiple serving instances: round-robin, least-load, prefix-aware, and hybrid cache-aware routing with queue-depth threshold sweep.
pythonsimulationlatencyinferenceload-balancingmulti-instancerequest-routingmooncakekv-cachellm-servingserving-infrastructureprefix-cachings-loramlsystems
-
Updated
Jul 19, 2026 - Python
KVCache Forge 的 CCF 2026 Mooncake 赛题工程日志、设计决策、实验记录与决赛路线图。
-
Updated
Aug 6, 2026 - Python
Self-contained PoC: cross-tenant timing oracle against shared content-hashed KV caches (Mooncake, Perplexity KV Messenger, DeepSeek 3FS).
-
Updated
May 17, 2026 - Python
Improve this page
Add a description, image, and links to the mooncake topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the mooncake topic, visit your repo's landing page and select "manage topics."