Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
About this project
A KVCache-centric Disaggregated Architecture for LLM Serving Paper Slides Traces Documentation Blog Slack as a rollout data-transfer backend for the fragmented, heterogeneous data moving between rollout and training in disaggregated RL. Blogs: [KVCache.AI, lmsys. — Aug 17, 2026: Mooncake is integrated into Speculators as a distributed backend for multi-node online training, efficiently moves hidden-state between vLLM inference workers and Speculators trainers through RDMA, eliminating the need for massive hidden-state storage in offline training. Benchmark on GB300 NVL72. — May 7, 2026: 🚀 vLLM officially features Mooncake Store — a deep dive into how…
Technologies
Project health
GitHub
Reviews
Built by
Maintain kvcache-ai/Mooncake? Claiming verifies admin access through your GitHub account and gives you control of this listing.