FlashMLA
FlashMLA: Efficient Multi-head Latent Attention Kernels
About this project
FlashMLA Introduction FlashMLA is DeepSeek's library of optimized attention kernels, powering the DeepSeek-V3 and DeepSeek-V3.2-Exp models. This repository contains the following implementations: Sparse Attention Kernels These kernels power DeepSeek Sparse Attention (DSA), as introduced in this paper. — Token-level sparse attention for the prefill stage — Token-level sparse attention for the decoding stage, with FP8 KV cache Dense Attention Kernels — Dense attention for the prefill stage — Dense attention for the decoding stage News — 2025.09.29 Release of Sparse Attention Kernels: With the launch of DeepSeek-V3.2, we are releasing the corresponding token-level sparse attention…
Technologies
Project health
GitHub
Reviews
Built by
Maintain deepseek-ai/FlashMLA? Claiming verifies admin access through your GitHub account and gives you control of this listing.