FlashMLA

by deepseek-ai · Developer Tools

FlashMLA: Efficient Multi-head Latent Attention Kernels

New0 ratings12,905 starsActive
Developer ToolsDeveloper ToolC++Cuda
GitHub⚑ Report
FlashMLA — image 1

About this project

FlashMLA Introduction FlashMLA is DeepSeek's library of optimized attention kernels, powering the DeepSeek-V3 and DeepSeek-V3.2-Exp models. This repository contains the following implementations: Sparse Attention Kernels These kernels power DeepSeek Sparse Attention (DSA), as introduced in this paper. — Token-level sparse attention for the prefill stage — Token-level sparse attention for the decoding stage, with FP8 KV cache Dense Attention Kernels — Dense attention for the prefill stage — Dense attention for the decoding stage News — 2025.09.29 Release of Sparse Attention Kernels: With the launch of DeepSeek-V3.2, we are releasing the corresponding token-level sparse attention…

Technologies

PythonC++CCuda

Project health

Recently updated
Last updatelast month
Contributors16
Latest release
Open issues & PRs127
LicenseMIT
On GitHubsince 2025

GitHub

12,905
stars
1,144
forks
16
contributors
127
open issues & PRs
C++
language
last month
last commit
View on GitHub ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review FlashMLA.

Built by

deepseek-ai
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain deepseek-ai/FlashMLA? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like