DeepSeek-V3
Paper Link ποΈ Table of Contents 1. Introduction 2. Model Summary 3. Model Downloads 4. Evaluation Results 5. Chat
About this project
Paper Link ποΈ Table of Contents 1. Introduction 2. Model Summary 3. Model Downloads 4. Evaluation Results 5. Chat Website & API Platform 6. How to Run Locally 7. License 8. Citation 9. Contact 1. Introduction We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for strongerβ¦
Technologies
Project health
GitHub
Reviews
Built by
Maintain deepseek-ai/DeepSeek-V3? Claiming verifies admin access through your GitHub account and gives you control of this listing.