DeepSeek-V2

by deepseek-ai · Developer Tools

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

New0 ratings5,038 starsActive
Developer ToolsDeveloper Tool
GitHub⚑ Report
DeepSeek-V2 — image 1
DeepSeek-V2 — image 2
DeepSeek-V2 — image 3
DeepSeek-V2 — image 4
DeepSeek-V2 — image 5
DeepSeek-V2 — image 6
1 / 6

About this project

Model Download Evaluation Results Model Architecture API Platform License Citation Paper Link 👁️ DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 1. Introduction Today, we’re introducing DeepSeek-V2, a strong Mixture-of-Experts (MoE) language model characterized by economical training and efficient inference. It comprises 236B total parameters, of which 21B are activated for each token. Compared with DeepSeek 67B, DeepSeek-V2 achieves stronger performance, and meanwhile saves 42.5% of training costs, reduces the KV cache by 93.3%, and boosts the maximum generation throughput to 5.76 times. We pretrained DeepSeek-V2 on a…

Project health

Inactive for over a year
Last updatelast year
Contributors7
Latest release
Open issues & PRs89
LicenseMIT
On GitHubsince 2024

GitHub

5,038
stars
550
forks
7
contributors
89
open issues & PRs
language
last year
last commit
View on GitHub ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review DeepSeek-V2.

Built by

deepseek-ai
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain deepseek-ai/DeepSeek-V2? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like