Model · DeepSeek · Text
DeepSeek V4 Pro
DeepSeek's flagship general-purpose MoE model. Broadly competitive with the strongest proprietary models, released under MIT at open-weights cost.
- Modality
- Text
- License
- MIT (Open)
- Parameter size
- 1.6T total, 49B activated
- Context window
- 1,048,576 tokens 1M-token context at 27% of DeepSeek V3.2's per-token FLOPs and 10% of its KV cache.
- Released
- August 13, 2026
- Last verified
- September 8, 2026
- Runs locally
- Yes
Strengths
- MIT licence — no commercial restrictions on the weights
- Broadly competitive with the strongest proprietary models (Terminal Bench 2.1 87.9 vs Claude Opus 4.8's 85.0; Cybergym 83.3 vs 78.3)
- 1M-token context at a fraction of the previous generation's compute and memory cost
- Gains are most pronounced in production environments
Weaknesses
- 1.6T total parameters puts self-hosting out of reach for almost everyone
- Trails on the hardest reasoning benchmarks (HLE 42.7 vs Claude Fable 5's 53.3 and Opus 4.8's 49.8)
- Trails Claude Opus 4.8 on repo-scale code work (NL2Repo 61.5 vs 69.7)
Try it
| Where | Type | Notes |
|---|---|---|
| Hugging Face | weights | MIT licence |
| DeepSeek Platform | hosted-api | API key required |
Used in solutions
Version history
- DeepSeek V4 Pro Aug 2026 Current
- DeepSeek V4 Flash Jul 2026
- DeepSeek V3 Dec 2024 Deprecated
Official sources
- Model card model
- Technical report paper
Change log
- — Corrected on three counts. Release date 2026-05-06 was a Hugging Face push timestamp, not a release — the GA DeepSeek-V4-Pro-0813 shipped 2026-08-13 per DeepSeek's API changelog, superseding the preview this entry described. Licence corrected from 'DeepSeek License' to MIT, and the false 'restrictions on commercial use cases' weakness removed. Added parameter size and 1M context window.
- — Initial entry. Sighted on Hugging Face (lastModified 2026-05-06).
Esc