12 models · on the curve since 2023
DeepSeek
DeepSeek lives at the sharp end of the output layer: open weights, frontier-adjacent quality, trained for a fraction of the cost. The durable asset is the efficiency itself, the ability to keep matching the frontier cheaply enough to reset everyone’s price expectations.
What would change this readThe cost advantage closes as rivals copy the efficiency tricks, or open weights stop tracking the closed frontier.
- Recent cadence
- 5 in last 12 mo
- 3 the year before
- Landmark density
- 4 / 12
- defining releases
- Weights
- Open weights
- DeepSeek-V4 flagship
Trains and serves on
Disclosed only for the previous generation: DeepSeek-V3 was trained on NVIDIA H800 GPUs, per its own technical report. For V4 there is no company statement at all. A Huawei-led team said in June 2026 it had completed full-parameter POST-training of V4-Pro on about 1,000 Ascend 910C chips — post-training, not pre-training, and not confirmed by DeepSeek. Ascend, Cambricon, Hygon and Moore Threads all completed day-zero inference adaptation at the V4 launch.
ConcentrationFrontier training silicon needs clearance from two governments at once — a US export licence for the NVIDIA part and a Chinese import approval for the same part — while Beijing directs routine inference onto domestic accelerators.
Flagship, per million tokens
$0.435 in $0.87 out
deepseek-v4-pro, cache-miss input
Was
$0.14 / $0.28
deepseek-chat, the V2-era flagship · Aug 2024
- Weights
- Full, under MIT — the most permissive licence of any lab on this page. V4-Pro, V4-Flash, R1 and the V3 line are all MIT; models before 2025 used a proprietary DeepSeek licence.
- Where you buy it
- Own API · HuggingFace · OpenRouter · Together AI · Fireworks · AWS Bedrock (R1, V3.1, V3.2)
- Structure
- Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Private and unlisted, grown out of the hedge fund High-Flyer, with founder Liang Wenfeng holding the large majority. The National AI Industry Investment Fund is a direct shareholder of the core entity at about 0.28%.
- Raised
- About RMB 50B (~$7.4B) in its first outside round, closed by July 2026, from Tencent, CATL, NetEase, JD.com and the National AI Industry Investment Fund
- Valuation
- About US$51.8B post-money (Jul 2026), implied by a regulatory filing rather than announced
- Revenue
- Not disclosed. Estimated at $400-500M annualised (The Information, 2026), largely API sales
Biggest single dependencyIts frontier training hardware requires simultaneous US export authorisation and Chinese state approval, and the domestic alternative has publicly demonstrated only post-training and inference at V4 scale — not pre-training.
What V4 was pre-trained on is genuinely unknown; the Huawei Ascend claim covers post-training only. Widely circulated figures of 2.788M H800 GPU-hours and a $5.5M training cost belong to V3 and are frequently misattributed to V4. The pricing page carries an unscheduled note that peak-hour rates will double, so the figures above can change without notice. Revenue is a single-source estimate, not a filing.
Hover any point to read the release. The y-axis is qualitative; only the x-axis carries dates.
- Cheap reasoning
- Reach for R1 when you want open reasoning at a fraction of closed-model cost (Jan 2025).
- Open frontier-adjacent
- Pick V3 for near-frontier quality you can host yourself (671B MoE, Dec 2024).
- Cost-sensitive scale
- Use DeepSeek where token economics decide the build (V2, the price-war trigger, May 2024).
- Apr 2026 DeepSeek-V4 (Pro/Flash) Million-token context preview
- Jan 2025 DeepSeek-R1 Open reasoning model — the global market shock
- Dec 2024 DeepSeek-V3 671B MoE trained at remarkably low cost
- May 2024 DeepSeek-V2 Cheap MoE that triggered China’s price war
Full release logall 12, newest first
- 2026
- Jul DeepSeek-V4-Flash Public beta of the Flash tier — a retrain of the April preview, not a new architecture. V4-Pro still pending
- Apr DeepSeek-V4 (Pro/Flash) ◆ Million-token context preview
- 2025
- Dec DeepSeek-V3.2 Plus a Speciale variant
- Sep DeepSeek-V3.2-Exp Sparse-attention experiment
- Aug DeepSeek-V3.1 Hybrid reasoning update
- May DeepSeek-R1-0528 Major R1 upgrade
- Jan DeepSeek-R1 ◆ Open reasoning model — the global market shock
- 2024
- Dec DeepSeek-V3 ◆ 671B MoE trained at remarkably low cost
- Jun DeepSeek-Coder-V2 Strong open coding model
- May DeepSeek-V2 ◆ Cheap MoE that triggered China’s price war
- 2023
- Nov DeepSeek LLM 67B base and chat models
- Nov DeepSeek Coder First model, code-specialised
- Founded
- 2023
- HQ
- Hangzhou, China
- Leadership
- Liang Wenfeng
- Parent
- High-Flyer
- Flagship
- DeepSeek-V4
- Weights
- Open (MIT-style)
- Also builds
- DeepSeek app
- Known for
- Low-cost open models
Founded July 2023 by Liang Wenfeng, spun out of the Chinese quant fund High-Flyer. Headquartered in Hangzhou, it is known for highly cost-efficient open-weight reasoning models — the R1 release in January 2025 triggered a global reaction over China’s frontier capabilities.