12 models · on the curve since 2023

DeepSeek

Value position · the site’s read Output layer

DeepSeek lives at the sharp end of the output layer: open weights, frontier-adjacent quality, trained for a fraction of the cost. The durable asset is the efficiency itself, the ability to keep matching the frontier cheaply enough to reset everyone’s price expectations.

What would change this readThe cost advantage closes as rivals copy the efficiency tricks, or open weights stop tracking the closed frontier.

Recent cadence
5 in last 12 mo
3 the year before
Landmark density
4 / 12
defining releases
Weights
Open weights
DeepSeek-V4 flagship

Founded 2023Hangzhou, China

The substrate underneath figures as of Aug 2026

Trains and serves on

Disclosed only for the previous generation: DeepSeek-V3 was trained on NVIDIA H800 GPUs, per its own technical report. For V4 there is no company statement at all. A Huawei-led team said in June 2026 it had completed full-parameter POST-training of V4-Pro on about 1,000 Ascend 910C chips — post-training, not pre-training, and not confirmed by DeepSeek. Ascend, Cambricon, Hygon and Moore Threads all completed day-zero inference adaptation at the V4 launch.

ConcentrationFrontier training silicon needs clearance from two governments at once — a US export licence for the NVIDIA part and a Chinese import approval for the same part — while Beijing directs routine inference onto domestic accelerators.

Flagship, per million tokens

$0.435 in $0.87 out

deepseek-v4-pro, cache-miss input

Was

$0.14 / $0.28

deepseek-chat, the V2-era flagship · Aug 2024

Weights
Full, under MIT — the most permissive licence of any lab on this page. V4-Pro, V4-Flash, R1 and the V3 line are all MIT; models before 2025 used a proprietary DeepSeek licence.
Where you buy it
Own API · HuggingFace · OpenRouter · Together AI · Fireworks · AWS Bedrock (R1, V3.1, V3.2)
Structure
Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Private and unlisted, grown out of the hedge fund High-Flyer, with founder Liang Wenfeng holding the large majority. The National AI Industry Investment Fund is a direct shareholder of the core entity at about 0.28%.
Raised
About RMB 50B (~$7.4B) in its first outside round, closed by July 2026, from Tencent, CATL, NetEase, JD.com and the National AI Industry Investment Fund
Valuation
About US$51.8B post-money (Jul 2026), implied by a regulatory filing rather than announced
Revenue
Not disclosed. Estimated at $400-500M annualised (The Information, 2026), largely API sales

Biggest single dependencyIts frontier training hardware requires simultaneous US export authorisation and Chinese state approval, and the domestic alternative has publicly demonstrated only post-training and inference at V4 scale — not pre-training.

What V4 was pre-trained on is genuinely unknown; the Huawei Ascend claim covers post-training only. Widely circulated figures of 2.788M H800 GPU-hours and a $5.5M training cost belong to V3 and are frequently misattributed to V4. The pricing page carries an unscheduled note that peak-hour rates will double, so the figures above can change without notice. Revenue is a single-source estimate, not a filing.

On the durability curve frontier era
2023202420252026 DeepSeek CoderDeepSeek-V4-Flash ARCHITECTURE the substrate · compounds  → OUTPUT per model · commoditises  → CROSSOVER the durability thesis  → DURABLE VALUE-CAPTURE ↑ the margin a release keeps · not raw capability TIME · FRONTIER MODEL RELEASES →
Output · value a model keeps Architecture · the compounding substrate DeepSeek release

Hover any point to read the release. The y-axis is qualitative; only the x-axis carries dates.

When to reach for them practical read
Cheap reasoning
Reach for R1 when you want open reasoning at a fraction of closed-model cost (Jan 2025).
Open frontier-adjacent
Pick V3 for near-frontier quality you can host yourself (671B MoE, Dec 2024).
Cost-sensitive scale
Use DeepSeek where token economics decide the build (V2, the price-war trigger, May 2024).
The landmark releases 4 of 12 models
  1. Apr 2026 DeepSeek-V4 (Pro/Flash) Million-token context preview Open1M ctxFlagship
  2. Jan 2025 DeepSeek-R1 Open reasoning model — the global market shock OpenReasoningMarket shock
  3. Dec 2024 DeepSeek-V3 671B MoE trained at remarkably low cost Open671B MoE
  4. May 2024 DeepSeek-V2 Cheap MoE that triggered China’s price war OpenMoEPrice war
Full release logall 12, newest first
  1. 2026
  2. Jul DeepSeek-V4-Flash Public beta of the Flash tier — a retrain of the April preview, not a new architecture. V4-Pro still pending OpenRetrainCurrent
  3. Apr DeepSeek-V4 (Pro/Flash) Million-token context preview Open1M ctxFlagship
  4. 2025
  5. Dec DeepSeek-V3.2 Plus a Speciale variant Open
  6. Sep DeepSeek-V3.2-Exp Sparse-attention experiment OpenSparse attn
  7. Aug DeepSeek-V3.1 Hybrid reasoning update OpenHybrid
  8. May DeepSeek-R1-0528 Major R1 upgrade OpenReasoning
  9. Jan DeepSeek-R1 Open reasoning model — the global market shock OpenReasoningMarket shock
  10. 2024
  11. Dec DeepSeek-V3 671B MoE trained at remarkably low cost Open671B MoE
  12. Jun DeepSeek-Coder-V2 Strong open coding model OpenCoding
  13. May DeepSeek-V2 Cheap MoE that triggered China’s price war OpenMoEPrice war
  14. 2023
  15. Nov DeepSeek LLM 67B base and chat models Open67B
  16. Nov DeepSeek Coder First model, code-specialised OpenCoding
Where they sit on weights versus the field
The durability take on DeepSeek from the archive
About DeepSeek at a glance
Founded
2023
HQ
Hangzhou, China
Leadership
Liang Wenfeng
Parent
High-Flyer
Flagship
DeepSeek-V4
Weights
Open (MIT-style)
Also builds
DeepSeek app
Known for
Low-cost open models

Founded July 2023 by Liang Wenfeng, spun out of the Chinese quant fund High-Flyer. Headquartered in Hangzhou, it is known for highly cost-efficient open-weight reasoning models — the R1 release in January 2025 triggered a global reaction over China’s frontier capabilities.

The frontier chart → Read the archive →