A Flagship Sheds Its Preview Label
DeepSeek took its flagship V4 Pro model out of preview and into general availability on 12 August 2026, closing a test period that had run since 24 April. The build carries the version tag 0813 and follows the smaller V4-Flash variant, which made the same move to official status on 31 July. Both models now sit inside DeepSeek's production API rather than a preview tier that developers were warned could change without notice.
V4 Pro keeps the one-million-token context window it carried through preview, with a maximum output of 384,000 tokens per request. DeepSeek built the model as a mixture-of-experts system with 1.6 trillion total parameters, of which 49 billion activate for any single token, the same architecture pattern the company has used since V3 to keep serving costs down while scaling total capacity.
The Efficiency Case Behind 1.6 Trillion Parameters
The headline number is not the parameter count but what DeepSeek did to serve it cheaply. A new attention mechanism the company calls Compressed Sparse Attention, paired with what it labels Heavily Compressed Attention, cuts the compute needed for each token of inference to 27 percent of what the prior V3.2 model required, and shrinks the memory-heavy key-value cache to 10 percent of the earlier footprint.
On benchmarks, V4 Pro resolves 80.6 percent of SWE-bench Verified tasks, matching Google's Gemini-3.1-Pro on the same test, and scores a 90.1 percent pass rate on GPQA Diamond and 93.5 percent on LiveCodeBench. It trails OpenAI's GPT-5.4 on Terminal Bench 2.0, scoring 67.9 against 75.1, and posts a Codeforces rating of 3,206, evidence that DeepSeek is now benchmarked against, not around, the leading American labs.
Ninety-Eight Percent Cheaper, and Not for Long
At general availability, V4 Pro is priced at 0.435 dollars per million input tokens on a cache miss, 0.003625 dollars per million on a cache hit, and 0.87 dollars per million output tokens, unchanged from its preview rate. The cheaper V4-Flash tier runs 0.14 dollars per million input tokens and 0.28 dollars for output. Research firm Artificial Analysis measured V4-Flash's effective cost at about 0.03 dollars per benchmark task, against 1.86 dollars for OpenAI's GPT-5.6 Sol and 3.15 dollars for Anthropic's Claude Fable 5, a gap of more than 98 percent.
In the same window as the V4 Pro release, DeepSeek's developer documentation page carried a separate notice: overall API pricing is set to rise significantly in the near future, though the company gave no date and no percentage. Developers were told to plan their usage accordingly, with the final pricing structure to be announced separately, an unusual admission that today's price sheet is not the one customers should be budgeting against.
The Line Item That Just Turned Variable
The timing matters more than the number. DeepSeek is warning of a repricing in the same release cycle that just gave V4 Pro production status, which tells procurement teams that the model's reliability graduated before its cost did. A concurrency cap of 500 simultaneous requests on the Pro endpoint, a fifth of the 2,500 ceiling on the cheaper Flash tier, is a second signal pointing the same way: capacity on the expensive, higher-quality tier is the constraint DeepSeek is managing first, and price is the lever it has left.
Enterprises that shifted workloads toward Chinese models specifically to escape unpredictable American frontier pricing now hold the same exposure from the other direction, with no floor and no calendar date attached to it. As an illustration only, not a forecast: a rise from 0.87 to even 1.50 dollars per million output tokens on a workload running billions of tokens a month turns a rounding error into a seven-figure annual swing for a US-based buyer, and DeepSeek has given no assurance the increase stops there. Treating a subsidized price sheet as a fixed line in a multi-year budget was always a bet that a discounted market would hold still; the vendor has now said, in its own documentation, that it will not.
Read next: The Model Changed but the API Name Did Not | Beijing Business Hours Now Price Your AI



