DeepSeek releases V4.1-Flash and folds V4-Pro traffic into it
The open-weight model activates only about 8 billion of 552 billion parameters per token; from 14 September DeepSeek routes all V4-Pro API traffic to it at Flash's lower price.
- Models & capabilities
- Open weights & ecosystem
- Notable
DeepSeek released DeepSeek-V4.1-Flash, an open-weight mixture-of-experts model with 552 billion total parameters that activates only about 8 billion per token during prefill and 16 billion during decode, published under the MIT licence on Hugging Face. The model supports a context window of up to one million tokens and adds native multimodal, image-understanding capability. DeepSeek said tests by multiple parties put V4.1-Flash ahead of its own V4-Pro flagship on performance, cost, speed and total runtime, and priced it well below V4-Pro’s rates — Techstrong.ai reported cuts of up to 32% from DeepSeek’s own August prices, reversing part of the peak-pricing increase it had introduced that month.
From 04:00 UTC on 14 September 2026, DeepSeek said it would automatically route all API requests naming deepseek-v4-pro to V4.1-Flash instead, billed at Flash’s cheaper rate, until a successor V4.1-Pro model ships — in effect retiring V4-Pro as a distinct served model rather than issuing a conventional side-by-side upgrade.
Techstrong.ai reported that shares of rival Chinese AI firms MiniMax and Z.ai fell more than 8% in Hong Kong trading the same day, with Alibaba down more than 2%, extending a pattern in which DeepSeek’s price cuts have repeatedly moved competitors’ stock. The release also came as DeepSeek was reported to be preparing an initial public offering on Shanghai’s STAR Market, continuing the company’s practice, established with V3 and R1, of releasing frontier-class open weights while competing aggressively on price.