Timeline

DeepSeek releases V3

DeepSeek's technical report put the final training run at 2.79 million H800 GPU-hours, or about $5.6 million at an assumed $2-per-hour rental rate.

  • Open weights & ecosystem
  • Models & capabilities
  • Major

DeepSeek released the weights of DeepSeek-V3, a 671-billion-parameter mixture-of-experts language model that activates 37 billion parameters per token, and published an accompanying technical report the following day. The company said it was trained on 14.8 trillion tokens and reported benchmark results it described as comparable to leading closed-weight models, while stating the model ran at roughly three times the inference throughput of its predecessor, DeepSeek-V2.

The technical report’s training-cost disclosure drew the most attention. DeepSeek said the final pre-training run used 2.788 million GPU-hours on Nvidia H800 chips — a China-market variant with reduced interconnect bandwidth, built to comply with US export controls — plus a further 119,000 GPU-hours for context-length extension and 5,000 for post-training. Assuming a rental price of $2 per GPU-hour, the company put total training cost at roughly $5.6 million, explicitly excluding the cost of prior research, ablation experiments and architecture development that had gone into reaching that configuration. The figure was strikingly low against the hundreds of millions of dollars typically associated with frontier-model training runs at US labs, though it measured only the final run and not DeepSeek’s cumulative research spending.

The company released the weights under a permissive licence allowing commercial use and self-hosting, continuing its practice, alongside firms such as Meta and Alibaba, of publishing capable open-weight models rather than serving them only through a closed API.

At release, DeepSeek-V3 drew comparatively little Western coverage — most contemporaneous attention focused on its benchmark scores and low reported cost as a data point in the debate over the efficiency of Chinese AI development under export controls. That changed within weeks: the same underlying architecture, extended with reasoning training into DeepSeek-R1 and released in January 2025, triggered a sharp sell-off in AI-related stocks after US investors concluded that a Chinese lab could match frontier performance at a fraction of the assumed compute cost, and the training-cost figures from this report became the most quoted numbers in that episode.