Timeline

DeepSeek releases DeepSeek-R1-0528 update

Released under an MIT licence, the update raised AIME 2025 accuracy from 70% to 87.5% by roughly doubling the average length of the model's reasoning traces.

  • Open weights & ecosystem
  • Models & capabilities
  • Benchmarks & progress
  • Notable

DeepSeek released DeepSeek-R1-0528, an update to its reasoning model that the company described as an improvement in benchmark performance, hallucination rate and front-end capabilities such as JSON output and function calling, with no change required to existing API usage. Weights were published on Hugging Face under an MIT licence, permitting commercial use and distillation, consistent with the open release of the original R1.

The reported gains were substantial for a point release: accuracy on AIME 2025 rose from 70.0% to 87.5%, GPQA Diamond from 71.5% to 81.0%, LiveCodeBench from 63.5% to 73.3%, and HMMT 2025 from 41.7% to 79.4%. DeepSeek attributed the improvement to greater “thinking depth” during inference — the model’s average reasoning trace on AIME questions roughly doubled, from about 12,000 to 23,000 tokens — rather than to any change in the underlying architecture or training data.

The update narrowed the gap to closed frontier reasoning models on several widely watched benchmarks: its AIME score put it close to OpenAI’s o3, and its GPQA score was within a few points of the same model. Combined with the open, commercially usable licence, it reinforced the pattern set by the original R1 release in January 2025 — that open-weight Chinese models could track closed Western frontier systems on reasoning benchmarks within months rather than years, at a fraction of the subscription cost of proprietary alternatives.