DeepSeek publishes DeepSeekMath, introducing GRPO
The 7B model reached 51.7% on the MATH benchmark without external tools, and its GRPO training method later underpinned DeepSeek-R1's reasoning training.
DeepSeekIdeas & essays · Models & capabilities