Moonshot AI launches Kimi K3
A mixture-of-experts design activating 104 billion of its 2.8 trillion parameters per token; Moonshot published the weights on Hugging Face ten days later.
- Open weights & ecosystem
- Models & capabilities
- Major
Moonshot AI launched Kimi K3, describing it as a 2.8-trillion-parameter mixture-of-experts model and the company’s largest to date. The model activates roughly 104 billion parameters per token across 896 experts, of which 16 are active at once, and supports a context window of up to 1 million tokens. Moonshot said the design achieved a 2.5-times improvement in scaling efficiency over its predecessor, Kimi K2, and reported strong scores on reasoning, coding and agentic-task benchmarks, including GPQA Diamond and BrowseComp. The model runs by default with an always-on “thinking mode,” with configurable low, high or maximum reasoning effort, and preserves its reasoning content across multi-turn conversations rather than discarding it after each response.
The launch followed a now-familiar pattern from Chinese labs: a capability announcement ahead of the full open-weight release, which Moonshot published on Hugging Face roughly ten days later, ahead of its own stated target date. Coverage described K3 as among the largest open-weight models released to date, and noted it scored competitively against Western frontier systems including Anthropic’s Claude Fable 5 on coding benchmarks. Commentators framed the release as continued evidence that Chinese labs were matching or exceeding US frontier capability on several benchmarks despite US export controls on advanced chips, by building larger, more efficient models rather than simply more compute.
The release also fit into a broader pattern that week of benchmark claims becoming harder to compare directly: K3 was among several systems reported to have scored full marks at the International Mathematical Olympiad days later, in results that were self-administered and graded by another AI model rather than official competition judges.