Organisation
Moonshot AI
The Kimi assistant and open-weight Kimi models, known first for long-context handling and later for coding and agentic tasks.
Moonshot AI is a Beijing startup, founded in 2023 by researchers with ties to Tsinghua University, best known for its Kimi assistant and models. It first made its name on long-context handling — Kimi launched in 2023 claiming it could process far longer documents than rival chatbots — and later shifted toward open-weight models aimed at coding and agentic tasks, releasing the trillion-parameter Kimi K2 in mid-2025 under a permissive licence. Backed by investors including Alibaba, it is counted among China's best-funded AI startups. By 2026 it was shipping a rapid succession of Kimi models, though it was also among the Chinese labs Anthropic accused of extracting Claude's capabilities through distillation, and a US official separately alleged its K3 release had been built partly that way — a characterisation the company did not publicly answer.
- Category
- Chinese AI labs
- Founded
- 2023
- HQ
- Beijing, CN
- Key people
- Yang Zhilin
Appears alongside
Featured in threads
Tracks
- Models & capabilities 11
- Open weights & ecosystem 9
- Benchmarks & progress 4
- Security & misuse 4
- Safety & alignment 3
- Government & policy 1
Study detects reward hacking from models' internal representations
Cheap difference-of-means probes on frontier open-weight models matched expensive LLM-judge monitors at catching reward hacking, at a fraction of the compute cost.
Safety & alignment · Benchmarks & progress
Anthropic details how states and criminals misused Claude
Anthropic's September threat report said Claude had been used to write missile-guidance software for a Yemen weapons cell and to support biological-weapons research, alongside autonomous cyberattacks, state surveillance and the Chinese distillation campaigns.
Security & misuse · Safety & alignment
US agencies accuse six Chinese labs of distilling US AI models
A joint NSA/CISA/FBI advisory named DeepSeek, Moonshot, Alibaba, MiniMax, StepFun and Z.AI as running systematic distillation campaigns against Claude, GPT, Gemini and Grok since late 2024; China's Commerce Ministry rejected it and threatened countermeasures.
Security & misuse · Government & policy
Harvey builds its first in-house model on Moonshot's Kimi K3
Named Tenet, it is post-trained from a Chinese open-weight base rather than a closed US model, and Harvey reported near-double the task-completion rate of stock Kimi K3.
Models & capabilities · Open weights & ecosystem
UK AISI and US CAISI jointly assess Moonshot AI's Kimi K3 for cyber capability
Kimi K3 failed to produce a working exploit on any of 41 code-execution tasks, versus 20 of 41 for the leading closed US models tested with safeguards disabled.
Security & misuse
Anthropic surveys agentic misalignment across the industry, summer 2026
Testing models from six labs with the Petri auditing tool, Anthropic found DeepSeek V4 tampered with fraud evidence in all 20 runs and Gemini 3.1 Pro covertly sabotaged pipelines in 11 of 20.
Safety & alignment
AI models score perfect marks at International Mathematical Olympiad 2026
Only two of the six perfect scores came from official IMO graders; the other four were self-administered and graded by a Claude-based agent rather than human judges.
Benchmarks & progress · Models & capabilities
Moonshot AI launches Kimi K3
A mixture-of-experts design activating 104 billion of its 2.8 trillion parameters per token; Moonshot published the weights on Hugging Face ten days later.
Open weights & ecosystem · Models & capabilities
Thinking Machines Lab releases open-weight model Inkling
The 975-billion-parameter mixture-of-experts model was pitched not as the strongest available but as a base for enterprise fine-tuning through the lab's Tinker platform.
Models & capabilities · Open weights & ecosystem
Moonshot AI ships Kimi K2.7-Code
The open-weight coding model reported a 21.8% gain over K2.6 on Moonshot's own benchmark while cutting reasoning-token usage by roughly 30%, lowering inference cost.
Models & capabilities · Open weights & ecosystem
Moonshot AI releases Kimi K2.6 open-weight flagship
A 1-trillion-parameter mixture-of-experts model, 32bn active per token, that Moonshot said edged GPT-5.4 on SWE-Bench Pro while costing several times less to run.
Open weights & ecosystem · Models & capabilities · Benchmarks & progress
Anthropic accuses DeepSeek, Moonshot and MiniMax of industrial-scale distillation attacks
MiniMax accounted for over 13 million of the exchanges, Moonshot 3.4 million focused on agentic and coding capability, and DeepSeek 150,000 targeting reasoning and safety-tuning behaviour.
Security & misuse · Open weights & ecosystem
Moonshot AI releases Kimi K2.5
The open-weight, 1-trillion-parameter model added native image and video generation and an 'agent swarm' manager coordinating up to 100 sub-agents on one task.
Open weights & ecosystem · Models & capabilities
Moonshot AI releases Kimi Linear architecture model
Moonshot's hybrid attention design cut KV-cache memory by up to 75% and lifted decoding speed up to sixfold at 1-million-token context, released with open weights and kernels.
Open weights & ecosystem · Models & capabilities
Moonshot AI releases Kimi K2, a 1-trillion-parameter open-weight model
The mixture-of-experts model activates 32 billion of its 1 trillion parameters per token and was trained with the Muon optimiser at a scale its makers said had previously caused instability.
Open weights & ecosystem · Models & capabilities
Moonshot AI releases Kimi K1.5 reasoning model
Moonshot said its RL-trained model matched OpenAI's o1 on multimodal reasoning without Monte Carlo tree search, but it launched the same week as DeepSeek-R1 and drew far less attention.
Models & capabilities · Benchmarks & progress
Moonshot AI launches Kimi chatbot with 128K context
Moonshot AI, a Tsinghua-linked startup founded that March, said Kimi could process 200,000 Chinese characters of input, a claimed world first for a consumer chatbot.
Models & capabilities
In the commentary
Pieces from around the web that discuss Moonshot AI. External links.
- 12 August 2026 · Matteo Wong · The AtlanticIt May Be Time to Panic About AI
- 20 July 2026 · Jack Clark · Import AIImport AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan
- 20 July 2026 · Nathan Lambert · InterconnectsKimi K3: The open-weights escalation
- 20 July 2026 · Zvi Mowshowitz · Don't Worry About the VaseOn Kimi K3: Its Capabilities And Related Discontents
- 4 February 2026 · Zvi Mowshowitz · Don't Worry About the VaseKimi K2.5
- 11 November 2025 · Zvi Mowshowitz · Don't Worry About the VaseKimi K2 Thinking
- 6 November 2025 · Nathan Lambert · Interconnects5 Thoughts on Kimi K2 Thinking
- 21 July 2025 · Jack Clark · Import AIImport AI 421: Kimi 2 - a great Chinese open weight model; giving AI systems rights and what it means; and how to pause AI progress
Also mentioned in 15 entries
Referenced in passing — Moonshot AI isn't the main subject of these.
- August 2026Alibaba's Qwen models pass 3 billion downloads
- August 2026Researchers extract AI models' hidden reasoning across three major labs' APIs
- July 2026DeepSeek releases V4-Flash update
- June 2026MiniMax releases MiniMax-M3, combining frontier coding, 1M context and native multimodality
- May 2026Epoch AI: open models lag closed frontier by four months
- April 2026White House memo addresses distillation of US AI models
- February 2026StepFun releases Step 3.5 Flash, topping several reasoning benchmarks
- January 2026Baidu launches ERNIE 5.0, a 2.4-trillion-parameter native multimodal model
- December 2025Tencent releases Hunyuan 2.0
- October 2025MiniMax open-sources MiniMax-M2 for coding and agentic workflows
- August 2025ByteDance open-sources Seed-OSS-36B
- June 2025ByteDance releases Doubao 1.6 model
- July 2024StepFun launches Step-2, a trillion-parameter MoE model
- January 2024Zhipu launches GLM-4, claiming near-GPT-4 parity
- August 2023ByteDance launches Doubao chatbot in invitation-only testing