Moonshot AI has released Kimi K3, a 2.8 trillion parameter mixture-of-experts model that has achieved performance parity with leading proprietary models like Claude Fable 5 and GPT-5.6. The model utilizes 896 total experts, with 16 activated per token, and features a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which the company claims provides a 2.5x improvement in scaling efficiency over its predecessor, Kimi K2. Despite its high performance, the model exhibits a 51% hallucination rate and a tendency to generate excessive tokens, which may increase operational costs. The release has sparked significant geopolitical debate, with some U.S. officials and industry figures characterizing open-weight models as a national security risk. While the model is currently available via API, Moonshot AI plans to release the full model weights on July 27, 2026. The high demand for Kimi K3 has already led to the temporary suspension of paid subscription plans due to GPU capacity constraints.
Kimi K3 is a 2.8 trillion parameter mixture-of-experts model featuring a 1-million-token context window. The model architecture activates 16 out of 896 experts per token to improve computational efficiency.
Kimi K3 demonstrates performance on coding benchmarks comparable to Claude Fable 5 and GPT-5.6. The model has a 51% hallucination rate and generates more tokens than necessary, potentially increasing costs.
Moonshot AI plans to release the full model weights to the public on July 27, 2026. High demand for the model has caused GPU capacity issues, resulting in the suspension of paid plans.
Chapter guide
Worth noting
- Benchmark results may be influenced by the specific harnesses used by Moonshot AI compared to competitors.
- The 51% hallucination rate is based on data from the Artificial Analysis Intelligence Index.
- The release of full model weights on July 27, 2026, is a stated plan and subject to change.
- The video contains a paid sponsorship from Mobbin.