On July 25, 2026, Kimi released K3, its most capable model to date. With 2.8 trillion parameters, it is the first open model to reach the 3T-class scale. K3 features native vision capabilities and a 1-million-token context window, built on Kimi Delta Attention and Attention Residuals architectures.
While K3's overall performance still trails the most powerful proprietary models—Claude Fable 5 and GPT 5.6 Sol—it demonstrated frontier-level performance across Kimi's evaluation suite, consistently outperforming other tested models. The model is available immediately on Kimi.com, Kimi Work, Kimi Code, and the Kimi API, with full model weights to be released by July 27, 2026.
K3 uses a Mixture of Experts (MoE) architecture with 896 experts, activating 16 per token via a Stable LatentMoE framework. Combined with refined training recipes, these innovations yield a 2.5× improvement in scaling efficiency over Kimi K2. The model excels in long-horizon coding, agentic knowledge work, and scientific research, as demonstrated by several case studies.
In coding benchmarks, K3 performed competitively with Claude Fable 5 in GPU kernel optimization and built a complete compiler called MiniTriton from scratch, rivaling Triton and torch.compile. It also designed a chip autonomously in 48 hours, achieving 8,700 tokens/s decode throughput in simulation.
K3's vision capabilities enable it to create interactive 3D experiences, edit videos, and produce motion-graphics explainers. In knowledge work, it completed a computational astrophysics reproduction task in two hours that would typically take one to two weeks. New features in Kimi Work—Widgets and Dashboard—allow users to create persistent, interactive visualizations.