Kimi K3
modelYour notes
Moonshot's next-generation flagship: a 2.8T-parameter Stable LatentMoE (16 of 896 experts active) built on Kimi Delta Attention (from Kimi Linear) and Attention Residuals, with Gated MLA, native vision, and a 1M-token context window — the architectural leap the K2 line (1T/32B) built toward. Runs at max thinking effort by default with low- and high-effort modes. AA Intelligence Index 44 (v4.3) — #4 overall and the highest-scoring Chinese model at its July 2026 release, still above Grok 4.5 (39 on v4.3), though AA flags verbosity (~2× median output tokens) and slow serving. Moonshot positions it as trailing only Claude Fable 5 and GPT-5.6 Sol. $0.30/$3/$15 per MTok (cache-hit/miss/output); live on Kimi.com, Kimi Work, Kimi Code, and the API at launch. Weights (~1.56TB) and the technical report shipped on schedule July 27 under a bespoke "Kimi K3 License" — the largest open-weight model to date (8-bit instruct checkpoint only; no base checkpoint). The report discloses 104B active parameters and the layer mix: 69 KDA + 24 Gated MLA layers with attention residuals.