Log in
Abstract lattice of dormant gray nodes with a sparse diagonal path of glowing teal nodes — most experts stay dark while only a few fire for each token

Dev Log /

Kimi K3: Open Weights Reach the Frontier

  • Models

On July 16, Moonshot AI released Kimi K3 and published every one of its weights. That alone makes it a milestone: at 2.8 trillion parameters, K3 is the largest open release yet — a mixture-of-experts model that activates 104 billion of those parameters per token, sees images natively, and holds a 1M-token context window. It is also the first open-weights model to sit within touching distance of the two systems that own the closed frontier, Claude Fable 5 and GPT-5.6 Sol.

A new ceiling for open weights

Over the past year, open models advanced fastest along one axis: test-time computation — reasoning effort, agent loops, more thinking per answer. The pre-trained foundations underneath mostly stayed in or near the 1T-parameter class, and as ever-stronger reinforcement learning piled onto foundations of similar scale, open progress risked converging while the gap to the strongest proprietary systems kept widening.

K3 pushes both axes at once. Its pre-training foundation is 3T-class — unprecedented for an open release — and large-scale RL is layered on top of it. Sparse MoE math keeps all of this servable: 896 routed experts, of which only 16 fire for any given token. Fewer than two experts in a hundred activate per token, so capability grows far faster than serving cost — the whole network’s knowledge carried on 104B-active economics. Refined data and training recipes add roughly 2.5× scaling efficiency over Kimi K2: each unit of training compute buys about two and a half times more capability than the previous generation’s recipe did.

Three inventions, three jobs

Three architectural ideas divide the work of moving information through a model this size.

Kimi Delta Attention (KDA) handles sequence. It carries long-range memory efficiently across a million tokens, and its decay curve is deliberately bounded — floored rather than free to fall toward zero — which removes a slow special-case path from the kernel and lets every chunk of attention run on dense tensor-core matrix multiplication. Every fourth layer swaps in Gated MLA for full global mixing across positions.

Attention Residuals handle depth. Each block can selectively reach back and draw on the outputs of earlier blocks, with learned per-layer weights deciding how much history to reuse — information flows across the network’s depth instead of dying between layers.

Stable LatentMoE handles width. Growing the routed expert space to 896 experts normally destabilizes training; Quantile Balancing derives each expert’s routing bias from score quantiles, keeping extreme sparsity balanced enough to train reliably.

Vision completes the design: a MoonViT-V2 encoder maps images into the same embedding space as text, so multimodality is built in from the first token rather than bolted on after the fact.

Where it lands

The benchmark spread reads like a closed-lab leaderboard with one bar in different colors:

  • Terminal-Bench 2.1: 88.3 — above Claude Fable 5 (88.0), half a point behind GPT-5.6 Sol (88.8).
  • ProgramBench: 77.8 — the best reported score, ahead of Sol and Fable 5.
  • SWE-Marathon: 42.0 — seven points clear of Fable 5 on GPU-kernel engineering.
  • FrontierSWE: 81.2 — second only to Fable 5 (86.6), well ahead of Sol (71.3).
  • BrowseComp: 91.2 — the top reported agentic-search score, above even Sol.
  • AutomationBench: 30.8 — first again.

The honest counterweights: DeepSWE sits at 67.5, behind both closed leaders, and GDPval professional-work Elo (1686) trails Fable 5 and Sol as well. Across Moonshot’s entire evaluation suite, though, the pattern holds steady — K3 trails exactly two systems and beats everything else measured, open or proprietary.

Selected Kimi K3 benchmark results: Terminal-Bench 2.1 88.3, FrontierSWE 81.2, DeepSWE 67.5, and BrowseComp 91.2, each shown against GPT-5.6 Sol, Claude Fable 5, Claude Opus 4.8, GPT-5.5, and GLM-5.2

Built for thousand-tool-call sessions

Post-training is aimed squarely at long-horizon work. Reinforcement learning ran across verifiable search, professional knowledge work, software and kernel optimization, vision-in-the-loop tool use, persistent assistant workflows, web development, and autonomous execution — each spanning multiple reasoning-effort levels. Environments push agents through hundreds to thousands of tool calls and millions of accumulated context tokens, drilling one loop until it holds: reason, act, observe, verify, adapt. Million-token rollouts ride on persistent sandbox states that survive interruption and resume mid-task.

Domain specialists then merge into one model. Multi-teacher on-policy distillation consolidates effort- and domain-specific policies into a single set of weights, so one model serves everything from quick answers to maximum-effort deep work — you turn one dial instead of swapping systems.

What you actually get

The full weights are on Hugging Face under moonshotai/Kimi-K3, vision encoder included, with the 1M-token context and reasoning-effort levels intact. The license deserves a plain statement: these are open weights published under Moonshot’s custom Kimi K3 license — not an MIT-style grant. Read its terms before embedding K3 in anything commercial.

Six months ago, choosing open weights meant accepting last generation’s capability. K3 retires that trade-off. Open territory now sits on the frontier itself, and the question is no longer whether open models can reach the top table — it is which lab pulls up the next chair.

Join iqshard

Choose the servers and Dev Log updates you want.

Subscribe to updates