
Dev Log /
Kimi K3 vs GLM-5.3: The Scoreboard Isn’t the Verdict
- Models
Five weeks separated 2026’s two biggest open-weights releases: Moonshot AI’s Kimi K3 on July 16, Z.ai’s GLM-5.3 on August 14. Both hold a million-token context, both sit near the top of the coding leaderboards, and both are built to drive software-engineering agents. Pick up any aggregator scoreboard and Kimi K3 looks like the clear pick. Work through matched metrics, independent evaluations, and what finished tasks actually cost, and a more interesting picture emerges — one that matters if you are choosing a model to build software with.
First, the scoreboard — then its footnotes
Third-party aggregators credit Kimi K3 on five of nine shared benchmarks. Their headline gap is ProgramBench: 77.8 versus 19.0. That comparison is not what it looks like — the two numbers measure different things. 77.8 is K3’s fully-solved rate; 19.0 is GLM-5.3’s almost-solved rate, labeled as such in Z.ai’s own launch-day table. On matched metrics the “gap” does not just shrink, it disappears: Z.ai’s table has GLM-5.3 ahead on ProgramBench almost-solved (19.0 vs 17.5) and dead level with K3 on NL2Repo — from-scratch repository reconstruction — at 58.0 apiece.
Sources disagree elsewhere too: SWE-Marathon scores K3 at 48.1 in Z.ai’s table and 42.0 in Moonshot’s own report. Public leaderboards mix harness settings and metric definitions, so treat cross-source deltas of a few points as noise, and read single-harness comparisons before verdicts.
Z.ai’s launch-day comparison ran both models through one harness. Every benchmark it reports for both models:
| Benchmark | GLM-5.3 | Kimi K3 |
|---|---|---|
| Terminal-Bench 2.1 | 88.2 | 88.3 |
| Terminal-Bench 3.0 | 28.3 | 17.4 |
| DeepSWE v1.1 | 66.9 | 67.5 |
| NL2Repo | 58.0 | 58.0 |
| ProgramBench (almost-solved) | 19.0 | 17.5 |
| SWE-Marathon v1.1 | 42.5 | 48.1 |
| PostTrainBench | 39.8 | 32.0 |
| Toolathlon Verified | 73.0 | 76.5 |
| AutomationBench v1.0.6 | 48.2 | 46.7 |
| HLE with tools | 62.5 | 59.8 |
| GDPval-AA v2 (Elo) | 1769 | 1682 |
| CyberGym | 84.5 | 80.0 |
| ExploitBench | 54.4 | 32.2 |
Read down those two columns: dead heats on repair-style coding, real GLM-5.3 margins on longer-horizon suites like Terminal-Bench 3.0 and PostTrainBench, and lopsided leads in security engineering. K3’s genuine edges live in tool orchestration (Toolathlon) and kernel-level marathon work. Nothing here resembles a class difference.
Independent evaluations level it further
Artificial Analysis’ Intelligence Index scores both models identically: 60. One independent lab went deeper, running each model through SWE-bench Verified — 500 real issues from open-source projects — and Terminal-Bench 2.1 with complete agent trajectories. Verified finished 94.2% to 93.8% in GLM-5.3’s favor; their Terminal-Bench run finished 86.5% to 80.9%.
More revealing than the aggregates is which tasks each model solved. Of the 500 Verified issues, 453 fell to both models and 34 fell to exactly one — GLM-5.3 took 18 exclusive solves, Kimi K3 16. Same ceiling, different blind spots. By category, GLM-5.3 dominated data-pipeline fixes 11–4 and ML tasks 5–1; K3 led observability work 4–1. Two models this close do not fail on the same problems — which is exactly why headline percentages undersell both of them.
The bill that arrives at the end
Sticker prices run $1.40/$4.40 per million input/output tokens for GLM-5.3 against $3/$15 for Kimi K3 — 2.1× and 3.4× apart. Agents multiply tokens, though: one task means many calls, long contexts, and cached reads ($0.26 vs $0.30 per million). The unit that matters is cost per completed task. On SWE-bench Verified trajectories it came to $0.18 for GLM-5.3 against $0.77 for Kimi K3 — four times as much for a marginally lower completion rate. The deep Terminal-Bench run told the same story: $0.46 versus $0.57 per task.
Equal-or-better completion at roughly a quarter of the spend is not a pricing footnote. It decides how much work an agent can simply be left to do.
Outside pure coding
GLM-5.3 carries non-coding edges too: Humanity’s Last Exam with tools goes 62.5 vs 59.8, GDPval professional-work Elo runs 1769 vs 1682, and offensive-security engineering is not close — 54.4 vs 32.2 on ExploitBench and 84.5 vs 80.0 on CyberGym. Kimi K3 answers outside coding as well: its 91.2 on BrowseComp is the top reported agentic-search score, and across general-agent benchmarks it trails only Fable 5 and GPT-5.6 Sol.
Where GLM-5.3 cannot follow
Vision. GLM-5.3 reads text only; Kimi K3 natively sees images. Screenshot-driven UI iteration, mockup-to-code workflows, render-and-inspect loops — that territory belongs to K3 alone, and its vision earns the trip. In one lab’s open-ended build-off, both models were asked to create the same game from scratch: K3 inspected its own rendered scenes and iterated on sprites, layout, and composition until its build looked genuinely shippable.
Keep one caveat beside that polish: seeing is not correctness. K3’s handsome game shipped a Stage 2 whose final jump was physically impossible — its validation had teleported past the gap. GLM-5.3’s plain-looking build was playable start to finish, because it kept re-testing traversal physics with scripted playthroughs. Whichever model drives your visual loop, keep functional checks inside it.
Which one belongs in your agent
For non-visual software engineering, the quality race is a tie — matched benchmarks deadlock, independent deep evals split hairs — and everything else points one way: GLM-5.3 finishes tasks at about a quarter of the cost, ships under an MIT license rather than a custom one, and carries about a month fresher training data. For screenshot-heavy visual iteration or search-heavy agent work, Kimi K3’s advantages are real; walk in knowing the ~4× cost per task and the license terms.
And if your stack can route between them, do: an oracle picking the better result per task reached 97.4% on SWE-bench Verified — proof that both models still solve things the other cannot.