Log in
Split composition: charcoal geometric circuit forms on the left face teal forms on the right, interlocking at a central seam

Dev Log /

Kimi K3 vs GLM-5.3: The Scoreboard Isn’t the Verdict

  • Models

Five weeks separated 2026’s two biggest open-weights releases: Moonshot AI’s Kimi K3 on July 16, Z.ai’s GLM-5.3 on August 14. Both hold a million-token context, both sit near the top of the coding leaderboards, and both are built to drive software-engineering agents. Pick up any aggregator scoreboard and Kimi K3 looks like the clear pick. Work through matched metrics, independent evaluations, and what finished tasks actually cost, and a more interesting picture emerges — one that matters if you are choosing a model to build software with.

First, the scoreboard — then its footnotes

Third-party aggregators credit Kimi K3 on five of nine shared benchmarks. Their headline gap is ProgramBench: 77.8 versus 19.0. That comparison is not what it looks like — the two numbers measure different things. 77.8 is K3’s fully-solved rate; 19.0 is GLM-5.3’s almost-solved rate, labeled as such in Z.ai’s own launch-day table. On matched metrics the “gap” does not just shrink, it disappears: Z.ai’s table has GLM-5.3 ahead on ProgramBench almost-solved (19.0 vs 17.5) and dead level with K3 on NL2Repo — from-scratch repository reconstruction — at 58.0 apiece.

Sources disagree elsewhere too: SWE-Marathon scores K3 at 48.1 in Z.ai’s table and 42.0 in Moonshot’s own report. Public leaderboards mix harness settings and metric definitions, so treat cross-source deltas of a few points as noise, and read single-harness comparisons before verdicts.

Z.ai’s launch-day comparison ran both models through one harness. Every benchmark it reports for both models:

BenchmarkGLM-5.3Kimi K3
Terminal-Bench 2.188.288.3
Terminal-Bench 3.028.317.4
DeepSWE v1.166.967.5
NL2Repo58.058.0
ProgramBench (almost-solved)19.017.5
SWE-Marathon v1.142.548.1
PostTrainBench39.832.0
Toolathlon Verified73.076.5
AutomationBench v1.0.648.246.7
HLE with tools62.559.8
GDPval-AA v2 (Elo)17691682
CyberGym84.580.0
ExploitBench54.432.2

Read down those two columns: dead heats on repair-style coding, real GLM-5.3 margins on longer-horizon suites like Terminal-Bench 3.0 and PostTrainBench, and lopsided leads in security engineering. K3’s genuine edges live in tool orchestration (Toolathlon) and kernel-level marathon work. Nothing here resembles a class difference.

Grouped bars comparing GLM-5.3 and Kimi K3 across twelve shared benchmarks from a single harness: near-ties on repair-style coding, GLM-5.3 ahead on longer-horizon and security suites, Kimi K3 ahead on Toolathlon and SWE-Marathon

Independent evaluations level it further

Artificial Analysis’ Intelligence Index scores both models identically: 60. One independent lab went deeper, running each model through SWE-bench Verified — 500 real issues from open-source projects — and Terminal-Bench 2.1 with complete agent trajectories. Verified finished 94.2% to 93.8% in GLM-5.3’s favor; their Terminal-Bench run finished 86.5% to 80.9%.

More revealing than the aggregates is which tasks each model solved. Of the 500 Verified issues, 453 fell to both models and 34 fell to exactly one — GLM-5.3 took 18 exclusive solves, Kimi K3 16. Same ceiling, different blind spots. By category, GLM-5.3 dominated data-pipeline fixes 11–4 and ML tasks 5–1; K3 led observability work 4–1. Two models this close do not fail on the same problems — which is exactly why headline percentages undersell both of them.

Venn diagram of SWE-bench Verified results: 453 tasks solved by both models, 18 solved only by GLM-5.3, 16 solved only by Kimi K3, and 13 solved by neither

The bill that arrives at the end

Sticker prices run $1.40/$4.40 per million input/output tokens for GLM-5.3 against $3/$15 for Kimi K3 — 2.1× and 3.4× apart. Agents multiply tokens, though: one task means many calls, long contexts, and cached reads ($0.26 vs $0.30 per million). The unit that matters is cost per completed task. On SWE-bench Verified trajectories it came to $0.18 for GLM-5.3 against $0.77 for Kimi K3 — four times as much for a marginally lower completion rate. The deep Terminal-Bench run told the same story: $0.46 versus $0.57 per task.

Equal-or-better completion at roughly a quarter of the spend is not a pricing footnote. It decides how much work an agent can simply be left to do.

Outside pure coding

GLM-5.3 carries non-coding edges too: Humanity’s Last Exam with tools goes 62.5 vs 59.8, GDPval professional-work Elo runs 1769 vs 1682, and offensive-security engineering is not close — 54.4 vs 32.2 on ExploitBench and 84.5 vs 80.0 on CyberGym. Kimi K3 answers outside coding as well: its 91.2 on BrowseComp is the top reported agentic-search score, and across general-agent benchmarks it trails only Fable 5 and GPT-5.6 Sol.

Where GLM-5.3 cannot follow

Vision. GLM-5.3 reads text only; Kimi K3 natively sees images. Screenshot-driven UI iteration, mockup-to-code workflows, render-and-inspect loops — that territory belongs to K3 alone, and its vision earns the trip. In one lab’s open-ended build-off, both models were asked to create the same game from scratch: K3 inspected its own rendered scenes and iterated on sprites, layout, and composition until its build looked genuinely shippable.

Keep one caveat beside that polish: seeing is not correctness. K3’s handsome game shipped a Stage 2 whose final jump was physically impossible — its validation had teleported past the gap. GLM-5.3’s plain-looking build was playable start to finish, because it kept re-testing traversal physics with scripted playthroughs. Whichever model drives your visual loop, keep functional checks inside it.

Which one belongs in your agent

For non-visual software engineering, the quality race is a tie — matched benchmarks deadlock, independent deep evals split hairs — and everything else points one way: GLM-5.3 finishes tasks at about a quarter of the cost, ships under an MIT license rather than a custom one, and carries about a month fresher training data. For screenshot-heavy visual iteration or search-heavy agent work, Kimi K3’s advantages are real; walk in knowing the ~4× cost per task and the license terms.

And if your stack can route between them, do: an oracle picking the better result per task reached 97.4% on SWE-bench Verified — proof that both models still solve things the other cannot.

Join iqshard

Choose the servers and Dev Log updates you want.

Subscribe to updates