Log in
A long dark circuit-board road carrying an unbroken glowing teal line to the horizon, mile-marker nodes lighting up along the way

Dev Log /

GLM-5.2: Built for Long-Horizon Work

  • Models

GLM-5.2 is Z.ai’s flagship model for long-horizon tasks, and at release it was the strongest open-weights model available for that kind of work. The headline capability is a context window that finally holds at scale: one million tokens of genuinely stable context, up from 200,000 on GLM-5.1, shipped under an MIT license with no regional restrictions. A million-token window is easy to advertise and hard to keep reliable across a long, messy, multi-hour coding session — GLM-5.2 was trained specifically to hold up under that sustained pressure, not merely to accept more tokens on paper.

A million tokens that hold under load

What a million tokens buys is the difference between a model that answers questions and one that does a job: whole codebases with their change histories, days of accumulated tool output, agent sessions that run for hours without anyone summarizing the middle to make room. Long context usually fails quietly — retrieval degrades, and the model loses the thread of an argument it made tens of thousands of tokens earlier. GLM-5.2 was trained on environments measured in hours and tens of hours precisely so the far end of the window stays as usable as the near end.

Head to head with the closed frontier

On FrontierSWE — open-ended engineering projects running from hours to tens of hours — GLM-5.2 scores 74.4%, essentially tied with Claude Opus 4.8’s 75.1% and ahead of GPT-5.5’s 72.6%. It is the highest-ranked open model on every long-horizon benchmark in Z.ai’s suite, and it beats GPT-5.5 outright on several: 62.1 vs 58.6 on SWE-bench Pro, 76.8 vs 75.3 on the MCP-Atlas tool-use benchmark, and 54.7 vs 52.2 on Humanity’s Last Exam with tools.

Grouped bars comparing GLM-5.2 and GPT-5.5 across six benchmarks: GLM-5.2 leads on the longer-horizon suites — FrontierSWE, SWE-bench Pro, MCP-Atlas, and Humanity's Last Exam with tools — while GPT-5.5 leads on the shorter single-shot DeepSWE and Tool-Decathlon

The honest counterweights: GPT-5.5 clearly wins the shorter, single-shot contests — 70.0 vs 46.2 on DeepSWE and 55.6 vs 48.2 on Tool-Decathlon. The pattern across the full spread is consistent, though: the shorter and more self-contained the task, the better GPT-5.5 fares; the longer and more agentic the task, the closer GLM-5.2 presses the closed frontier — until, on genuinely long-horizon work, it draws level and in places passes it.

IndexShare: making a million tokens affordable

Stable long context is an architecture story, not just a training one. GLM-5.2 introduces IndexShare, which lets one lightweight indexer serve every four sparse-attention layers instead of each layer recomputing its own — a 2.9× reduction in per-token FLOPs at the full 1M context length. The speculative-decoding layer got the same treatment: reusing top-k indices and sharing the KV cache across prediction steps lifts acceptance length by roughly 20%, which converts directly into faster generation at the same output quality.

What it serves like

Those architecture gains show up in serving economics, not just in a paper. Normalized against GLM-5.1 at a 32K context, GLM-5.2 holds 4.69× the throughput at 200K tokens — the exact point where GLM-5.1’s window ends and it cannot serve the request at all — and keeps scaling to 6.97× at the full million. The successor does not merely serve longer requests; it serves them faster than its predecessor served requests a fraction of the size.

A dial for effort

GLM-5.2 also introduces effort-level control — thinking budgets that trade capability against latency and cost. Z.ai positions its capability under comparable token budgets between Claude Opus 4.7 and Claude Opus 4.8: quick answers on the low settings, deep agentic work at Max, one model covering the whole range instead of separate fast and deep systems.

Pure open

The weights ship under a standard MIT license with no regional restrictions — weights anyone can take, deploy, fine-tune, and build products on without negotiating terms. For a model at this capability level, that is the rarest spec on the page.

Put it together — a million-token context that holds under real engineering pressure, benchmark results that draw level with leading closed models exactly as tasks get longer and more agentic, serving throughput that makes the long context economical to run, and an unencumbered license — and GLM-5.2 reads less like a routine version bump and more like a statement: the gap between open and closed is closing fastest precisely where the work that matters happens.

Join iqshard

Choose the servers and Dev Log updates you want.

Subscribe to updates