Dev Log
Notes, published by the iqshard development team.
Kimi K3 vs GLM-5.3: The Scoreboard Isn’t the Verdict
Scorecards give Kimi K3 the edge over GLM-5.3. Matched metrics, independent evals, and finished-task costs tell a more useful story for anyone building software.
- Models
GLM-5.3 Just Shipped, and It Is Not a Small Update
GLM-5.3 is a post-training-only upgrade that nearly doubles AutomationBench, more than doubles ExploitBench, and closes the gap to GPT-5.6 Sol on agentic coding.
- Models
Why More Context Doesn't Mean Better Answers
A large context window tells you how much a model can accept, not how much it can use well. Better answers come from smaller, targeted context
- Mechanics
Kimi K3: Open Weights Reach the Frontier
Moonshot AI releases Kimi K3 with full open weights — a 2.8T-parameter MoE with native vision and a 1M-token context, and the closest an open model has come to the closed frontier.
- Models
Why an Agent Uses Far More Tokens Than a Chatbot
A chatbot and an agent run on the same kind of model and the same tools. What makes an agent burn far more tokens is not how it works, but what it is for
- Mechanics
GLM-5.2: Built for Long-Horizon Work
GLM-5.2 holds a stable 1M-token context and leads open-weights models on long-horizon, tool-using work — the kind of work that resembles a real job.
- Models
Input, Output, Cached: The Three Kinds of Token in a Request
A request is made of three kinds of token — input, output, and cached — and each is handled differently. How they work, and why caching reuses text across users
- Mechanics
Three Variations on the Leaky Bucket: Queue, Meter, and Token Budget
Three rate-limiter designs apply the same leaky-bucket concept as a queue, an activity meter, or a replenishing token budget.
- Mechanics
What Exactly Are Tokens?
Tokens are how AI models read and write, and how you are charged. A plain guide to input, output, and cached tokens for people who sign the invoices
- Mechanics
The NVIDIA DGX B300: An AI Factory in Ten Rack Units
Eight Blackwell Ultra GPUs, 2.304 TB of HBM3e, and NVLink stitching eight accelerators into one very large GPU — inside the NVIDIA DGX B300.
- Infrastructure
Join iqshard
Choose the servers and Dev Log updates you want.