Log in
A crowded context window full of overlapping documents beside a smaller focused window where three relevant sources lead cleanly to one answer

Dev Log /

Why More Context Doesn't Mean Better Answers

  • Mechanics

A million-token context window sounds like a million tokens of memory. Put in the whole codebase, every project document, and the complete conversation, and the model should have everything it needs. In practice, having the information somewhere in the window is not the same as finding it, following it, or reasoning over it correctly. A context window measures capacity. Better answers depend on what fills it.

If tokens are new to you, start with what they are. Context is the next part of the mechanism: it is the set of those tokens the model can use when producing its answer.

The window is a limit, not a memory

Context includes more than the question you type. It holds the system instructions, conversation history, documents, tool definitions, search results, file contents, and every answer the model has already produced. All of it occupies the same working space.

The advertised context window tells you how many tokens the model can accept in that space. It does not promise equal attention to every token inside it. Researchers at Stanford demonstrated the gap by moving the same relevant document through a long prompt. Models used it best near the beginning or end. Put it in the middle and performance dropped, producing the U-shaped result now called "lost in the middle." The evidence never disappeared. Its position changed how reliably the model used it.

Anthropic describes the broader constraint as an attention budget. Every new token consumes some of it. As the input grows, the model has more relationships to track and more candidates competing for attention. That produces a gradient rather than a sudden failure: the model still works, but precision can fade before the window is close to full.

Two context windows with the same capacity: one packed with stale history, broad documents, and tool output around a small relevant signal; the other containing only the task, constraints, and relevant sources

Nothing breaks when quality starts to slip

That gradual decline is called context rot. It rarely arrives as an error message. The model starts overlooking an earlier constraint, reaches for a similar but incorrect fact, or gives a generic answer where it had been specific before. A coding agent may produce code that compiles while solving the wrong problem.

Chroma isolated the effect across 18 models. Its researchers kept the tasks fixed and changed the input length, so a longer prompt did not also mean a harder question. Performance became less reliable as the input grew, even on controlled retrieval and text-replication tasks. When the question and answer used different language, degradation came faster. Plausible distractors became more damaging as the context length increased.

This does not create one universal danger line. Different models and tasks degrade at different rates. Straightforward retrieval can work over very long inputs. Reasoning across scattered evidence is harder. The useful conclusion is not "never use long context." It is that the model-card limit cannot tell you how much context your task will use well.

Better evidence beat more evidence

Hexaware tested that distinction on a production codebase. Across eight coding tasks and three model families, it compared three ways to provide context: the entire codebase, short summaries plus a few full files, and only the full files needed for the task.

A chart comparing three coding-context strategies: targeted full files scored 88.5 percent weighted correctness at about 10,000 input tokens, the full codebase scored 80.2 percent at about 332,000 tokens, and compressed context scored 68.8 percent at about 19,000 tokens

The targeted set won. It averaged 88.5% weighted correctness from about 10,000 input tokens. Loading the full codebase used about 332,000 tokens and scored 80.2%. The compressed version was smaller than the full codebase, but it scored 68.8% because summaries removed exact function names, parameters, and module paths. The models retained enough shape to write plausible code, then invented identifiers that did not exist.

This was one codebase, eight tasks, and a controlled set of model calls, not a universal ranking. But it reveals the important distinction. Smaller did not win merely by being smaller. Targeted, exact context won. It removed competing material without throwing away the details the task required.

The wrong context can outweigh the right context

The same experiments compared retrieval quality. With the correct files, hallucinations ran at roughly 1%. Loading the full codebase raised that to about 30%. Supplying the correct files plus two plausible but wrong files reached roughly 33%.

The right answer was present in every one of those noisy prompts. The model did not fail because it lacked information. It failed because adjacent evidence competed with it. Similar-looking code was more dangerous than obviously irrelevant padding because it offered a believable wrong path.

That is why dumping every search result into a prompt is not harmless insurance. Neither is carrying old tool output through every step of an agent run. Agents use far more tokens because their jobs run longer, but a long job does not make every token from every earlier step useful to the next decision.

Build a working set, not an archive

Good context engineering starts with the smallest sufficient set, not the shortest possible prompt. Give the model the task, the constraints that govern it, and the exact evidence it needs. Keep full detail where names, figures, or wording matter.

Then retrieve the rest when the work reaches it. A path or document title can remain outside the window until the model needs the contents. Search results can be cleared after they have led to the relevant source. Long projects can keep decisions and current state in structured notes, then begin a clean session from that record instead of dragging the whole transcript forward.

Summaries help when the gist is enough. They hurt when the discarded detail is the thing the answer depends on. The same is true of sub-agents: isolating exploration can keep the main working set clean, but the handoff must preserve the findings that matter.

Use the capacity as headroom

A large context window remains useful. It lets a model open a long document, carry more state, and delay the point where something must be retrieved or compacted. But it is headroom, not a target.

Anthropic's rule is the practical one: find the smallest set of high-signal tokens that maximizes the chance of the result you want. Enough context to answer correctly. Little enough that the answer does not have to fight its way through an archive to find it.

Join iqshard

Choose the servers and Dev Log updates you want.

Subscribe to updates