
Dev Log /
Input, Output, Cached: The Three Kinds of Token in a Request
- Mechanics
A request to a language model is made of three kinds of token, and the model treats each one differently. There is the text you send in, the text that comes back, and the text the model has already seen before. They are not interchangeable: the model reads input, writes output, and reuses what it has cached, and each of those is a different amount of work. Lump all three together as "tokens used" and you have thrown away the thing that tells them apart.
If tokens are new to you, start with what a token is — this picks up where that leaves off, with what happens to them inside a request.
Reading is cheaper than writing
The split starts with how a model handles the two halves of a request.
When you send a prompt, the model reads all of it at once. Every token in your input is processed together, in a single pass, before the model writes anything. Reading in parallel like this is fast and efficient, which is why input is the cheaper side of a request.
Writing works the other way. The model produces output one token at a time, and each new token depends on every token before it. It cannot write the tenth word of a sentence until it has written the ninth. That step-by-step generation is far more work than reading, so output tokens cost more — commonly two to six times the input rate, depending on the model, according to Turing Post's survey of provider pricing.
That gap has a practical edge. A rambling, verbose answer is not just harder to read — it is the expensive part of the request, generated one costly token at a time. Asking for a short, structured answer instead of a long explanation is a real saving. So is trimming the input: because the model reads everything you send, padding a prompt with documents it does not need means paying to have noise read to you, and then paying again for the longer answer that noise tends to provoke.
The third price: text the model has already seen
Now the interesting one. Cached tokens are input the model has processed before and can reuse instead of reading again from scratch.
To see why that saves so much, look at what the model builds when it reads. As it works through your prompt, it computes an internal representation of every token — a structure called the key-value cache, or KV cache. That computation is the cost of reading. If the next request begins with the exact same text, the model can reuse the KV cache it already built for that stretch rather than recomputing it. The work is already done.
This matters most when the same text leads every request. Consider a typical setup: a 4,000-token set of standing instructions, a 10,000-token reference document, and then a short question that changes each time. Without caching, the model reprocesses all 14,000 tokens on every single call. With caching, those 14,000 tokens are read once and reused on every request after the first. Only the short, changing question is new work.
The pricing reflects the saved effort. Anthropic charges a small premium the first time text is written to the cache, then serves cached reads at roughly a ninety percent discount from its normal input rate. Google's caching for Gemini works along the same lines, and OpenAI's varies by model. Kelly, writing for finance leaders, put the practical range at fifty to ninety percent off, depending on the model. For any workload that repeats a long preamble across many calls, that is not a rounding error — it is most of the input cost.
There is one catch worth knowing. A cache does not live forever. Providers hold it for minutes to hours, then discard it. If your requests are frequent, the cache stays warm and the discount holds. If they are sparse or bursty, the text may have aged out of the cache by the time you send the next request, and you pay the full read again.
One cache, many users
The version of this that surprises people is that a cache need not belong to a single user.
Think about how a popular AI coding tool works. Every one of its users sends requests that begin with the same long set of instructions — the same system prompt describing how the tool should behave. That preamble is identical across thousands of people. If each user's copy were cached separately, the same text would be computed thousands of times over.
Serving infrastructure that shares a cache across users computes that shared prefix once and makes it available to every request that starts with it. As Turing Post describes it, if a thousand agents share the same system prompt, the KV cache for that prompt is built a single time and reused across all of them. The more a prefix is shared, the more the saving compounds — the first user pays to warm it, and everyone after reads it cheaply.
This is why caching strategy has become something inference providers compete on rather than an afterthought. When many workloads pass through the same machine and share common instructions, the shared work is done once for all of them instead of once per user.
Why the three-way split is the whole point
A single number, "tokens used", hides all of this. Two requests with identical token totals can mean very different amounts of work depending on how those tokens divide across the three types. A request that is mostly cached input is light. A request of the same size that is mostly fresh output is not.
That is why the three counts are kept apart wherever the work is measured honestly. The mix is what distinguishes one request from another, and the total on its own tells you almost nothing. Read the split, and a request stops being an opaque lump of "tokens" and becomes something you can actually reason about.