Counting tokens
A request is three token counts, not one. Credits turn that into a single comparable number — and one tied directly to dollars.
Not all tokens are equal
The server still has to count something. Not to bill you and not to cut you off, but to tell one request from another.
Every request produces three separate token counts, not one. Cached input is text the server has already seen and can re-read cheaply. New input is text it has to read for the first time. Output is text it has to generate — the expensive one, by a wide margin. If tokens themselves are new to you, start with what a token is.
So a request is a triple, not a total. Adding the three together throws away the only thing that distinguishes them, which is why "tokens used" describes almost nothing on its own. Two requests with identical token totals can demand several times more work from the server than one another, depending on how those tokens split.
Counting requests is no better. A short question and a full agentic coding turn are both one request, and they are nowhere near the same amount of work.
Why that becomes a problem
The scheduler does not predict which queued request will be heavy and move it ahead of another. Requests stay in arrival order inside each participant's queue, and every preferential dispatch costs one credit before the work begins.
The difference appears when the request finishes. Only then does the server know how much cached input it read, how much new input it processed and how much output it generated. Turning those three counts into one number lets the scheduler charge the actual work back to the bucket that admitted it.
That accounting changes what happens next, not what already ran. Heavy work leaves less priority behind; light work leaves more. A credit is the comparable unit that makes that retrospective accounting possible: a weighted sum of the three token counts, with each type weighted by how much work it actually costs.
Where those credits come from, and what happens when you have spent them, is a separate story. Credits that come back tells it.
What a credit is worth
The weights track the standard pay-as-you-go retail rates for the model, so a credit is tied directly to dollars. One credit buys a different number of tokens depending on which kind you spend it on, and roughly the same amount of money either way.
| Token type | Credits per token | Tokens per credit | Retail rate | Retail per credit |
|---|---|---|---|---|
| Cached input | 0.0001 | 10,000 | $0.26/M | $0.0026 |
| New input | 0.0004 | 2,500 | $1.40/M | $0.0035 |
| Output | 0.0012 | 833 | $4.40/M | $0.0037 |
Output costs twelve times what cached input costs, which is the whole reason a single token total tells you nothing. A credit lands between a quarter and a third of a cent of inference at retail, whichever way you spend it.
Work out a request
Pick a shape of request, or set your own, and see what it comes to in credits and at retail.
Request profile
Tokens per request
52,850
Credits per request
6.66
Retail price per request
$0.0152
| Entitlement | Total requests | Total tokens | Equivalent retail value |
|---|---|---|---|
| 8,800 credits | 1,321 | 69.8M | $20.03 |
| 44,000 credits | 6,607 | 349.2M | $100.16 |