
Dev Log /
What Exactly Are Tokens?
- Mechanics
Ask what an AI model costs and the answer comes back in a unit most people have never had to think about: the token. Every prompt you send and every answer you get back is measured in tokens, and the meter runs on both. If your company is starting to see AI on its invoices, this is the number underneath them.
A token is the smallest piece of text a model handles at once. It is tempting to call it a word, and for a rough count that is close enough. But it is not quite right, and the gap is where a lot of surprise costs hide.
Not quite words
Models do not read the way you do. Before a model sees your text, the text is broken into tokens: whole words when they are short and common, fragments when they are long or rare. A tokenizer might keep "the" and "revenue" intact, then split "tokenization" into "token" and "ization". Punctuation and numbers usually stand on their own. A comma is a token. A dollar sign is a token.
NVIDIA gives a clean example. The word "darkness" splits into "dark" and "ness"; "brightness" splits into "bright" and "ness". The shared "ness" carries the same internal number both times, which is part of how the model learns that two words are related. You see one word. The model sees two pieces.
Because of this, a token count and a word count are not the same. A useful rule of thumb from OpenAI: in ordinary English, one token runs about four characters, or roughly three-quarters of a word. A hundred tokens land near seventy-five words. Short everyday sentences track close to word-for-word. Dense text with long technical terms, code, or unusual formatting drifts well above it.
The best way to feel the difference is to watch it move. Type into the box below and change a few words. Swap a plain word for a long one. Add a price, or a comma, or a run of spaces. The count reacts.
An estimate for illustration. Every model uses its own tokenizer, so the exact count varies — the only precise number comes from running the model itself.
The estimate here is deliberately rough, because there is no single right answer. Every model family carries its own tokenizer trained on its own vocabulary, so the same sentence lands at slightly different counts depending on who you send it to. Anthropic recently noted that its newer models produce roughly thirty percent more tokens for the same text than its older ones, purely because the tokenizer changed. The words did not get longer. The unit did.
Input, output, and the shortcut in between
Here is where the bill takes shape. A single request has two sides, and they are priced differently.
Input tokens are everything you send: your question, your instructions, and any documents or data you paste in. The model reads all of it at once, in a single pass. Reading is the cheaper side.
Output tokens are what the model writes back. These come out one token at a time, each one depending on the last, so generating them takes more work than reading. That difference shows up on every price sheet. Output tokens commonly cost several times more than input tokens — Turing Post puts the range at two to six times, depending on the model.
That split has a practical edge. A long, rambling answer costs more than a short, precise one, so asking for a tight format instead of a wordy explanation is a real saving, not a style preference. And because the model reads everything you send, padding a prompt with documents it does not need is paying to have noise read back to you. More context can also make the answer worse, even when every token fits inside the model's window.
Cached tokens are the shortcut. When you send the same opening text again — the same standing instructions, the same reference document, the same long preamble — the model can reuse the work it already did on that stretch instead of processing it from scratch. Providers charge far less for those reused tokens, often a fraction of the normal input rate, and the savings compound in any workload that repeats the same context across many calls. Anthropic, Google, and OpenAI all offer some version of it. The main catch is that a cache does not live forever; leave a gap long enough and the discount lapses.
Why the unit matters to a budget
Tokens sound like a technical detail. They are really a pricing detail, and treating them as one changes how the numbers behave.
The word count of a task rarely predicts its token count. A page of legal text, a spreadsheet, a screenshot, a stretch of code — each carries far more tokens than its length suggests, because each tokenizes inefficiently. A prompt padded "just in case" is not free insurance. Kelly, writing for finance and HR leaders, put a number on it: a prompt twenty percent less efficient than it needs to be can drive costs well past double, as the model works through material it was never going to use.
None of this requires you to count tokens by hand. It requires knowing that the count is there, that it does not match your instinct for how long a piece of text is, and that the three kinds — input, output, cached — each move the bill in their own direction. The teams that stay ahead of their AI spend are the ones who learned to read that meter early, rather than meeting it for the first time on an invoice.
Start with the box above. Once you can see a sentence turn into tokens, the pricing stops being a mystery and starts being arithmetic.