Log in
A small chatbot robot with one speech bubble on the left, and a larger multi-armed agent robot juggling tools inside a looping cycle on the right

Dev Log /

Why an Agent Uses Far More Tokens Than a Chatbot

  • Mechanics

A chatbot and an agent are built from the same kind of language model. These days they even work in similar ways: both can call tools, run code, and take several steps before they hand anything back. Yet a chatbot answers you in a few thousand tokens, while an agent can run through millions to finish a single job. The gap does not come from how they work. It comes from what they are for.

New to tokens? Start here. For why input, output, and cached tokens are counted apart, this covers it. This picks up from there.

Under the hood, the line has blurred

It used to be simple. The first ChatGPT, in 2022, did exactly one thing: it took the conversation so far and predicted what came next. No tools, no code, no browsing. You typed, it wrote back, and that was the whole machine.

That is not what a chatbot is anymore. Ask a modern one to crunch some numbers and it will write a Python script, run it in a sandbox somewhere on the server, read the result, and format it into a table for you. Ask a follow-up and it will edit the script and run it again. Somewhere behind the chat box, a small tool-using loop is doing real work.

So the loop is no longer the thing that separates a chatbot from an agent. Anthropic defines an agent as "an LLM autonomously using tools in a loop," and by that description a modern chatbot turn qualifies. If the machinery is shared, the cost gap has to come from somewhere else.

Two panels built from the same parts, a language model, tools, and an internal loop. The chatbot panel answers your message with you in the loop; the agent panel pursues your goal and runs unattended

What a chatbot is, and what an agent is

Strip away the shared parts and the difference is the job each one is pointed at.

A chatbot is an interactive tool for a conversation. You send a message, it works out a reply, and it hands the reply back to you. It might use tools to get there, and it might take a few internal steps, but the unit of work is the exchange. It is scoped to the conversation, you are in the loop, and every turn is bounded: it answers, then waits for you. A support reply-bot is a chatbot. So is a self-service portal with a chat window on the front. So is the assistant that just wrote you that Python script.

An agent is handed a goal, not a message. A coding tool like Claude Code is the clear case. You point it at a task, and it works across a persistent environment, your source tree, your git history, your file system, and it keeps going. It reads code, makes an edit, runs the tests, reads the failures, tries again. It does this for minutes, hours, sometimes days, and you are not there for each step. You set the goal and it works until the goal is met.

The word "chatbot" has drifted a long way. The 2022 version and the 2026 version share a name and not much else. But even the newest one, tools and sandbox and all, is still doing the original job: answering you, in a conversation.

Two different jobs

Both exist because they answer different needs, and the shape of each job is what sets the token count.

A conversation is bounded by you. You ask, it answers, you read the answer, you ask the next thing. Between turns, nothing runs. Even a demanding turn is one scoped task with a clear end: produce the reply, then stop and wait. The work can only get so big before it has to hand control back.

A goal is bounded only by being finished. "Migrate this service to the new API" or "track down the flaky test and fix it" has no natural stopping point until the thing is actually done. The agent chooses the next step itself, takes it, checks the result, and continues. A single goal can hold dozens or hundreds of the scoped tasks a chatbot would handle one at a time, except the agent runs them back to back, on its own, without pausing for you.

A timeline. A chatbot is a short cluster of bounded turns with gaps where you reply, measured in thousands of tokens. An agent is one long continuous run stretching across hours, measured in millions of tokens

Why the tokens add up

Now the count. A chatbot reply stays small because the conversation keeps it small: one question, one answer, a pause. A whole exchange is a few thousand tokens.

An agent run gets large because the goal makes it large. Surveys of agent frameworks put a single agent at roughly four times the tokens of a standard chat exchange, and a multi-agent system at fifteen times or more. Those are per-comparison figures; stretched across a real run they compound. One documented autonomous coding session went for twenty-five hours and consumed thirteen million tokens on its way to thirty thousand lines of code.

There is no single multiplier to quote, because the number depends entirely on how big the goal is and how long the agent runs. That is exactly the point. The token count tracks the work, and the work tracks the use case.

A long run does not make every earlier token useful to every later step. Search results, file reads, and failed approaches accumulate as an agent works, so keeping its context focused matters as much as making the window larger.

The takeaway

A chatbot and an agent can be built from the same model and reach for the same tools. What separates them is the job. A chatbot holds a conversation, scoped to a task, and waits for you between turns. An agent takes a goal and works until it is done, however long that takes and however many steps it runs. One stays small because a conversation stays small. The other grows without a fixed ceiling, because finishing a goal on its own is a great deal of work, and every step of that work is counted in tokens.

Join iqshard

Choose the servers and Dev Log updates you want.

Subscribe to updates