AI agents, rather than people, are now the main consumers of AI tokens on OpenRouter.

According to figures cited in the article, Feb. 6, 2026 marked the point when agent usage first caught up with human usage on the platform. By Aug. 10, that gap had widened sharply. Agent-related token consumption rose from about 0.51 trillion to 7.3 trillion, a 14x increase, while human usage climbed to 1.4 trillion, up 2.8x. On that basis, agents were consuming about 5.2 times as many tokens as humans.
The piece says the shift happened within roughly six months. Peter described it as feeling like a 10-year jump, even though the Feb. 6 crossover was only half a year earlier.
How OpenRouter separates human and agent traffic
The article spends time on where the numbers came from and how they were calculated.
OpenRouter is described as one of the world’s largest model gateways, connecting about 70 model providers on one side and developers on the other. It processes 28 trillion tokens a week. Chris Clark, the company’s co-founder and COO, estimated that this represents about 1% of global inference volume, with half of that traffic coming from the United States.
Within that 1%, OpenRouter does not classify traffic by who the user is. It looks at how the API key is being used. The company tracks behavior and scores it across seven signals, then sorts traffic into Agentic, Mixed, and Human categories.
Those signals include how often tools are called, how many seconds pass between responses, and how many back-and-forth turns happen in a session. The logic is straightforward: human users tend to ask a question, pause, read, think, and type again. Agents keep moving, call tools repeatedly, and run in loops. The rhythm is different enough to make the traffic distinguishable.
The article notes that this is still an estimation model rather than a direct identity check, but says the directional reading is unlikely to be wrong.

It also stresses that token counts do not measure how many people are using AI. They show how much work AI is doing on a user’s behalf. A single human objective can now trigger dozens of model calls and tool invocations.
Agents generate far more calls per task
The article argues that the real gap comes from the way humans and agents use models.
A typical human request is simple: open a chat window, type a prompt, wait for a response, copy the result, and close the session. From start to finish, the exchange may involve fewer than 10 turns and only a few thousand tokens.
An agent request looks very different. A user sets a goal, and the system keeps going on its own: reading files, calling tools, writing results, reading again, and revising until the task is done. The human front-loads objectives, rules, and context into the initial prompt, and the model keeps reading from and writing against that context as the task progresses.
That changes the order of magnitude quickly. Someone who spends the morning copying and pasting in a chat interface can move to an agent-driven workflow in the afternoon and end up with a completely different token bill. Peter’s point, as quoted in the article, is that the real unlock comes when people let agents run autonomously for long stretches.
Cached tokens are cheaper, but total costs are still rising
A 14x jump immediately raises the issue of cost, and the article addresses that directly.
Peter added that agent requests are naturally multi-turn, and on average nearly 70% of the tokens in a single agent request come from cached prompts. Cached tokens are usually priced far below standard input tokens.

The reason is that agents work against the same goal over repeated cycles. The first pass loads a large amount of context, such as coding rules, operating manuals, or tool lists, at full price. Later turns only update what changed, while the cached portion is billed at a small fraction of the normal input rate.
In effect, much of the steeply rising token curve comes from reusing the same context again and again. The article says a16z put the figure at more than 85% on a total-volume basis, while Peter’s original post put the average at nearly 70% per request.
That lower unit price has not stopped budgets from being hit as usage surges.
The article cites Uber CTO Praveen Neppalli Naga, who said earlier this year that the company’s engineers burned through the full-year Claude Code budget in just four months. According to media accounts referenced in the piece, a two-hour demo by Naga himself cost $1,200.
EY offered another benchmark. It estimated that the cost of the same customer service interaction rose from about $0.04 in 2023 to about $1.20 in 2026, a 30x increase. The article says this does not mean models became more expensive. What changed was the workflow. Instead of a straight line where a customer asks one question and the system gives one answer, the same request may now trigger ticket lookups, inventory checks, record reviews, and multiple rounds of revision before a final response is produced.
Goldman Sachs, the article says, estimated that agentic AI could push token consumption 24x higher by 2030.
Peter said that when a cost line rises 14x in half a year, people start examining it closely. The article links that pressure to growing attention on open-weight models and low-cost token providers.

It also points to memory as part of the equation. Cached tokens may be cheaper, but they still need memory. If an agent runs for hours without restarting from scratch, memory is what holds the working context together. The piece argues that this is one reason high-bandwidth memory has become so tight.
The dividing line is workflow, not model choice
The article’s next point is that the key split is not which model a company uses. It is whether the model has been wired into real work.
The same model might consume only a few thousand tokens a day if it is used like a search box. Connect it to internal files, tools, and a process that can keep running, and tens of millions of tokens in a day are no longer unusual.
To illustrate the gap, the article cites OpenAI enterprise data. Among the top 10% of companies by depth of usage, active users generated 8.3 times as many output tokens as users at a typical company. Back in January, that gap was only 2.6x. In roughly six months, the spread had widened to more than three times its earlier level.
The same batch of statistics showed that 21% of active users at top companies used plugins every week, versus 9% at typical companies. Skill adoption stood at 19% versus 3%. Inside OpenAI itself, that figure was 95%.
The article says that once common tasks are packaged into reusable skills, agents do not have to rediscover the process each time. Fewer repeated searches and fewer rework cycles translate into lower costs.
Human review remains the bottleneck
The final section turns to the question of who signs off on all this work.

Rising agent traffic does not mean AI is acting fully on its own. Many tasks are still initiated by people, and human approval often remains in the middle of the workflow.
In comments under Peter’s original post, one developer serving clients in the service sector said those business owners never intended to let go entirely. Purchasing, sending, and signing still had to come back to a person at every step.
But routing work back to a person does not guarantee that someone actually checks it. Another industry participant said an agent deleted 18 videos from an advertising account and no one noticed until a full day later. The article presents that less as a spectacular model failure and more as a review failure: no one was watching that step.
Token spending can jump 14x in half a year. Headcount for validation cannot. The article argues that this may be one of the most overlooked and expensive lines in the current growth cycle.
Repeated execution in areas such as analysis, research, first drafts, and proposal writing may become abundant. What stays scarce, the piece says, is verification, judgment, and deciding what is worth doing in the first place. AI may now be the dominant consumer of tokens, but it still cannot sign on a human’s behalf.
The article closes on that point: the question ahead is no longer simply whether to use AI, but who reviews the work AI has completed and who decides the next move.

