Dell's 2028 token estimate jumped 57× — and today's run rate is already higher. Agentic and reasoning workloads are quietly rewriting the bill.

Via I/O Fund: AI Token Demand is Shattering Forecasts
While Wall Street argues whether AI will monetize, the usage meter is not waiting for consensus. According to I/O Fund's July 30, 2026 analysis, annual token processing is now discussed in quadrillions — and those forecasts keep getting revised upward after they are already wrong.
For an owner-led firm, that is not a chip-stock story. It is a cost-control story. If your team moves from chatbots to agents and reasoning models without caps, routing, or ROI checks, the bill grows because each task quietly consumes more tokens — not because someone "bought more AI seats."
Dell COO Jeffrey Clarke captured the miss in October 2025: the company once modeled about 1 quadrillion tokens by 2028; the revised figure was 57 quadrillion — and he still expected to be low. I/O Fund notes Dell also said inference-demand expectations rose at least 100× in under a year.
The punchline for operators: I/O Fund cites current processing near 370 trillion tokens per day (roughly 135 quadrillion annualized), already about 2.4× Dell's revised 2028 estimate — with years left on the calendar. If the industry underestimates demand this badly, your internal "we budgeted for ChatGPT" line item is not a plan.
The underestimation is not one vendor's bad spreadsheet. I/O Fund stacks several independent misses that point the same direction.
Goldman Sachs (May 2026, as cited by I/O Fund) projected about 47 quadrillion tokens per month in 2028 — roughly 565Q per year, around 10× Dell's figure — and saw monthly processing rising from 1.7Q in mid-2025 toward nearly 120Q by mid-2030. Goldman also estimated agentic workloads could account for roughly 80%+ of that total. I/O Fund argues even those numbers may be conservative versus a then-current monthly run rate near 11Q.
Tirias Research, per the same piece, once expected about 20 trillion annual tokens by end-2024; actual usage was estimated at 667 trillion — more than 33× the original call. Anthropic's CEO Dario Amodei planned for 10× growth in 2026, then reportedly saw revenue and usage climb about 80× annualized in Q1, straining compute.
Big-tech surfaces show the same slope: Google's monthly token processing rose to 3.2 quadrillion in May 2026 — about 330× since May 2024 and 7× since May 2025 — while Gemini alone was processing 22 billion tokens per minute (about 1Q per month) by the Q2 2026 call. Microsoft processed over 100 trillion tokens in a quarter (about 5× year over year). OpenRouter's weekly volume jumped from 5T to 25T in six months; Fireworks AI reported 40 trillion tokens per day by mid-July 2026.
The driver that matters for SMEs is workload mix. Anthropic estimates multi-agent systems can use up to 15× more tokens than chatbot requests. Stanford researchers, cited by I/O Fund, say coding agents can consume about 1,000× more tokens than code-reasoning chats. On OpenRouter, reasoning-model tokens went from near 0% early in 2025 to around 60% by late 2025 — while prompt tokens per request rose about 4× and completion tokens roughly tripled versus early 2024. More capable tools, used casually, become a consumption problem long before they become a strategy.
You do not need a hyperscaler forecast desk. You need operating habits that treat tokens like any other variable cost.
None of this requires picking a winner in the AI infrastructure trade. It requires refusing to run unbounded usage as if it were a flat SaaS seat.
"We thought… that inference would drive by 2028, 1 quadrillion tokens. Now it's 57 quadrillion, and I'm sure we're wrong." — Dell COO Jeffrey Clarke (Oct 2025), via I/O Fund
AgentsROI.ai is a stack-agnostic Managed Intelligence Provider for owner-led SMEs. We do not sell you a model. We help you see, govern, and operate the AI you already have so it pays for itself.
Start with a Workflow ROI Audit when you need a prioritized map of where AI saves money — and where agentic token burn is quietly eating the gains. That is the paid entry diagnostic for firms that suspect usage is outrunning value.
Pair that with Managed AI Operations so monitoring, optimization, and usage guardrails do not depend on the owner remembering to open a dashboard. When model choice and fallbacks are the bottleneck, Model Selection & Continuity Planning matches capability, cost, and privacy to the job — with a plan for when a provider throttles or reprices.
Vendor-neutral by design: cloud, hybrid, or local — whatever fits the work. The goal is simple: AI that pays for itself, with someone accountable for the meter.
The industry just admitted — again — that inference demand was underestimated by orders of magnitude. Your firm does not control Dell's spreadsheet or Goldman's 2030 curve. You do control whether agents run without budgets, whether reasoning models are the default for every draft, and whether anyone can explain last month's AI invoice in plain English.
If token spend is opaque, or agent pilots are expanding without caps, book a Workflow ROI Audit. We will show you what is burning tokens, what is worth it, and what to govern next — so the forecast miss stays Wall Street's headache, not yours.
This article summarizes publicly reported information and is for general informational purposes only. It does not constitute legal, tax, financial, investment, security, or compliance advice. AgentsROI.ai is not a law firm, accounting firm, or registered investment adviser. Facts, pricing, statistics, and product capabilities cited here reflect the sources listed at the time of writing and may change. Readers should verify current information independently and consult qualified professionals regarding obligations specific to their industry, jurisdiction, and circumstances—including applicable New York State and New York City requirements. AgentsROI.ai may have commercial relationships with vendors mentioned; where material, such relationships are disclosed. Nothing in this article is an endorsement of any specific AI product, model, or provider. This article is not investment advice and does not recommend buying or selling any security.