Most Agentic AI Spend Goes to Fixing Answers, Not Getting Work Done

McKinsey says refinement loops eat most agentic AI spend. Falling token prices did not stop the bill from rising.

calender-image
July 21, 2026
clock-image
7 min read
Most Agentic AI Spend Goes to Fixing Answers, Not Getting Work Done
Free weekly briefingThe Business AI Briefing for people who run the Business — 5 min, zero hype.
Get the briefing free →

Via MarketScale: 60% of agentic AI costs go to response refinement: McKinsey

Your agents are spending more time rewriting themselves than finishing the job

Token prices fell. Enterprise AI bills still rose. That is not a mystery if most of the spend is agents checking, revising, and regenerating their own answers.

MarketScale summarized a July 2026 McKinsey report showing response refinement consumes about 60 percent of total agentic AI costs. Separately, McKinsey Enterprise AI FinOps Survey data cited in the piece found 93 percent of surveyed enterprise participants already exceed their AI budgets.

For an owner-led firm, the lesson is not that agents are useless. It is that autonomy without cost-per-outcome measurement turns a cheap model into an expensive loop. If you cannot see where tokens go by workload phase, you are buying refinement theater and calling it productivity.

Why the bill rose while unit prices collapsed

MarketScale points to three forces McKinsey called out: more AI workloads at scale, consumption-based pricing that rewards longer outputs, and frontier models used for routine tasks cheaper models could handle.

Menlo Ventures data cited in the McKinsey report showed enterprise large language model spending tripled over a 12-month period by the end of 2025. Stanford HAI 2025 AI Index figures in the same coverage show inference cost for GPT-3.5-level capability collapsing from about 20 USD per million tokens to 0.07 USD through 2024 — more than a 99 percent drop. Falling unit prices plus rising total spend is a volume and architecture problem, not a pricing paradox.

McKinsey forthcoming 2026 State of AI survey results cited by MarketScale found one in five organizations already constrained AI use because of AI-related operating costs. Boards expected cheaper models to keep total spend tame. What they underestimated was refinement loops multiplying across agents, tasks, and retries.

Blog Image

What smart firms measure before they scale agents

Tokens are the bill, not the value — a framing MarketScale attributes to David Tepper of Pay-i via the McKinsey coverage. Cost per million tokens is the wrong KPI if humans still rewrite half the output.

  • Attribute spend by phase. Split retrieval, reasoning, generation, and refinement. If you cannot see the 60 percent McKinsey flagged, you cannot cut it.
  • Match model to task. Map each workflow to the smallest capable model. Frontier models on routine drafts are a documented cost driver.
  • Track cost per completed task. Include human correction time. An agent that retries less but needs less cleanup can be cheaper than a thrifty-looking loop.
  • Cap autonomy until value clears cost. Require a written answer: does the completed work exceed the full operating cost of generating it?
  • Kill quiet pilots. If nobody owns the metric, the refinement loop owns your budget.

Response refinement consumes 60 percent of total agentic AI costs, according to McKinsey as reported by MarketScale.

How AgentsROI helps turn agent spend into a decision

AgentsROI starts with a Workflow ROI Audit: find where AI saves real time and money, and where refinement loops are burning cash without a completed outcome. You get a prioritized, costed roadmap instead of another vague pilot.

Then Model Selection and Continuity Planning matches the model to the job — cost versus capability versus privacy — with fallbacks so a price change or model sunset does not break the workflow.

Vendor-neutral by design. The point is not to pick a brand. The point is to stop paying for invisible retries you cannot defend to a board or a bookkeeper.

Stop buying refinement loops you cannot measure

If 60 percent of agentic spend is response refinement and most surveyed enterprises are already over budget, the next hire is not another agent. It is a measurement habit. Audit the loops, right-size the models, and keep only the workflows where completed work beats full cost.

Talk to AgentsROI about a Workflow ROI Audit before the next consumption invoice surprises you.

This article summarizes publicly reported information and is for general informational purposes only. It does not constitute legal, tax, financial, investment, security, or compliance advice. AgentsROI.ai is not a law firm, accounting firm, or registered investment adviser. Facts, pricing, statistics, and product capabilities cited here reflect the sources listed at the time of writing and may change. Readers should verify current information independently and consult qualified professionals regarding obligations specific to their industry, jurisdiction, and circumstances—including applicable New York State and New York City requirements. AgentsROI.ai may have commercial relationships with vendors mentioned; where material, such relationships are disclosed. Nothing in this article is an endorsement of any specific AI product, model, or provider.