AI Cost Pullback [What It Changes]
TL;DR
AI cost pullback is the moment a team scales back or freezes AI agents because the cost of running them beats the value those agents prove, usually because nobody can see the meter in real time.
As of 24 June 2026, KPMG found only 26% of large-firm leaders had full, real-time visibility into what AI costs to operate, while 35% still named token and inference literacy as a barrier.
McKinsey's 2026 State of AI survey (25 August 2026) found about 20% of organizations already constrained AI use because of operating costs, even as most still plan to spend more.
A 2-truck HVAC shop, a 10-person agency, or a solo clinic should treat this as a billing problem, not a model-brand problem: cap tokens, default to a cheap model, and only spend frontier rates on steps that change a customer outcome.
Key Takeaways
The constraint that broke is visibility, not curiosity: dashboards are common, token-level controls are not.
List prices keep falling on small models, but agent loops, retries, and web tools multiply the bill.
Large firms are scaling agents faster than small ones; a small shop that copies the large-firm stack without a cap copies the overrun.
Seat licenses and token meters are different products; mixing them without a workflow owner is how a monthly tool budget disappears in a quarter.
Pullback is a routing change: keep the workflow, swap the model, and measure cost per completed job.
What AI cost pullback actually is
AI cost pullback is a deliberate scale-back of AI agents after run costs outrun proven benefit, usually because the team cannot see token and inference spend as the work happens.
That definition is not a Wall Street story. It is a shop-floor story. A 2-truck HVAC company that lets an agent rewrite every job note, a 10-person marketing agency that lets an agent crawl the web on every brief, or a solo clinic that lets an agent summarize every chart is buying the same meter the $1 billion firms just learned to fear. The difference is that those firms have a finance staff. You probably do not.
The U.S. Small Business Administration Office of Advocacy counts 34,752,434 small businesses in the United States, and according to that same Advocacy FAQ, 99.9% of U.S. businesses are small. Those shops already buy software the way they buy parts: a known monthly number. Token billing is the opposite. It moves with how chatty the agent is, how often it retries, and whether someone turned on web search. If you already care about the state of small-business automation, this is the new line item on that same bill.
The Census Bureau's Statistics of U.S. Businesses program still publishes firm, establishment, and payroll counts by enterprise size, which is the backbone for those small-business totals. Nothing in that series will warn you when an agent burns a month of margin on a Friday afternoon. That warning has to live in the workflow.
What the 2026 surveys actually measured
On 24 June 2026, KPMG LLP released its U.S. AI Quarterly Pulse Survey, fielded 28 April to 25 May among 204 U.S. C-suite and business leaders at organizations with $1 billion or more in annual revenue. Only 26% of organizations see live AI run costs. That sentence is the core of AI cost pullback, and it sits in KPMG's own release.
According to KPMG, 18% of those organizations now orchestrate multiple AI agents across workflows, up from 9% in the prior quarter. Agent deployment itself held roughly steady at 53% versus 55%. The work got more connected. The bill got harder to see.
The same KPMG pulse reports 66% of organizations have monitoring dashboards and 61% have approval processes, yet only 36% have direct token or usage controls. Leaders still plan a weighted-average $202 million of AI investment over the next 12 months, nearly unchanged from $207 million the quarter before. Spend intent did not fall. Control did not catch up.
Yahoo Finance republished the CFO Dive write-up of that survey on 25 June 2026 and quoted KPMG's Rahsaan Shears: "AI agents are changing both the operating model and the economics." Axios reported the same 26% visibility figure the same week and quoted Shears saying clients were going through budgeted amounts faster than they anticipated.
McKinsey's State of AI in 2026, dated 25 August 2026, widens the lens. According to McKinsey, about 20% of respondents report that AI-related operating costs, including token costs, constrained their AI use. About 20% of firms constrained AI use over opex. That is pullback in survey form, published in the same McKinsey report.
According to that McKinsey survey, 40% of respondents at organizations with more than $1 billion in revenue report scaling AI agents, up from 27% a year earlier, while the share at smaller organizations stayed flat at 22%. Enterprise-level EBIT impact attributed to AI sat at 37%, unchanged from the prior year, even as 80% of respondents said AI improved their own productivity. Individual speed is not the same as a firm-level return.
A companion McKinsey piece on agentic economics (13 July 2026) reports that 93% of respondents to a McKinsey survey exceeded their AI budgets. The same article says McKinsey itself was processing roughly five trillion AI tokens a month as of May 2026, that about 10% of users accounted for about 65% of token consumption, and that concise prompts and outputs can cut token use 30% to 40% on some workflows. It also states that agentic tasks can consume roughly 1,000 times more tokens than simple chat, and that about 60% of an agentic task's costs sit in refining answers rather than producing the first draft.
That is the mechanism. The first answer is cheap. The retries, the tool calls, and the "are you sure?" loops are the invoice.
Fortune reported on 26 May 2026 that Uber had already used its entire 2026 AI coding-tools budget in four months after ranking teams on an internal leaderboard by tool usage. Uber president and COO Andrew Macdonald told the Rapid Response podcast, in Fortune's account, "That link is not there yet" between those usage stats and shipping 25% more useful consumer features. Uber CEO Dara Khosrowshahi said on an earnings call, as quoted in the same Fortune story, that about 10% of the company's committed code is built by autonomous agents. If a company of that size cannot draw the line from tokens to features, a 10-person shop should not assume the line will draw itself.
| Metric | Q2 2026 figure | Comparison figure |
|---|---|---|
| Organizations deploying AI agents | 53% | 55% prior quarter |
| Multi-agent orchestration across workflows | 18% | 9% prior quarter |
| Real-time visibility into AI run cost | 26% | — |
| Monitoring dashboards in place | 66% | — |
| Approval processes in place | 61% | — |
| Direct token or usage controls | 36% | — |
| Leaders citing cost-management / token literacy as a barrier | 35% | — |
| Planned AI investment, next 12 months (weighted average) | $202 million | $207 million prior quarter |
| Employee resistance to AI agents | 20% | 5% prior quarter |
| Survey sample (U.S., $1B+ revenue) | 204 leaders | Fielded 28 Apr–25 May 2026 |
Sources: KPMG Q2 AI Quarterly Pulse Survey; Yahoo Finance / CFO Dive.
| Survey signal (2026) | Figure | Population |
|---|---|---|
| Constrained AI use because of operating costs (incl. tokens) | ~20% | McKinsey global State of AI respondents |
| Large orgs ($1B+) scaling AI agents | 40% | Up from 27% prior year |
| Smaller orgs scaling AI agents | 22% | Flat vs prior year |
| Respondents attributing some EBIT impact to AI | 37% | Unchanged vs prior year |
| AI high performers (≥5% EBIT, "significant" impact) | ~6% | All respondents |
| Enterprise-scale AI deployment | 44% | Up from 38% prior year |
| Skipped buying software because coding agents can build it | 32% | All respondents |
| Individual productivity gain from AI | 80% | All respondents |
| Respondents exceeding AI budgets (separate McKinsey survey) | 93% | McKinsey agentic-economics survey |
Sources: McKinsey, The state of AI in 2026; McKinsey, Is that AI agent worth it?.
How the meter actually works
A token is a billing slice of text, not a sentence. Providers charge for tokens in and tokens out, then add extras when the agent searches the web, runs code, or stores a cache. OpenAI's public API price list posts GPT-5.6 Luna at $0.20 input and $1.20 output per million short-context tokens, GPT-5.6 Terra at $2.00 and $12.00, and GPT-5.6 Sol at $4.00 and $20.00. The same ladder appears on OpenAI's business pricing page. Cached input on Luna is $0.02 per million tokens. Web search is $10.00 per 1,000 calls.
Anthropic's Claude API table lists Claude Haiku 4.5 at $1 input and $5 output per million tokens, Claude Sonnet 5 at $2 and $10, and Claude Opus 5 at $5 and $25. Cache reads are a fraction of input: 0.1× on most models. The Claude consumer plans page is a different product: Pro at $20 per month or $17 per month on annual billing, Max from $100 per month, and Enterprise at $20 per seat per month plus usage billed at API rates. A seat is not a token cap.
Google's Gemini Developer API prices Gemini 3.8 Flash at $0.75 input and $3.75 output per million tokens through 31 December 2026, then $1.50 and $7.50 from 1 January 2027. Google Cloud Agent Platform pricing posts the same $0.75 / $3.75 introductory rates for Gemini 3.7 Flash and Gemini 3.6 Flash through 31 December 2026. Batch and flex paths cut those Flash rates in half on that Cloud page.
Microsoft Azure OpenAI pricing matches the OpenAI short-context ladder on Global GPT-5.6-luna at $0.20 input and $1.20 output per million tokens, and it offers a Batch API at a 50% discount on Global Standard. The same Azure page warns that pricing increases for Microsoft Foundry EU Data Zone and non-U.S. regional deployments start 1 September 2026. Amazon Bedrock publishes its own per-model token tables for the same idea: you pay for what the agent reads and writes, on the cloud bill you already have.
OpenAI's Batch API states a 50% cost discount versus synchronous APIs with a 24-hour completion window. Anthropic posts the same 50% batch cut: Haiku 4.5 batch is $0.50 / $2.50 per million tokens. Overnight jobs — invoice coding, CRM backfill, transcript cleanup — belong on batch. Live chat does not.
| Model (short context, standard) | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| OpenAI GPT-5.6 Luna | $0.20 | $1.20 |
| OpenAI GPT-5.6 Terra | $2.00 | $12.00 |
| OpenAI GPT-5.6 Sol | $4.00 | $20.00 |
| Anthropic Claude Haiku 4.5 | $1.00 | $5.00 |
| Anthropic Claude Sonnet 5 | $2.00 | $10.00 |
| Anthropic Claude Opus 5 | $5.00 | $25.00 |
| Google Gemini 3.8 Flash (through 31 Dec 2026) | $0.75 | $3.75 |
| Azure GPT-5.6-luna Global | $0.20 | $1.20 |
| OpenAI Batch API vs sync | −50% | −50% |
| Anthropic Batch Haiku 4.5 | $0.50 | $2.50 |
Sources: OpenAI API pricing; OpenAI business pricing; Anthropic Claude API pricing; Gemini Developer API pricing; Azure OpenAI pricing; OpenAI Batch API.
The FinOps Foundation treats this as the same discipline cloud bills already required: teams collaborate, everyone owns usage, and variable cost is a feature if you can see it. The FinOps Framework says business value should drive technology decisions and that FinOps data should be accessible, timely, and accurate. FOCUS, the FinOps Open Cost and Usage Specification, exists because AI, cloud, and SaaS invoices do not share a native language. A small shop does not need a FinOps department. It needs one person who can read a token dashboard the way they read a fuel report.
Datadog Agent Observability is one of the tools built for that read: it traces prompts, tool calls, latency, and token usage, and it bills on LLM spans (one provider call) rather than every surrounding workflow step, with on-demand charges after the first 100,000 LLM spans. You do not need that vendor. You do need the habit: every agent step should show tokens, model name, and cost before the next step runs.
Seat products still exist beside the meter. Microsoft 365 Copilot is sold as an in-app assistant with agents you add from a store, not as a raw token table on the marketing page. Mixing an uncapped API agent with a seat license without naming an owner is how two bills describe the same work and neither one is trusted.
Why this broke now
Three things arrived at once.
First, agents stopped being a single chat box. KPMG recorded multi-agent orchestration doubling to 18%. Each extra agent resends context. McKinsey describes long-lived context as a recurring cost because models are stateless, so prior work gets paid for again.
Second, incentives rewarded volume. KPMG named "token-maxxing" — gamifying token use with incentives and leaderboards — and said 41% of leaders would consider it while 22% opposed it and 37% stayed neutral, in the 24 June 2026 release. Fortune described Uber's leaderboard as the setup that exhausted a year of coding-tool budget in four months. If you pay people to spend tokens, they will spend tokens.
Third, the physical layer got expensive even as unit prices fell. According to the International Energy Agency, data centres accounted for around 1.5% of world electricity consumption in 2024, or 415 terawatt-hours, and consumption is set to more than double to around 945 TWh by 2030. The IEA Energy and AI report frames that as the supply side of the same meter: training and running models live in power-hungry data centres. A typical AI-focused data centre, in the IEA executive summary, consumes as much electricity as 100,000 households, and global data-centre investment amounted to half a trillion dollars in 2024. You do not pay that invoice directly. You pay it as tokens, seats, and cloud.
Regulation is arriving on a calendar, not a vibe. The EU AI Act overview states that transparency rules take effect in August 2026, GPAI model rules took effect in August 2025, and high-risk system obligations start 2 December 2027. Logging, documentation, and human oversight are cost lines, not slogans. The OECD AI Principles, updated May 2024 and now with 47 adherents, put accountability and robustness next to innovation. NIST's AI Risk Management Framework, released 26 January 2023 as NIST AI 100-1, is voluntary and organizes work into Govern, Map, Measure, and Manage. NIST AI 600-1, the generative-AI profile from July 2024, tells organizations to treat generative systems as a distinct risk set. ISO/IEC 42001:2023 is the first AI management-system standard; ISO lists it as a 51-page standard published December 2023. None of those documents set your token price. All of them assume you can explain what the system did and what it cost.
USTA analysis: the 30-point control gap and the 16.7× model ladder
USTA analysis. This section uses only figures already cited above. No new survey numbers.
Input A: KPMG reports 66% of organizations have monitoring dashboards. Input B: the same release reports 36% have direct token or usage controls. Dashboard-to-control gap = 66 − 36 = 30 percentage points. That 30-point band is the operational shape of AI cost pullback: most large firms can watch a chart, and a minority can stop the meter. A small shop that buys a dashboard and skips the cap has copied the expensive half of that pattern.
Input C: OpenAI lists GPT-5.6 Sol output at $20.00 per million short-context tokens. Input D: GPT-5.6 Luna output is $1.20 per million. Sol/Luna output ratio = 20.00 ÷ 1.20 = 16.67×. For a month with 10 million output tokens on one workflow: Luna cost = 10 × $1.20 = $12.00; Sol cost = 10 × $20.00 = $200.00; delta = $188.00. That $188 is not a forecast. It is list-price arithmetic on 10 million output tokens. If the step is "rewrite a dispatch note," Luna is the default. If the step is "decide whether to void an invoice," Sol may be worth the 16.7×. The pullback move is to stop paying 16.7× on every step.
What a small team should change this week
Name one owner. If two people can start an agent and nobody can stop it, you have already failed the KPMG visibility test described in the 24 June 2026 pulse.
Put a hard monthly budget on every API key. OpenAI's business pricing FAQ tells customers they can set a monthly budget in billing settings after which requests stop, with a possible delay, and that overage is still the customer's problem. Use that control. Do not rely on a spreadsheet.
Default every workflow to the cheap model. Route only the exception path to a frontier model. Teams already routing documents through US Tech Automations workflows will plug this in as a model swap, not a rebuild: the intake form still lands in the same queue, the CRM still gets the same fields, and only the model ID on that node changes.
Cap tools separately from tokens. OpenAI charges $10.00 per 1,000 web-search calls. An agency agent that "just checks the web" on every draft can outspend the language model. Turn search off unless the step requires a live fact.
Move batch work to batch. Nightly form-to-CRM cleanup, invoice coding, and transcript roll-ups match the OpenAI Batch 50% discount and 24-hour window. Live customer replies stay on the synchronous path.
Kill leaderboards that rank token spend. KPMG and Fortune both describe volume games. Rank completed jobs, closed tickets, or posted invoices instead.
Reuse context with a cache instead of resending the handbook on every turn. Anthropic prices a cache hit at 0.1× input on most models (0.025× on Fable 5.1 and Mythos 5.1). A clinic's policy pack and an HVAC firm's price book are cache candidates.
Keep humans on the money steps. NIST AI 100-1 treats AI systems as socio-technical. A bookkeeper still signs the return. If you run accounting practice workflows or a Drake versus ProConnect tax path, the agent can draft; the licensed person still files. The same rule applies to executive-assistant automation: let a cheap model sort the inbox, not spend Sol rates rereading the whole thread.
If you need a reporting spine for the new line item, treat AI spend like any other operating cost you already compare in tools such as Fathom versus Jirav: a named account, a monthly cap, and a variance note when the agent retries.
HVAC dispatch notes that land in a US Tech Automations document step should call a cheap model by default, with a spend cap on the expensive one. That is the whole pullback, expressed as a node setting.
Honest limits
The KPMG sample is 204 U.S. leaders at $1 billion-plus firms. It is not a survey of 2-truck shops. Directionally it still matters, because those firms have more staff to watch the bill and still only 26% see it live. Industry slices inside the same release are stark: 54% of technology leaders said operating costs were fully visible, 31% of banking leaders said the same, and 4% of asset-management and private-equity leaders said costs were fully visible.
McKinsey is a global management survey, not a census. The 20% "constrained use" figure is the pullback signal; the 80% personal-productivity figure is why executives keep spending anyway. Those two numbers can both be true.
List prices change. The tables above are the pages as fetched for this article. OpenAI states GPT-5.6 Sol promotional pricing is available at least through 21 November 2026. Google states Flash promotional rates through 31 December 2026. Do not freeze a quote from a screenshot six months old.
This article does not claim a 49% global "scaled back agents" figure. The 9 August 2026 Forbes URL in the original source pack did not load, so that figure is omitted.
Signal vs Speculation
Demonstrated fact (sourced): As of 24 June 2026, KPMG documented 26% real-time cost visibility, 35% cost-literacy barriers, 36% token controls, 66% dashboards, and a doubling of multi-agent orchestration to 18% among 204 U.S. leaders at $1 billion-plus firms. As of 25 August 2026, McKinsey documented about 20% of organizations constraining AI use over operating costs, 40% of large organizations scaling agents versus 22% of smaller ones, and 37% reporting some EBIT impact. Vendor list prices and 50% batch discounts are on the OpenAI, Anthropic, Google, and Azure pages cited above. IEA documented 415 TWh of data-centre electricity in 2024.
Our read: If the 30-point dashboard-to-control gap holds, small and mid-size firms will feel AI cost pullback as a series of quiet freezes — a receptionist agent turned off after a surprise invoice, a research agent limited to ten runs a day — rather than a press-release strategy shift. Over the next 12–36 months, the shops that keep agents will be the ones that treat intelligence like overtime: allowed on named jobs, capped by default, audited monthly. The shops that copy large-firm "use more tokens" programs will copy Uber's four-month budget burn at a scale they cannot absorb. We expect seat-plus-meter bundles to keep spreading, which will make mixed bills worse unless one owner reconciles them. We do not expect unit token prices to save a badly routed workflow; McKinsey already shows agent loops wiping out cheaper units.
Glossary
AI cost pullback: Scaling back or freezing AI agents because operating cost exceeds proven benefit, usually after the team loses sight of the live meter.
Token: The unit providers use to bill model input and output; not the same as a word or a seat.
Inference: The act of running a trained model on a new request; this is the usage that shows up on the monthly bill.
Agentic workflow: A multi-step path where software calls tools, retries, and hands work between models without a person at every click.
Prompt cache: Stored prompt prefix billed at a discount on later hits so the same handbook is not prepaid every turn.
Batch API: An asynchronous job window, typically 24 hours, sold at about 50% off synchronous rates for work that can wait overnight.
FinOps: The practice of making variable technology spend visible and owned across finance, engineering, and the business.
Token-maxxing: Ranking or rewarding people for consuming more tokens, which KPMG flagged as a way to confuse activity with value.
FAQs
What is AI cost pullback?
AI cost pullback is when a team reduces or pauses AI agents because running them costs more than the benefit those agents can prove. It is a billing and visibility event, not a verdict on whether AI is useful.
Why do cheaper tokens still produce a bigger bill?
Because agents retry, resend context, and call tools. McKinsey reports that agentic tasks can use about 1,000 times more tokens than simple chat and that about 60% of an agentic task's cost sits in refinement. Cheaper units times many more units is a larger invoice.
How should a small shop meter agents?
Give one person the API keys, set a monthly budget on the provider, log model name and tokens on every step, and default to the cheapest model that completes the job. OpenAI documents a monthly budget control in billing settings. Turn that on before the first production run.
Do I need a frontier model for every step?
No. OpenAI prices Luna output at $1.20 per million tokens and Sol output at $20.00. That is a 16.7× gap on the same million tokens. Use the expensive model when the decision is expensive. Use the cheap model when the task is a rewrite, a sort, or a draft.
What is the difference between a dashboard and a token cap?
A dashboard shows spend after it happens. A cap stops requests. KPMG found 66% of large organizations have dashboards and 36% have token or usage controls. The gap is the pullback risk.
How do batch APIs change the bill?
They cut list prices in half for jobs that can wait up to 24 hours, per OpenAI and Anthropic. Overnight backfills belong there. Live customer chat does not.
Should we rank staff on how much they use AI?
Not on tokens. KPMG described token-maxxing as a way to reward activity over outcomes, and Fortune reported a usage leaderboard sitting next to a four-month budget burn. Rank completed work.
What to do next
AI cost pullback is already a practice at large firms even when they still call AI a priority. Small teams do not need a $202 million plan. They need a cap, a cheap default model, and a named owner. Map a capped agent path on agentic workflows, or start from the home page and open the same builder from there.
Teams that want that routing without a custom stack can inspect the agentic workflow builder from US Tech Automations. Keep the workflow. Swap the model. Watch the meter.
About the Author

Helping businesses leverage automation for operational efficiency.
Related Articles
See how AI agents fit your team
US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.
View pricing & plans