93% of Enterprises Are Exceeding Their AI Budget. Most Don’t Know Why.

Most enterprise AI programmes are not failing on capability. They’re failing on cost visibility.

According to McKinsey’s Enterprise AI FinOps survey, 93% of organisations report exceeding their AI budget, and 46% are over by 10 to 30%. This isn’t a rounding error in a technology line item. It’s a structural blind spot, and it gets worse, not better, as AI adoption scales.

The pattern McKinsey describes is consistent with what happens whenever a new consumption-based technology outpaces the governance built to manage it: spend becomes invisible before it becomes a problem, and by the time it’s a problem, the fix is an emergency contract renegotiation rather than a planned decision.

Why AI spend breaks the old cost-management playbook

Many CIOs default to comparing AI spend management to the early days of cloud computing, and the comparison holds at the governance level: unclear ownership, fragmented purchasing, and reactive budget conversations are familiar territory. But McKinsey identifies five features specific to AI that make it a materially harder problem than cloud ever was.

Usage is genuinely unpredictable. The same task can generate dramatically different token volumes depending on which model it invokes and how many agent chains it triggers, and agentic workflows multiply model calls per interaction in ways that quickly outpace any budget built on prior usage. This isn’t a forecasting weakness that better spreadsheets will fix. Token usage for the same task can vary by up to 30 times in agentic coding workflows, according to Stanford Digital Economy Lab research cited in the report.

Consumption pricing removed the safety net. Seat-based software pricing hid true usage cost behind a flat fee. As large language model providers move to consumption pricing, that safety net disappears, and organisations now pay directly for every high-usage scenario they create.

Governance maturity hasn’t caught up. There is, in most organisations, no tagging structure, no sourcing strategy and no playbook for AI spend specifically. Without these, forecasting, allocation and control at scale simply aren’t possible.

Model and tool selection defaults to the expensive option. When the trade-offs between quality, latency, risk and cost aren’t clear, users default to premium models or familiar tools rather than the cheapest option that would do the job.

Citizen developers are quietly multiplying consumption. Employees can now build AI-powered workflows and autonomous agents with minimal engineering effort, outside central IT. This democratises innovation, but it also means unmanaged token consumption is growing in places finance and IT aren’t watching.

The scale of the blind spot

The headline numbers describe a problem that gets bigger with adoption, not smaller. As organisations move from isolated use cases to enterprise-wide deployment, AI spend increases nearly fourfold, and 62% of organisations have already moved past experimentation into active deployment. Most firms surveyed expect spend to increase by at least a further 25% over the next 12 months.

The visibility problem underneath that growth is specific and quantifiable. McKinsey’s experience across engagements has found that 20 to 30% of AI spend is typically unaccounted for, fragmented across cloud providers, foundation-model vendors, software platforms, experimentation environments and individual business units, with no single source of truth. In one organisation cited in the report, an exercise to consolidate spend visibility uncovered costs spread across enterprise copilots, foundation-model contracts, AI-enabled software features, API-based services, experimentation environments and business-unit purchases. What began as a routine budget exercise turned into the discovery of an entirely fragmented, previously invisible portfolio.

FinOps maturity across the four capabilities that would catch this early is consistently low: only 25% of organisations have mature spend visibility and attrition tracking, 24% have mature token consumption tracking, 21% have mature cost allocation (chargeback or showback), and just 20% have mature forecasting and budgeting. Organisations with high forecasting maturity specifically save 10% more on AI spend than their peers on average, which makes the 20% maturity figure the single most important number in the report for anyone deciding where to invest first.

Where the money actually goes: the optimisation levers that work

McKinsey’s analysis identifies roughly 40 levers that drive the best outcomes in AI spend optimisation, but six categories account for most of the impact, and about a third of organisations surveyed have already captured savings of 20 to 30% through active optimisation:

  • Model routing. Few tasks warrant an expensive frontier model. Routing each task to the lowest-cost model that can still deliver the required quality is the single highest-leverage lever available.
  • Reducing bloated token inputs. Shortening prompts, limiting context windows, passing only relevant document or tool-output sections, and summarising conversation history before resending it to the model. Prompt caching alone, reusing static prompt context, can cut repeated input-token costs by up to roughly 90%, particularly for RAG systems and agents with large, stable prefixes.
  • Controlling output-token consumption. Defining expected response formats, using structured outputs, and capping response length by use case so models stop generating text nobody reads.
  • Managing agent and workflow volume. Setting limits on retries, tool calls, agent loops and escalation paths, and redesigning repetitive or poorly structured workflows upstream rather than simply adding more agents to a broken process.
  • Batching and caching. Batching non-urgent, high-volume requests where latency doesn’t matter, and caching repeated prompts, stable context and reusable intermediate results.
  • Selective use of open-weight models. Treating open-weight models as a genuine alternative to frontier models for simpler tasks that don’t need the most expensive reasoning capability available.

Organisations that apply these levers thoughtfully save 20 to 30% on AI cost overall, according to McKinsey’s experience, by developing the capability to optimise spend, improve accountability, and redirect the resulting savings toward the highest-value opportunities rather than simply banking them.

The control plane: making spend observable, attributable and governable

McKinsey’s central architectural recommendation is the AI control plane: an enterprise management layer that sits between users, applications, agents and the models or platforms they consume. The analogy the report draws is a useful one for a board conversation: if ERP systems became the system of record for financial transactions, the AI control plane may become the system of record for intelligence consumption.

It operates across three vectors. Visibility and attribution capture request-level telemetry across users, applications, workflows, models, prompts, agents, tokens, API calls, latency and cost, tagged to a business unit, product, use case, workflow, owner and cost centre, so organisations can move from aggregate vendor invoices to genuine unit economics: cost per task, per case, per code review, per customer interaction. Policy and governance enforce rules in real time, including approved model access, role-based permissions, token budgets, spend thresholds, context-window limits and premium-model approvals, catching avoidable spend before it happens rather than in a retrospective invoice review. Routing and optimisation direct workloads to the right model or provider based on cost, quality, latency, risk and availability, and surface opportunities like caching, prompt standardisation, agent-loop reduction and model right-sizing.

The unit of governance, McKinsey is explicit on this point, should be the completed business outcome, not the token cost. A control plane exists to connect the two, through showback and chargeback mechanisms that make AI consumption visible against the business activity that generated it.

Sourcing has fundamentally changed, and buy-versus-build is the wrong question now

Traditional software sourcing was built for a predictable world: seat-based licensing, single-vendor strategy, fixed demand forecasts, long-term commitments, periodic benchmarking. AI sourcing runs on the opposite logic: consumption-based pricing, a multimodel ecosystem, dynamic demand management, flexible commercial structures and continuous benchmarking. Applying the old sourcing playbook to the new consumption model is, per McKinsey’s data, one of the more common reasons organisations lose control of spend.

Leading organisations are seeing unit cost reductions of 10 to 20% through two specific levers: building an AI accountability model that gives visibility into who is consuming AI resources and what value that consumption is generating, and moving past the binary buy-versus-build question entirely. With proprietary frontier models, provider-hosted and enterprise-hosted open-weight models, and edge or local models all now viable for different workloads, the real sourcing question is finding the right mix of buy, build, host, route and switch, continuously re-evaluated against performance, cost, risk and business value, rather than settled once at procurement.

What CIOs should prioritise now

McKinsey’s recommendation, consistent with the report’s overall argument, is not to cut spend reflexively. The organisations creating the most value are the ones building the capability to forecast demand, govern consumption and continuously optimise, not the ones simply switching things off.

Six actions define that approach:

  1. Build forecasting capability before costs erupt. Model demand under different adoption, pricing and workload scenarios, and include projected consumption curves, sensitivity analyses and cost-per-outcome metrics in every AI business case. Organisations with high forecasting maturity already save 10% more than peers, and only 20% currently have this capability.
  2. Move to a business-value accountability model. Connect AI usage directly to the products, workflows and business units generating demand, and evolve measurement toward business metrics like cost per claim processed or revenue per AI-enabled workflow, rather than raw token counts.
  3. Anchor initial AI consumption capability around the highest-value workflows, where impact will be greatest and where the resulting savings can be reinvested elsewhere.
  4. Maintain flexibility in sourcing and technology. Today’s optimal model or vendor is unlikely to be optimal in six months. Architectures should support multiple model providers and the ability to shift workloads between proprietary and open-weight models without a re-platforming exercise.
  5. Build a permanent AI FinOps capability before spending reaches scale, staffed as an ongoing cross-functional team responsible for forecasting, monitoring, sourcing evaluation and cost-per-outcome measurement, not a temporary project office that disbands once the initial cost spike is dealt with.
  6. Embed governance directly into the architecture. As adoption expands, automation is the only practical way to govern consumption at scale. AI gateways, control planes, policy engines and automated guardrails should route requests to lower-cost models by default, enforce budget thresholds, and escalate automatically when cost or risk exceeds a predefined limit, so the economically efficient choice is the default choice rather than something users have to remember to make.

Why this is a Phase Zero problem again

The pattern here runs parallel to what we see across enterprise AI governance generally, and to the identity and access findings we covered in our recent piece on the agentic IAM governance gap: organisations that build the visibility and accountability foundation before scale consistently outperform organisations trying to retrofit control after the spend has already happened.

McKinsey’s own data makes the case directly. Forecasting maturity, not spend cuts, produces the 10% saving. Optimisation discipline, not blanket restriction, produces the 20 to 30% saving. A control plane, not a policy memo, produces the accountability.

This is precisely the sequencing argument behind Phase Zero: establish the tokenomics foundation, the control plane, the forecasting discipline and the named accountability, before the fourth or fifth agentic workflow goes live, rather than discovering the fragmented spend picture after the budget has already been exhausted mid-year.


FAQ

What percentage of companies exceed their AI budget? 93% of organisations exceed their AI budget, according to McKinsey’s Enterprise AI FinOps survey. 46% are over budget by 10 to 30%, and 8% are over by more than 30%.

Why is AI spend harder to manage than traditional cloud or software spend? AI usage is fundamentally less predictable: the same task can generate dramatically different token volumes and trigger different agent chains, and token usage for identical tasks has been shown to vary by up to 30 times. Combined with a shift to consumption-based pricing and low FinOps maturity, AI removes the safety net that seat-based software pricing used to provide.

How much of enterprise AI spend typically goes untracked? McKinsey’s experience across engagements shows 20 to 30% of AI spend is often unaccounted for, fragmented across cloud providers, foundation-model vendors, software platforms, experimentation environments and individual business units with no single source of truth.

What is an AI control plane? An AI control plane is the enterprise management layer that sits between users, applications, agents and the AI models they consume, making usage observable, attributable and governable across visibility and attribution, policy and governance, and routing and optimisation. McKinsey compares its likely role to what ERP systems became for financial transactions.

How much can companies save by optimising AI spend? Companies that actively optimise their AI consumption typically save 20 to 30%, according to McKinsey’s analysis, primarily through model routing, prompt caching (which can cut repeated input-token costs by up to roughly 90%), output-token controls, and reducing unnecessary agent loops.

What should a CIO do first to get AI spend under control? Build forecasting capability before costs erupt. Only 20% of organisations currently have mature forecasting and budgeting practices for AI, yet organisations with high forecasting maturity save 10% more than their peers on average, making this the single highest-leverage starting point.

Stay informed on all things AI...

Join Our Webinar Cloud Migration with a twist

Aug 18, 2022 03:00 PM BST / 04:00 PM SAST