AI Economics Series Article 2 | The Unit Economics of Enterprise AI

Share
đź’ˇ
Why read this

Learn that the next competitive advantage will belong to organisations that can measure, govern and optimise the cost of intelligence.

For the past two years, most executives have asked one question about artificial intelligence: what can it do? That was the right question during the experimentation phase. The question that matters now is what it costs to do it repeatedly, safely and at enterprise scale.

Capability is accelerating. Stanford’s 2026 AI Index reports organisational AI adoption at 88%, and performance on SWE-bench Verified, a coding benchmark, rose from 60% to near 100% in a single year. AI has not become perfect — but capability is no longer the scarce commodity it was two years ago. Access to powerful models is widely available. The scarce capability is managerial: knowing where AI creates value, what it costs to operate, and when its economics break.

This is the second phase of the AI economy. The first ran on possibility; this one runs on unit economics.

From model access to economic discipline

The early enterprise AI conversation was dominated by model choice — which provider, which benchmark, which assistant. That discussion still matters, but a model is not an operating model, and a proof of concept is not a business case.

Enterprises are discovering that AI behaves less like traditional software and more like a variable-cost production system. Every prompt, retrieval call, reasoning step, tool call, agentic loop and human review step carries a cost. Some of those costs sit visibly on invoices; others hide in cloud platforms, integration teams, risk controls, rework, data preparation and delayed process redesign.

AI expenditure is also scaling before many organisations have developed the financial-management capability to govern it. Gartner forecasts worldwide AI spending—including infrastructure, services, software and models—at $2.59 trillion in 2026, up 47% year over year. Because infrastructure represents a large share of that total, the figure should not be read simply as enterprise application spending. It does, however, illustrate the scale of the economic system now forming around AI.

Agentic AI raises the stakes further. A chatbot answers; an agent acts — and acting means planning, retrieval, reasoning, tool use, validation, retry logic and sometimes escalation. A single request becomes a chain of computational and operational events, so the cost of AI rises not only with the number of users but with the complexity of the work delegated to it.

The hidden problem: AI cost is not linear

In traditional enterprise technology, scale usually improves the economics — more users, better amortisation of fixed costs. AI does not behave that way. Some costs decline with scale, while others rise with usage, complexity and governance burden.

Pricing itself is part of the challenge. OpenAI’s public GPT-5 pricing varies widely by tier: GPT-5, GPT-5 mini and GPT-5 nano carry materially different input and output token costs. That is less a criticism of one vendor than evidence of a broader reality — AI cost is multi-dimensional, usage-sensitive and highly dependent on design choices.

The biggest mistake organisations make is treating token cost as the whole picture, when it is only the visible tip of the stack. The real cost of an AI application is shaped by how often the capability is used, which model handles which task, how much context and history gets passed in, and how many reasoning steps and agentic loops — retries, searches, tool calls, delegated sub-tasks — sit behind each answer. On top of that come the costs that never appear on a model invoice: connecting AI to enterprise systems and data, the control burden of validation, human review, auditability and compliance, and the rework generated by hallucinations, poor outputs and wrong assumptions.

So “what is the price per million tokens?” is the wrong opening question. It is like assessing a trading platform by the price of compute cycles while ignoring market data, controls, latency, exception handling, regulatory reporting and operational risk.

The emerging divide: adopters versus economic operators

A distinction is emerging between organisations that deploy AI and organisations that operate it economically.

Read against those numbers, most AI investment cases should not be priced as full automation. They should be priced as assisted, controlled, partially automated operating models. A fully automated process might support a large headcount-reduction assumption; a human-in-the-loop process delivers speed, quality, capacity and risk benefits, but not the same labour-substitution economics. Treating the second as if it were the first produces inflated ROI, disappointed boards and, eventually, funding fatigue.

AI is becoming an operating cost, not an innovation budget

The next mistake is organisational. Many firms still manage AI through innovation budgets, transformation portfolios or technology experiments. That made sense when use cases were isolated. It will fail as AI enters everyday work — because once AI is embedded into sales, engineering, operations, customer service, compliance, risk and finance, it becomes an operating cost that needs forecasting, allocation, controls and optimisation.

This is where AI economics becomes a board issue. Not because every director needs to understand tokenisation, but because boards need to know whether the organisation can answer five basic questions.

Where is AI spend accumulating? Not only by vendor, but by business unit, application, process, user group, model and workflow.

Which usage is creating measurable value? Usage volume is not value. A million prompts may signal adoption, waste or both.

Which use cases have positive unit economics? A use case that looks impressive in a demo can turn uneconomic at thousands of users, long documents, expensive models and mandatory review.

What stops runaway consumption? Agentic systems can burn resources through retries, searches, tool calls and long-context reasoning. Consumption needs guardrails.

And who owns the economics? If finance owns the budget, technology owns the platform, risk owns the controls and the business owns the process, no one owns the full equation.

Organisations that can answer these make better investment decisions: they stop funding impressive but low-value pilots, route simple tasks to cheaper models, reserve expensive reasoning for high-value work, cut duplicated context, build reusable components, and price governance into the operating model from day one.

The Four-Layer Enterprise AI Economics Model

Building on emerging AI FinOps practice, I propose a four-layer model for enterprise AI economics: consumption, orchestration, control and value.

Every serious enterprise AI initiative now needs a unit economics model with at least four layers.

1. The consumption layer

The direct cost of model usage: input, output and cached tokens, context windows, audio, image, video, embeddings, retrieval and batch processing. It is the most measurable layer, and the easiest to underestimate.

The management question here is less “which model are we using?” than “what is the cheapest, most reliable and most governable model pathway for this class of work?” For high-complexity reasoning, frontier proprietary models (latest as of publication date) -          OpenAI’s GPT-5.6 range, Anthropic’s Fable 5, Opus 4.8 and Sonnet 5, and Google’s Gemini 3.5 generation alongside Gemini 3.1 Pro — may justify premium consumption where the task genuinely requires superior reasoning, coding, multimodal capability, long-context analysis or agentic tool use. OpenAI positions GPT-5.6 as its most advanced model for coding and agentic tasks; Anthropic describes Claude Fable 5 as its most capable widely released model, with Claude Opus 4.8 suited to complex agentic coding and enterprise work; Google calls Gemini 3.1 Pro its most advanced model for complex tasks, handling text, audio, image, video and code-repository inputs.

The selection pathway should also include open-weight alternatives, particularly where cost control, customisation, data sensitivity, latency, resilience or regulatory sovereignty matter. Meta’s Llama 4 Scout and Maverick are open-weight, natively multimodal models; Mistral Large 3 is an open-weight, general-purpose multimodal and multilingual model; DeepSeek-R1 and DeepSeek-V3 are positioned as open-source models with strong reasoning and competitive performance; and Google frames Gemma 4 as one of its most intelligent open-model families.

Put together, this is a more sophisticated economic architecture: frontier models for scarce, high-value reasoning; smaller or cheaper proprietary models for routine scale workloads; and open-weight models hosted in sovereign or private environments where the organisation needs greater control over the model, the data, the inference environment and the evidence trail. Sovereign AI is increasingly defined as the ability to build and run AI using local or controlled infrastructure, data, talent and business networks — NVIDIA defines it as a nation’s capacity to produce AI using its own infrastructure, data, workforce and business networks. Cloud providers also offer regional data-routing patterns that can help meet jurisdictional requirements. Those options are not economically neutral: Anthropic, for example, documents a pricing premium for certain regional and multi-region endpoints. Sovereignty and residency should therefore be costed as design requirements rather than treated as free configuration choices.

In mature AI organisations, model choice becomes dynamic. Simple classification, extraction and summarisation tasks should not automatically run on frontier reasoning models; expensive models earn their place where incremental accuracy, reasoning or risk reduction justifies the cost; open-weight models come into play where workload volume, sensitive data, jurisdictional control or fine-tuning changes the economics. The consumption layer stops being about price per token and becomes about designing a model portfolio that balances capability, cost, sovereignty, resilience and control.

2. The orchestration layer

The cost of turning model calls into workflows: agents, planning, tool use, memory, retrieval-augmented generation, workflow engines, API calls, logging, monitoring and exception handling.

This is where cost escalates quickly. A user sees one answer; the system may have run ten searches, read twenty documents, called three tools, weighed alternatives and retried failed steps. Without telemetry, the organisation cannot see what that answer actually cost.

Agentic systems require particular attention because one visible outcome may conceal many searches, model calls, tool invocations and retries. In a 2026 study of agentic coding tasks, Bai and colleagues found token use dramatically higher than in conventional code-chat and reasoning tasks, with substantial variation between repeated runs of the same task. The results should not be generalised mechanically to every agent, but they show why agentic applications need task-level telemetry, cost prediction and explicit consumption budgets before they are scaled.

3. The control layer

Human review, approval workflows, model evaluation, audit trails, policy checks, data-loss prevention, red-teaming, compliance testing, security review and operational resilience. In regulated industries the control layer is part of the product, not an optional extra. And an application that cannot be explained, audited or controlled is not cheap — it is deferred risk. 

4. The value layer

The numerator in the ROI equation: revenue uplift, cost avoidance, productivity gain, working-capital benefit, risk reduction, cycle-time compression, customer retention, error reduction and management capacity.

This layer is often the weakest, because organisations measure activity rather than outcome. They count users, prompts, pilots and generated documents; far less often do they measure whether the process became faster, cheaper, safer or more profitable.

The discipline of AI economics is the discipline of connecting all four layers. Consumption without value is waste; value without controls is fragile. Controls that lack cost visibility turn into expensive bureaucracy, and orchestration nobody owns sprawls into technical debt.

The infrastructure constraint will feed back into enterprise economics

AI economics is not confined to enterprise invoices. It is also shaped by the physical infrastructure required to deliver intelligence: chips, data centres, power, cooling, grid capacity and specialist equipment.

The International Energy Agency projects that global data centre electricity consumption will double to around 945 TWh by 2030 in its base case — growing by around 15% per year from 2024, more than four times faster than total electricity consumption from all other sectors. The supply-chain pressure is already visible: Reuters reported on 9 July 2026 that AI data centre demand is intensifying shortages of critical US grid equipment — transformers, circuit breakers, switchgear — and cited Wood Mackenzie projections that US data centre capacity could rise from around 24 GW today to 110 GW by 2030.

The practical consequence is that AI costs will not be set by model competition alone. Even if model prices keep falling, the total cost of reliable, compliant, low-latency, regionally hosted AI may not fall at the same rate. Software abundance and physical constraint will shape the economics together.

Extending FinOps into enterprise AI

Cloud computing created the need for FinOps because elastic infrastructure challenged traditional budgeting and ownership models. The FinOps discipline has since expanded into AI cost and value management. The enterprise challenge is now to apply those principles below the vendor-invoice level—to models, prompts, agents, workflows, controls and business outcomes.

The required capability is a management system that forecasts, observes, allocates and optimises AI expenditure against measurable value. It should include a complete cost taxonomy, task-level unit costs, model-routing policy, agent-consumption guardrails, business-value telemetry, investment controls and post-implementation reviews.

It should include:

A cost taxonomy separating model cost, platform cost, orchestration cost, data cost, integration cost, control cost and human operating cost.

A unit-cost model for each AI application — cost per task, per completed workflow, per customer interaction, per document processed, per decision supported, per exception.

A model-routing strategy that uses cheaper models where possible, reserves expensive reasoning for high-value or high-risk work, and deliberately evaluates open-weight, sovereign-hosted options where data control, cost or resilience change the business case.

Consumption guardrails limiting runaway agent loops, excessive context, unnecessary retries and inappropriate use of premium models.

Business-value telemetry that measures outcome, not just usage.

Investment governance requiring AI initiatives to show expected consumption, value drivers, risk controls, scaling assumptions and break-even points before funding.

Post-implementation reviews comparing forecast usage and value against what was consumed and realised.

None of this is bureaucracy for its own sake. It is how AI moves from experimentation to industrialisation.

The strategic opportunity

Organisations that master AI unit economics gain three advantages. They scale faster, because they know which use cases deserve capital and do not have to rely on storytelling, vendor claims or isolated productivity anecdotes. They operate cheaper, because their systems are designed for economic efficiency from the start — no premium models for commodity tasks, less redundant context, reusable prompts and components, cost-aware agents, and a model portfolio rather than a single-model default. And they govern better, because cost visibility and control visibility reinforce each other: the same telemetry that shows spend also shows workflow behaviour, exception rates, escalation patterns and risk exposure.

The next wave of competitive advantage will come from operating intelligence as an economic resource — which is a very different thing from merely having AI.

Conclusion: intelligence is becoming measurable

The AI economy is entering a more serious phase. The winners will not be the organisations with the most pilots, licences or innovation narratives, but the ones that understand the economics of intelligence — that ask harder questions before funding initiatives, measure cost at the level of work rather than vendors, and treat governance as part of the economic model rather than an afterthought. They will know that agents are not free labour but metered systems with variable consumption and control requirements, and that frontier, smaller proprietary and open-weight models belong in a single economic architecture. Above all, they will recognise that AI value is created not when a model produces an output, but when a business process becomes measurably better.

The board-level question has moved on from whether the organisation has an AI strategy to whether it understands the unit economics of that strategy.


Sources

Stanford HAI — 2026 AI Index Report

Gartner — Forecasts Worldwide AI Spending to Grow 47% in 2026 (May 2026)

KPMG — Global AI Pulse Survey, Q2 2026

KPMG UK — AI adoption accelerates in the UK (July 2026)

Bain & Company — Your AI Budget Is Growing. Your Returns Aren’t. Here’s Why.

Reuters — US power companies scramble to secure equipment as surging data center demand strains supplies (9 July 2026)

IEA — Energy and AI

Read more