AI Economics Series Article 4B | The Margin Trap: Why Cheaper Tokens Don’t Mean Cheaper AI
A companion to Article 4A, The Great AI Capex Gamble, which closed by asking whether intelligence can be monetised at scale before the bill comes due. This piece looks at where the margin will actually settle.
Artificial intelligence can become far cheaper and still disappoint investors.
That sounds like a contradiction. For decades the technology industry has been trained to see falling unit costs as the path to software-like margins. Build once, distribute endlessly, and let scale do the rest. Many early AI business cases borrowed that logic: train the model, serve more customers, spread the cost across a larger base, then wait for the margin curve to bend upwards. AI is not following the script.
A unit of intelligence is getting cheaper. Supplying intelligence at industrial scale is not. The system behind it is capital-heavy, energy-hungry, chip-constrained and increasingly competitive on price. That gap may decide who makes money from AI, who merely funds its expansion, and who discovers too late that high usage is not the same as high margin.
Stanford’s 2025 AI Index captures the speed of the cost decline. It found that the inference cost of a system performing at GPT-3.5 level fell more than 280-fold between November 2022 and October 2024. It also recorded annual hardware cost declines of around 30%, annual energy-efficiency improvements of around 40%, and a narrowing gap between open-weight and closed models on some benchmarks. (Stanford HAI, 2025 AI Index)
A more recent study drawing on pricing data from Artificial Analysis and Epoch AI sharpens the picture. Tracking what it costs to reach a given benchmark score over time, its authors find that price falls by five to ten times per year for frontier models across knowledge, reasoning, mathematics and software-engineering benchmarks. They flag the opposite pull as well. Because newer frontier systems are larger and lean harder on multi-step reasoning, the raw expense of running them is climbing too, by a factor of roughly 3 to 18 times per year depending on the model. (Gundlach et al., The Price of Progress, arXiv)
So there is a paradox. AI keeps getting cheaper in capability-per-dollar terms, while the most ambitious AI systems may still become more expensive to operate in absolute terms. Enterprise buyers can only welcome that. For model providers, infrastructure owners, application vendors and investors it is less comfortable, because a product can get cheaper for two different reasons: scaling, or commoditising.
Intelligence is being repriced
The repricing is already visible in public model price cards. The figures below are standard API prices as of 12 August 2026, per million tokens. OpenAI’s GPT-5.6 Sol is priced at $5 for input and $30 for output, with cached input materially lower and higher rates for long-context requests. (OpenAI GPT-5.6 Sol model documentation; OpenAI API pricing) Google’s Gemini pricing shows the same laddering: Gemini 3.1 Pro Preview at $2 for input and $12 for output for prompts up to 200,000 tokens ($4 and $18 above that), and Gemini 3.5 Flash-Lite at $0.30 and $2.50 respectively. (Google Gemini API pricing) Anthropic’s Claude pricing spans premium, mid-tier and lower-cost models: Claude Opus 5 at $5 for input and $25 for output, Claude Sonnet 5 at $2 and $10, and Claude Haiku 4.5 at $1 and $5, with batch processing offering a 50% discount on input and output tokens. (Anthropic Claude pricing)
The numbers will keep moving, so the exact prices matter less than the shape of the market: a premium tier, an efficient mid-tier, a low-cost scale tier, and a growing open-weight alternative. Buyers can route work between them, and vendors know it. Pricing power will be hard to defend unless capability, workflow ownership or trust is properly differentiated.
Nor is this mere vendor discounting. It reflects technical progress: smaller models, distillation, mixture-of-experts architectures, better inference stacks, prompt caching, batch processing, smarter routing and higher utilisation of accelerated compute. Intelligence is becoming a more liquid input, and the management question changes with it. It is no longer “can we afford AI?” but “which form of intelligence, at what level of quality, under what control model, for which business outcome?” That is where margin is decided.
The margin trap
Falling token prices can seduce executives into the wrong conclusion. In classic software, lower marginal cost improves margins: more usage means better economics. AI is messier. Every inference still consumes compute, and more advanced reasoning consumes more tokens. Agentic systems search, plan, retry, call tools, read documents, maintain context and sometimes loop through failed attempts before producing anything useful. Multimodal work adds audio, image and video costs. Enterprise deployment adds logging, monitoring, evaluation, access controls, audit trails and human exception handling.
So the price of a token may fall while the number of tokens required to complete a business task rises. Trade coverage through 2026 had already put a name to that pattern: The Information tracked organisations “tokenmaxxing” their way through reasoning models and agentic workflows under a near-identical heading, the AI margin squeeze. (The Information — ‘Tokenmaxxing’ and the AI Margin Squeeze) The mechanism is the one this piece works through in detail: the unit cost of intelligence declines, but the operational system around intelligence becomes more complex and more heavily used.
The industry’s capital commitments make the squeeze more visible. Reuters reported in February 2026 that Alphabet, Amazon, Meta and Microsoft were expected to invest about $650 billion in AI-related infrastructure in 2026, up from $410 billion in 2025, based on Bridgewater Associates analysis. (Reuters) McKinsey estimates that data centres worldwide will require about $6.7 trillion in capital outlays by 2030 to keep pace with demand for compute, including about $5.2 trillion for AI-processing data centres. (McKinsey)
That capital has to be serviced: depreciated, powered, cooled, connected, secured and utilised. Falling prices help demand grow, but they also make it harder to earn attractive returns unless usage expands faster than pricing falls.
Scarcity margins do not last forever
For now, parts of the lower stack still enjoy scarcity economics. NVIDIA’s fiscal 2026 results show why: full-year revenue of $215.9 billion, up 65% year on year, with fiscal-year GAAP gross margin of 71.1% and fourth-quarter GAAP gross margin of 75.0%. (NVIDIA newsroom) Ordinary infrastructure does not earn margins like that. Scarcity does.
But scarcity attracts substitution. Hyperscalers are building custom silicon while model providers optimise inference. Enterprises are introducing routing layers. Open-weight models are becoming good enough for more practical workloads, and prompt caching and batch processing chip away at the need for expensive live inference. Procurement teams, meanwhile, are learning to compare price, latency, quality and data-control trade-offs far more aggressively than they did a year ago.
The model layer will not disappear. Frontier models will remain strategically important where performance genuinely matters. That means complex reasoning, software engineering, multimodal analysis, planning, research and high-value judgement. Most enterprise work needs reliable, governed, adequate intelligence at the lowest sustainable cost, not the most expensive intelligence available.
That creates at least three markets, not one. At the top sits premium frontier intelligence: expensive, but defensible when the task demands better reasoning, deeper context or lower error tolerance. Below it is the efficiency layer, made up of smaller proprietary models, mini models, flash models, haiku models and specialists handling high-volume work. That layer will be intensely competitive, because buyers can test performance and switch. Third comes the open-weight and sovereign layer, which enterprises, public-sector bodies and regulated firms will increasingly consider where data residency, hosting control, customisation, auditability, resilience or vendor dependency matter. Open-weight models will not displace frontier systems everywhere, but they will cap pricing power in many routine and sensitive workloads.
The profit pool will therefore be uneven. Chips, cloud, models, applications, data owners and workflow platforms will all compete for margin. Some will own scarcity, some distribution, some trust. Others will discover they are reselling a commodity with a nice interface.
Jevons comes for AI
Cheaper intelligence will create more demand. That is the optimistic half of the story. The half many business cases miss is that the demand may consume the savings. A workflow that once made one model call may make twenty. A simple assistant becomes an agent that retrieves policy, checks systems, drafts a response, validates it, logs its evidence and escalates uncertainty. Customer-service tools graduate from answering questions to investigating cases. Coding assistants go from autocomplete to planning, testing, debugging and generating pull requests. Compliance assistants move from summarisation to control testing and exception routing.
The organisation sees a better capability. The cost system sees a longer chain of work.
This is why token price is the wrong economic unit. The better measure is cost per resolved case, per decision supported, per control tested, per client interaction, per accepted production-grade code change. The firms that win will engineer the cheapest reliable outcome rather than buy the cheapest model.
Enterprise AI needs a margin architecture
Every serious AI programme now needs an economic architecture sitting beside the technical one, built around a single question: what is the least expensive system that produces the required business outcome at the required level of quality, risk, latency and control?
That points towards a portfolio. Reserve frontier models for scarce, high-value reasoning. Let mid-tier models handle repeatable knowledge work, and small models classify, extract, route and summarise. Test open-weight models where sovereignty, privacy, customisation or cost control change the equation, and deploy specialist models where the task is narrow enough to justify them.
Cost-aware architecture then becomes a source of margin. Prompt caching, semantic caching, batching, shorter context windows, retrieval optimisation, model routing, fallback logic, evaluation gates and human-in-the-loop thresholds look like technical housekeeping. They are economic levers. A poorly designed AI system turns cheap intelligence into expensive automation. A well-designed one turns the same intelligence into operating leverage.
The application squeeze
Application vendors face a different version of the same problem. Summarisation, drafting, search, chat interfaces, workflow assistance and basic analytics can now be bolted onto almost any product. Those features may improve the product without supporting premium pricing for long, because competitors can reproduce them with the same underlying models.
The risk is highest for thin AI wrappers: little proprietary data, weak workflow ownership, limited distribution, no regulated trust. If the model provider builds the feature into its own interface, or the platform embeds it natively, the wrapper loses oxygen. Falling token prices may reduce its cost base, but falling differentiation can reduce its pricing power faster.
Durable application economics will come from harder assets, among them workflow depth, proprietary data, embedded distribution, domain trust, regulatory permission and measurable outcomes. The margin sits in converting model capability into process ownership rather than merely exposing it. That is why enterprise AI value is likely to migrate towards the system layer, meaning model plus data plus workflow plus controls plus behaviour change. The model matters. The system captures the value.
The case against the squeeze
The trap thesis is not the only plausible outcome, and the strongest version of the alternative deserves stating. The same research cited earlier that shows running costs climbing 3 to 18 times a year for the largest reasoning systems also shows unit efficiency improving 5 to 10 times a year. Some operators build cost governance in from the start, routing to smaller models by default, reserving frontier reasoning for the cases that need it, caching and batching aggressively. For them, efficiency gains can outrun volume growth. On that reading the margin squeeze says less about AI economics generally than about undisciplined deployment, and it should ease as the market’s better operators pull away.
Heavy capital intensity cuts two ways. It squeezes near-term returns, but it also raises the barrier to staying at the frontier. As smaller model providers run out of capital or get absorbed into larger ones, the number of firms able to compete on frontier capability narrows, and pricing power tends to concentrate rather than dissolve. That is the pattern capital-intensive industries have generally followed once a shakeout runs its course.
Cloud computing offers a precedent for both dynamics. Hyperscalers spent the better part of a decade absorbing thin margins on infrastructure build-out before cloud became one of their most profitable segments: AWS posted an operating margin of 37.7% in the first quarter of 2026, on operating income of $14.2 billion against $37.6 billion in segment revenue. (Amazon — First Quarter 2026 Results) A return like that would have looked implausible during AWS’s own early build-out years. If AI infrastructure follows a similar maturation curve, with utilisation improving and software layers adding margin on top of increasingly commoditised hardware, today’s squeeze may prove to be a phase rather than a ceiling.
None of this changes the argument for most enterprises building on top of these systems today. If consolidation and margin recovery arrive, they will most likely accrue to whichever infrastructure and model providers survive the shakeout, not to the businesses buying their tokens. The bull case describes a reasonable path for the industry’s largest suppliers. Buyers face a different equation, and the discipline this piece argues for still has to be built regardless of how the supply side eventually settles.
The board question
Boards should not spend much time asking whether AI is getting cheaper. It is. The live question is where the margin will settle: with the chip provider, the cloud provider, the model company, the application vendor, the owner of proprietary data, whoever controls distribution, or the enterprise that redesigns the workflow and can deliver a trusted outcome in a regulated environment.
Falling intelligence costs will not distribute value evenly. They may widen the gap between firms that merely consume AI and firms that convert it into proprietary operating advantage. For most enterprises the goal is not to become a model company. It is to become an AI-margin company: an organisation that repeatedly turns cheaper intelligence into lower cost, faster decisions, better controls, improved customer outcomes and new revenue, without letting usage growth overwhelm the business case.
That takes discipline. Every AI use case should define the value driver, unit cost, model route, consumption pattern, human exception rate, operating-control model and scaling assumption. Otherwise pilots will continue to look impressive and production will continue to disappoint.
Exhibit: the AI margin equation AI gross margin is not revenue minus token cost. It is closer to: Revenue per business outcome minus model inference minus retrieval minus orchestration minus data movement minus monitoring minus compliance minus human exception handling minus depreciation or cloud margin. This is why the margin squeeze matters: token prices can fall while total delivery cost stays stubborn. |
The next phase of AI economics
The first phase of generative AI was about capability. Could the model write, reason, code, summarise, search and converse? The second was about adoption, and whether organisations could find useful applications and persuade employees to use them. The third phase is about economics: can AI scale without damaging the business case that justified it?
Many strategies will struggle here, with good demos, enthusiastic users and rising consumption but weak unit economics. The stronger firms will build cost governance into the architecture from the beginning. They will know when to use frontier intelligence, when to route to cheaper models, when to host privately, when to cache or batch, and when not to automate at all.
Intelligence will become cheaper. That looks increasingly likely. Profitability will depend on something less automatic: scarcity, differentiation, utilisation and control. Cheaper does not mean free, and usage does not mean margin. Intelligence becomes profitable only when it is embedded in a system that captures value.
Access to AI is becoming table stakes. The next competitive advantage is economic control over it.
Sources
Stanford HAI — 2025 AI Index Report
Gundlach et al. — The Price of Progress (arXiv)
OpenAI — GPT-5.6 Sol Model Documentation
Reuters — Big Tech to invest about $650 billion in AI in 2026, Bridgewater says (23 February 2026)
McKinsey & Company — The cost of compute: A $7 trillion race to scale data centers
NVIDIA — Financial Results for Fourth Quarter and Fiscal 2026
The Information — ‘Tokenmaxxing’ and the AI Margin Squeeze
Amazon — Amazon.com Announces First Quarter Results (Q1 2026)