Global spending on AI (artificial intelligence) products is hitting a shocking $2.52 trillion per year, and its unexpected costs are breaking corporate budgets. Recent reports and data from CloudZero report that the rapidly increasing spending on AI tools has severely strained company balance sheets; one in five organizations are overshooting their projected infrastructure spending by 50% or more.
The Shift from Training to Inference
While the initial technology boom focused mainly in the multi-billion dollar training of these AI models, the real long-term economic reality lies in inference, the ongoing, live deployment of those networks to answer everyday user input and requests, or “prompts.” This shift is a pivot from traditional Software-as-a-Service (SaaS) subscription models into the territory of Inference Economics. In this new region, technology companies are moving from software providers to compute wholesalers, measuring and charging for each machine-generated “thought.”
Understanding Inference
In economics, an AI “thought” represents a discrete, resource-heavy unit of computation called inference. Unlike other software that pulls pre-written code from a cloud database, artificial intelligence generates every response by dynamically calculating probabilities across its neural network (basically, the AI’s “brain”). The components of these automated thoughts are tokens, the pieces of data representing words, pixels, and code that the AI reads and generates. Because every single token requires specialized graphic processing units (GPUs) to process complex floating-point math, every generated response has a physical expense.
The Economic Shift in SaaS Pricing
This tangible cost floor is causing a structural shift for AI companies creating SaaS products. Companies like Salesforce and Zendesk are moving towards a usage or output based fee, pivoting from the “per-seat” model to align their revenues with their expenses. For example, Zendesk now charges around $1.50 for every customer problem that its AI solves completely. By billing based on conversation or output, these software firms are treating AI as a digital unit of labor rather than a static tool. This shift highlights an economic shift: AI is removing one of software’s key advantages, the near-zero cost of creating copies.
The Marginal Cost of AI
Because an AI-generated output requires real, physical resources, it has a noticeable marginal cost, the additional expense that a company incurs to produce one more unit of output. Every time Claude or ChatGPT answers a prompt, it draws on massive amounts of electricity from local power grids, uses water to cool data centers, and goes through expensive computing chips.
Test-Time Compute
Furthermore, the newest models have introduced a concept called the test-time compute. Instead of outputting a fast answer, these models are designed to pause and double-check their facts, correcting mistakes before responding. This internal calculation process triggers thousands of hidden calculations for a single question, and because of these heavy resource demands, intelligent software can never scale down to be free.
Cloud vs. Edge Inference
To remedy exploding data center bills, the tech industry is splitting its infrastructure into two different strategies: the cloud and the edge. With cloud inference, heavy, highly complex AI tasks stay in the cloud, running on centralized servers managed by big tech companies. This gives businesses access to the most powerful models available, but it leaves them vulnerable to unpredictable usage spikes and constant network fees. With edge inference, tech companies are pushing simpler, everyday AI tasks to the “edge,” meaning the AI runs locally on specialized chips called Neural Processing Units (NPUs) built directly onto the device. This gives enterprises a choice: they must audit their workflow and decide which high-value tasks are worth the expensive bill and which ones can be offloaded onto the user’s own device.
The Jevons Paradox
Many corporate financial managers assume that because the market price of a single AI token has dropped dramatically, AI budgets will shrink. However, an economic rule known as the Jevons Paradox is occurring. The Jevons Paradox states that when technological progress makes a resource cheaper and more efficient, we don’t actually end up saving money. Instead, the lower price tag causes mass demand, and total consumption of that resource skyrockets.
Because basic AI text generation is now so cheap, engineers are no longer running just single queries. They are using cheap tokens to build automated systems where multiple AI agents “talk” to each other in recursive loops all day. By making individual tokens cheap to run, the industry has almost guaranteed that total corporate computing bills will skyrocket.
The Inference Class and Cost Agility
The biggest winners of the AI boom will be the Inference Class: practical, agile software companies that focus heavily on Financial Operations, tracking and optimizing cloud costs in real time. They buy raw AI power at cheap wholesale prices, package it smartly, and use it to solve specific, high-value problems for customers. Ultimately, economic resilience will be redefined by Cost Agility, the ability of a software platform to automatically switch out its underlying AI models without taking its system offline. In the digital economy, financial strength will belong to companies that both build strong infrastructure and manage daily AI costs with discipline and optimization.
Subscribe to the Future Economists newsletter by entering your email below to stay updated on our next piece and get insights on new technology and the economics behind it. 👇
Sources and References
- CloudZero: FinOps for AI
- CloudZero: FinOps Maturing, Efficiency Falling
- FinOps: CloudZero Member
- Zendesk: How AI Agent Pricing Works (Automated Resolutions)
- Salesforce Newsroom: Welcome to the Agentic Enterprise — Agentforce 360
- FinOps Foundation: A One Word Change — How the Community Made Our Mission Evolution Inevitable
- CloudZero: AI Pricing Explained — What AI Actually Costs and How Providers Charge for It in 2026
- Ren, Li, Yang & Islam: Making AI Less “Thirsty” — Uncovering and Addressing the Secret Water Footprint of AI Models (arXiv)
- IEA: Key Questions on Energy and AI — Data Centre Electricity Consumption, 2015–2025
- OpenAI: Trading Inference-Time Compute for Adversarial Robustness
- Canalys: Global Cloud Infrastructure Spending Reaches $95.3B in Q2 2025, Driven by AI
- W.S. Jevons: The Coal Question (Library of Economics and Liberty)