AI Monetization

17 min read

45 AI infrastructure and inference cost statistics

Written by

Pranathi Tipparam

Current benchmarks on AI spending, compute, power demand, inference pricing, and the revenue infrastructure behind rapidly changing AI economics

AI economics are moving in two directions at once. Gartner's latest AI spending forecast puts worldwide AI spending at $2.7 trillion in 2026, while Epoch AI estimates that the cost of achieving a given level of AI benchmark performance has fallen about 13x per year since 2023.

That combination creates an unusual monetization environment: infrastructure investment is expanding quickly while the cost of delivering specific levels of model performance can fall just as rapidly. AI companies therefore need pricing and billing infrastructure that can follow changing token economics, compute costs, model choices, credits, and customer usage. Orb's AI billing platform supports usage-based and hybrid models built on raw usage events, pricing controls, and finance workflows.

Key takeaways

  • AI spending continues to accelerate: Gartner forecasts $2.67 trillion in worldwide AI spending for 2026, up 49.5% year over year
  • Infrastructure dominates AI investment: Gartner forecasts approximately $1.48 trillion in AI infrastructure spending in 2026
  • Inference is overtaking training: Gartner expects inference to account for 55% of AI-optimized IaaS spending in 2026, with $23.3 billion allocated to inference workloads
  • Power requirements are expanding rapidly: The IEA projects global data-center electricity consumption to reach roughly 945 TWh by 2030
  • Inference economics deflate quickly: Epoch AI estimates that the cost of achieving a given AI performance level has fallen about 47% per quarter since 2023
  • Enterprise AI demand remains strong: Menlo Ventures estimates enterprise generative-AI spending reached $37 billion in 2025, 3.2x its 2024 level

AI infrastructure spending and market growth

AI spending now stretches across infrastructure, software, services, security, applications, and devices. Looking at those categories separately helps distinguish the economics of physical AI infrastructure from the much broader market forming around it.

1. Worldwide AI spending is forecast at $2.67 trillion in 2026

Gartner expects worldwide AI spending to reach $2.7 trillion in 2026 across infrastructure, software, services, cybersecurity, applications, models, and related categories.

AI is becoming a spending layer across the technology stack rather than a market confined to model developers, linking product economics to a broad ecosystem of compute and infrastructure providers.

2. AI spending is forecast to grow 49.5% in 2026

The same Gartner AI forecast projects 49.5% year-over-year growth from approximately $1.787 trillion in 2025.

Growth at that pace, from an already large base, points to sustained demand for chips, cloud capacity, power, networking, models, and AI software.

3. AI infrastructure spending reaches $1.48 trillion in 2026

Gartner forecasts AI infrastructure spending of approximately $1.484 trillion in 2026, representing more than half of its total AI spending forecast.

That balance underscores how capital-intensive AI remains beneath the software layer, with chips, servers, networking, devices, and data-center capacity shaping the economics that reach application companies.

4. AI-optimized IaaS spending reaches $42.3 billion

Worldwide AI-optimized IaaS spending is projected at approximately $42.276 billion in 2026.

The category is especially relevant to companies renting accelerated compute because cloud pricing can flow directly into model-serving and product-delivery costs.

5. AI-optimized IaaS spending grows 96.4%

Gartner expects 96.4% growth in AI-optimized IaaS during 2026, from roughly $21.529 billion in 2025.

Nearly doubling in one year shows how quickly demand for specialized cloud infrastructure is expanding as more AI workloads reach production.

6. Inference IaaS spending reaches $23.3 billion

Gartner forecasts inference spending of approximately $23.3 billion within AI-optimized IaaS in 2026.

Because inference is tied to live product activity, rising inference spend strengthens the connection between customer usage, infrastructure cost, and monetization.

7. Training IaaS spending reaches $19 billion

Training remains substantial, with Gartner projecting about training spending of $19 billion in 2026.

Training often creates concentrated bursts of compute demand, while inference accumulates continuously as customers use deployed products.

8. Inference reaches 55% of AI-optimized IaaS spending

Inference is expected to account for 55% of spending on AI-optimized IaaS in 2026, rising to 59% in 2027.

As AI moves deeper into production, economics increasingly depend on request frequency, token volume, model choice, context length, latency, and workflow complexity.

9. AI-optimized IaaS reaches $66.1 billion in 2027

Gartner projects IaaS spending for AI-optimized workloads to reach approximately $66.143 billion in 2027, another 56.5% increase from 2026.

Even with slower percentage growth than 2026, the market would still add nearly $24 billion in annual spending.

AI data-center capacity and capital requirements

AI models ultimately depend on physical infrastructure. Accelerators, high-bandwidth memory, networking, cooling systems, power generation, and data-center construction all shape the cost and availability of compute.

10. AI data centers require $5.2 trillion through 2030

McKinsey estimates that meeting worldwide AI demand will require $5.2 trillion of data-center investment through 2030.

That capital requirement shows the physical scale behind AI services that may appear to customers as little more than an API call or software interface.

11. AI capacity demand reaches 156 GW by 2030

McKinsey models approximately 156 GW of demand for AI-related data-center capacity by 2030.

At that scale, access to power generation, transmission, land, cooling resources, and permitting becomes an important part of AI infrastructure economics.

12. About 125 GW must be added between 2025 and 2030

Around 125 additional GW would need to be added between 2025 and 2030 under McKinsey's scenario.

Much of the capacity envisioned for the end of the decade therefore has yet to be built, making infrastructure delivery a constraint alongside semiconductor supply.

13. Global AI power capacity reached about 30 GW

Epoch AI estimates AI power capacity reached roughly 30 GW globally in the final quarter of 2025.

It is a modeled estimate rather than a facility census, but it provides a useful benchmark for comparing current scale with the much larger capacity requirements projected for 2030.

14. Global AI compute is growing 3.4x per year

Epoch AI's latest compute growth estimate puts annual growth in computing power from the global AI-chip stock at approximately 3.4x since 2022, with a 90% confidence interval of 3.2x to 3.7x.

At that pace, compute availability can change significantly within a single product or pricing cycle.

15. Global AI compute doubles every 6.8 months

Expressed another way, the current Epoch compute trend implies a doubling of global AI compute roughly every 6.8 months.

The pace shows why infrastructure assumptions can become outdated quickly as capacity expands and newer hardware reaches the market.

16. Frontier data-center compute has doubled about every seven months

At the upper end of the market, the computing capacity of the largest known AI data centers has recently grown at a pace equivalent to doubling every seven months.

This is a frontier measure rather than an industry average, but it shows how quickly the maximum feasible scale of AI infrastructure is moving.

17. Compute represents 54% to 62% of costs in a three-company sample

Epoch estimates that R&D and inference compute costs represented 54% to 62% of total costs and operating expenses across three AI companies it analyzed.

The sample is too small to generalize across the industry, but it illustrates how compute can become one of the dominant economic inputs at frontier AI businesses.

18. HBM represents 63% of AI-chip component costs

High-bandwidth memory accounted for an estimated 63% of component costs for AI chips in Q4 2025, up from 52% in Q1 2024.

The shift shows why AI-chip economics depend heavily on memory capacity and bandwidth, not only the accelerator processor itself.

Electricity demand from AI infrastructure

AI growth is also becoming an energy story. More efficient accelerators help, but deployment is increasing fast enough that overall data-center electricity demand is still projected to rise sharply.

19. Data centers could consume about 945 TWh in 2030

The IEA expects data-center electricity consumption to reach roughly 945 TWh globally by 2030 in its base case, more than double the estimated 415 TWh consumed in 2024.

Power availability can increasingly influence data-center location, expansion schedules, capacity planning, and compute costs.

20. Data centers approach 3% of global electricity consumption

By 2030, data centers could account for just under 3% of worldwide electricity demand under the IEA's base-case projection.

The global share masks greater local pressure in regions where large clusters of facilities compete for grid capacity.

21. Data-center electricity demand grows about 15% annually

The IEA projects electricity consumption growth from data centers of roughly 15% per year between 2024 and 2030.

That is more than four times the growth rate expected from all other sectors combined in its base case.

22. Accelerated-server electricity use grows 30% annually

Electricity consumption from accelerated servers is projected to rise about 30% per year, with the IEA identifying AI adoption as the primary driver.

Accelerated computing is therefore growing twice as quickly as overall data-center electricity consumption.

23. Accelerated servers contribute almost half of new demand

Accelerated servers are expected to generate almost half of the net increase in global data-center electricity consumption through 2030.

The mix makes AI-oriented acceleration one of the principal forces changing the future power profile of data centers.

24. Conventional-server electricity grows about 9% annually

Electricity consumed by conventional servers is still projected to grow around 9% annually.

That remains substantial but is markedly slower than the 30% growth projected for accelerated servers.

AI hardware efficiency and scaling

Physical expansion is only one side of the equation. AI hardware is simultaneously becoming faster, more energy-efficient, and more cost-effective, while frontier models consume increasingly large amounts of compute.

25. Machine-learning hardware performance improves 43% annually

Stanford's AI Index reports approximately 43% annual improvement in machine-learning hardware performance measured in 16-bit floating-point operations.

At that pace, performance doubles roughly every 1.9 years, allowing newer hardware to execute substantially more AI computation.

26. Hardware price-performance costs fall about 30% annually

The same AI Index finds that the cost of equivalent machine-learning hardware performance has improved by about 30% per year.

This creates downward pressure on inference costs even though energy, networking, software, and model-development expenses also shape end-user prices.

27. Hardware energy efficiency improves about 40% annually

Machine-learning hardware energy efficiency has improved by roughly 40% per year, according to Stanford.

Efficiency gains allow more computation per unit of electricity, although total power consumption can still rise when aggregate AI demand grows faster.

28. Frontier language-model training compute grows about 5x per year

Epoch AI estimates training compute for frontier language models has grown roughly fivefold per year since 2020.

The trend shows how developers continue applying more compute to push the frontier even as older capability levels become cheaper to reproduce.

29. Frontier training compute doubles every 5.2 months

The same trend implies a compute doubling time of approximately 5.2 months for frontier language-model training.

That widening scale can make the economics of developing frontier models very different from serving mature capabilities.

30. LLM training datasets double about every eight months

Stanford reports that LLM dataset sizes used in notable training runs have doubled approximately every eight months.

Larger datasets add storage, preprocessing, networking, and data-management costs alongside the accelerators performing the training.

31. AI-chip performance per dollar improves 49% annually

Epoch AI estimates that every dollar spent on AI chip performance has purchased roughly 49% more theoretical compute performance per year since 2023.

The measure is adjusted to constant 2025 dollars and reflects both better chips and purchasing shifts toward newer generations.

32. AI-chip price performance doubles every 1.7 years

At the same rate, Epoch hardware data implies performance per dollar doubles approximately every 1.7 years.

Infrastructure economics can therefore change materially before a long-term product or commercial strategy has finished its lifecycle.

33. GPU memory bandwidth rises 28% annually

GPU memory bandwidth has increased about 28% per year since 2008, according to Epoch's hardware tracking.

That matters because many AI workloads are constrained by moving model weights and intermediate data, not simply raw arithmetic performance.

34. Five hyperscalers own 71% of global AI compute

Amazon, Google, Meta, Microsoft, and Oracle collectively controlled an estimated 71% of AI compute as of Q4 2025, up from 63% in Q1 2024.

The estimate is based on H100-equivalent compute and highlights how much global AI capacity is concentrated among a relatively small number of infrastructure operators.

How quickly AI inference costs are falling

Rapid infrastructure expansion is happening alongside striking declines in the cost of reaching previously expensive model capabilities. That combination is reshaping what AI products can offer and how companies can price them.

35. GPT-3.5-level MMLU inference fell from $20 to $0.07 per million tokens

Stanford found that the cheapest model reaching GPT-3.5-equivalent performance of 64.8 on MMLU fell from $20 to $0.07 per million tokens between November 2022 and October 2024.

Holding capability approximately constant makes the comparison more useful than simply comparing prices across unrelated models.

36. The MMLU inference price fell more than 280x

That movement amounts to a 280-fold decline in the price of reaching GPT-3.5-equivalent MMLU performance.

Cost changes of this magnitude can make previously expensive capabilities viable for much higher-volume product use.

37. Inference prices have fallen 9x to 900x annually depending on the task

Stanford cites Epoch estimates showing annualized inference price declines ranging from 9x to 900x depending on the benchmark and performance threshold.

The wide range matters because AI companies cannot assume every workload will become cheaper at the same pace.

38. Equivalent AI performance has become 47% cheaper each quarter

Epoch AI estimates the cost of reaching a given benchmark-performance level has fallen approximately 47% each quarter since 2023 across five mathematics, science, and game benchmarks.

By holding performance approximately constant, the analysis captures cost deflation rather than simply comparing increasingly capable models.

39. Equivalent performance costs about 13x less each year

Compounded across a year, Epoch's estimate works out to roughly a 13x annual reduction in the price of equivalent benchmark performance.

Epoch cautions that the dataset spans only about three years, but the pace remains far faster than typical SaaS pricing cycles.

40. New state-of-the-art performance falls 66% in price each quarter

For capability levels immediately after they first appear at the frontier, Epoch estimates 66% quarterly declines in price across its benchmark average.

The pattern suggests newly scarce capabilities can become much cheaper once competing models and optimized serving approaches reproduce them.

41. Two-year-old performance levels fall 32% in price each quarter

Even two years after a performance level first appears, Epoch finds prices declining about 32% per quarter, or roughly 4.7x per year.

Cost deflation therefore continues well after a capability leaves the frontier.

Enterprise AI demand and monetization

Cheaper inference does not automatically mean lower total spending. Lower unit costs can unlock additional workloads, while wider deployment and heavier usage push aggregate consumption upward.

42. Enterprise generative-AI spending reached $37 billion

Menlo Ventures estimates enterprise AI spending reached $37 billion in 2025.

Its market model combines a survey of 495 U.S. enterprise AI decision-makers with bottom-up estimates covering applications, models, training infrastructure, and supporting infrastructure.

43. Enterprise AI spending grew 3.2x year over year

Menlo estimates spending increased 3.2x year over year, from $11.5 billion in 2024 to $37 billion in 2025.

The increase shows how aggregate AI spending can rise rapidly even while the unit cost of individual inference workloads falls.

44. Foundation-model APIs captured $12.5 billion in enterprise spend

The model API layer captured an estimated $12.5 billion in enterprise generative-AI spending in 2025.

APIs naturally produce measurable consumption units such as tokens and calls, placing usage at the intersection of infrastructure expense and product monetization.

45. PLG accounts for 27% of enterprise AI application spend

Product-led growth accounts for an estimated 27% of AI spend at the application layer, compared with 7% in traditional software.

A larger self-service contribution makes usage, credits, limits, and spend visibility more important earlier in the customer journey.

How Orb supports changing AI economics

AI economics can move faster than traditional software pricing cycles. Compute expands, model costs fall, and customer consumption can rise quickly as products move into production.

Orb gives AI companies a common revenue foundation for adapting to those changes.

Raw usage and pricing

Orb's standard architecture retains raw usage events, which can represent tokens, model calls, agent runs, compute time, API requests, completed actions, and other measurable outcomes.

Those events can feed billable metrics, backfills, corrections, simulations, and billing investigations. Orb’s dimensional price groups also support pricing across multiple usage dimensions—such as region, instance type, and environment—using a single pricing configuration for dimension combinations.

Pricing evolution

Orb's pricing simulations apply proposed pricing to historical usage so teams can compare customer and revenue effects before rollout.

Price evolution then supports controlled deployment of approved changes, while Orb's AI billing capabilities can combine usage charges, seats, platform fees, prepaid credits, minimum commitments, overages, and customer-specific terms.

Usage visibility and scale

Orb's Experience Kit can power customer-facing usage experiences, while Spend Controls supports monitoring, threshold alerts, and automated workflows.

For high-volume workloads, Orb's Enterprise platform is regularly stress-tested at 250,000+ events per second with idempotency guarantees, with Hosted Rollups available for configured higher-volume event streams.

Billing context remains auditable

Frequent price evolution still requires a reliable historical record.

Orb connects raw usage events, pricing configurations, invoices, corrections, and finance workflows so companies can change monetization while retaining context around the usage and commercial terms behind customer charges.

Frequently asked questions

Why are AI inference costs falling so quickly?

Several forces are improving at the same time. Hardware delivers more performance per dollar and per unit of energy, model architectures become more efficient, smaller models reach capability levels previously associated with larger systems, and inference software continues to improve. Epoch AI estimates that the cost of achieving a given benchmark-performance level has fallen about 47% per quarter since 2023, although the pace varies significantly across tasks and capability levels.

Does cheaper inference mean AI spending will fall?

Not necessarily. Lower unit costs can make new AI workloads economical, while existing products may generate much higher usage as adoption expands. Gartner still forecasts rapid growth in AI infrastructure spending, and Menlo estimates enterprise generative-AI spending grew 3.2x from 2024 to 2025. Total expenditure therefore depends on both the price of each unit and how many units organizations consume.

Why is inference becoming more important than training?

Training happens when models are created, fine-tuned, or updated, while inference occurs whenever a deployed model processes a live workload. Once AI products reach production, every customer interaction can contribute to ongoing infrastructure consumption. Gartner expects inference to account for 55% of AI-optimized IaaS spending in 2026, putting it ahead of training within that infrastructure category.

How can AI companies adapt pricing as costs change?

AI pricing works best when product usage can be measured separately from the commercial rules applied to it. Historical consumption then gives teams a way to test how alternative rates, metrics, credits, or packages would affect real customers. Orb supports that process through raw usage events, pricing simulations, price evolution, dimensional pricing, credits, and hybrid pricing structures.

What billing capabilities matter for AI products?

AI products can benefit from granular event ingestion, flexible billable metrics, raw usage events, hybrid and credit-based pricing, dimensional pricing, backfills, simulations, customer-facing usage visibility, and finance workflows. Spend controls and scalable ingestion become increasingly relevant as consumption grows. Keeping those capabilities connected around the same usage foundation also makes it easier to change pricing without losing the history behind customer charges.

Contact Sales

Ready to try a billing platform built for modern growth?

See how AI companies are removing the friction from invoicing, billing and revenue.