Choosing your usage-based value metric: the layer cake pricing model


Market data revealing how AI companies are transforming their pricing strategies to capture value from tokens, API calls, and compute usage
The AI pricing landscape is undergoing a fundamental shift. With global AI spending forecast to total nearly $1.5 trillion in 2025 and average monthly corporate AI budgets projected to rise to $85,521, businesses face an urgent challenge: for AI products whose costs and delivered value scale with tokens, API calls, compute, or other consumption metrics, seat-only pricing can create real misalignment between what a customer pays and what a customer consumes. Stripe notes that per-token and per-API-call consumption metrics are common among AI model companies. Companies building AI monetization strategies need billing infrastructure that handles variable usage patterns, from token-based LLM charges to compute-intensive inference workloads. The statistics below show why usage-based and hybrid pricing models are prevalent across current surveys and pricing-page analyses of AI and SaaS vendors, and what this means for revenue operations.
The economics of delivering AI are what force the pricing question. When the cost of serving a customer moves with consumption, flat pricing quietly transfers that volatility onto the vendor's margin. The statistics in this section show how quickly that pressure has built and how vendors are responding.
Flexprice's June 2026 analysis of 50 AI pricing pages found that 46 of the 50 tools included a usage component in their pricing, and only four charged a flat per-seat fee. This is a pricing-page analysis of a 50-tool sample rather than a market-wide census, but within that sample, pure seat-based pricing is now the exception rather than the rule.
In Revenera's Monetization Monitor, a survey of 501 product leaders at global technology companies, 70% of those developing AI products or features said delivery costs are eroding their margins. This is the mechanism behind the pricing shift: when unit costs scale with consumption, pricing that does not scale with consumption compresses margin as adoption grows.
52% of surveyed AI producers plan to introduce new pricing or monetization models specifically to mitigate cloud spend. Executing that change without a multi-quarter engineering project is exactly what usage-based billing infrastructure exists to make possible.
CloudZero's March 2025 survey of 500 U.S. software engineers at manager level and above found that average monthly AI spend was projected to rise 36%, from $62,964 in 2024 to $85,521 in 2025. This is a projection rather than an observed year-end result, but the trajectory suggests companies moving beyond pilot programs into production deployments where consumption compounds.
The share of organizations planning to invest over $100,000 per month in AI rose from 20% to 45% in CloudZero's research. The important qualifier is that these are organizations planning that level of spend, not necessarily organizations already reaching it. Either way, the shift toward higher spending tiers creates demand for enterprise-grade billing with robust reporting.
Subscription remains the single most common AI monetization model at 42% of respondents, but Revenera projects the share of pure subscription pricing to decline as usage-based approaches grow. That is the transition state most AI vendors are living in right now: a subscription base that no longer covers a consumption-driven cost structure.
57% of Revenera respondents identified rising cloud spend as the biggest blocker to ARR growth. When infrastructure cost becomes the constraint on revenue growth, aligning price with consumption stops being a pricing preference and becomes a margin requirement.
The shift from seat-based to consumption-based pricing represents one of the most significant changes in software business models. AI companies are leading this transformation because their variable cost structures require pricing that aligns with actual usage.
Metronome's January 2025 survey of 100 SaaS companies found that 85% had adopted usage-based pricing in some form. This is a survey of 100 companies rather than a census of the software industry, but it is a direct measurement rather than a figure passed through secondary reporting.
In a Salesforce Ventures survey of more than 300 startup and enterprise executives, more than half of AI sellers use hybrid pricing, while 23% use pure consumption pricing. Hybrid gives customers a predictable floor while letting vendors capture value from variable consumption above it.
Stripe reports that among AI company leaders, 56% use hybrid pricing and 38% use pure usage-based pricing. Only a small minority still rely on traditional seat-based models alone.
The AI pricing landscape remains highly dynamic: 87% of AI sellers surveyed by Salesforce Ventures plan to change their pricing within the next 12 to 18 months. That cadence is the strongest practical argument for billing systems that support price evolution with simulation and auditable, controlled rollouts.
Flexprice's analysis of 50 AI and SaaS pricing pages found 38 of the 50 used tiered pricing. Tiering is rarely the whole model on its own; in this sample it usually sits alongside usage components, which means billing systems have to evaluate several pricing mechanics against the same stream of events.
Despite rapid AI adoption, many organizations struggle with the financial mechanics of AI pricing. These challenges create opportunities for billing solutions that bring transparency and predictability to AI consumption.
McKinsey's Enterprise AI FinOps survey of May 2026, covering 120 enterprise participants and 75 qualified respondents across five major industries, found that 93% of respondents report exceeding their AI budgets. Budget overruns at that rate are not a procurement failure so much as a measurement failure: consumption is being incurred faster than it is being observed.
Financial leaders face significant challenges with AI pricing. In a survey of 614 CFOs reported by Salesforce Ventures, 71% said their company struggles to monetize AI effectively, indicating a gap between AI capabilities and revenue capture.
The disconnect extends beyond monetization challenges. 68% of CFOs said their current pricing model no longer works for AI products and services, signaling the need for pricing infrastructure overhauls.
78% of IT leaders surveyed reported unexpected SaaS charges due to consumption-based or AI pricing models. Note the precise framing: these are unexpected SaaS charges attributed to consumption-based or AI pricing, which is a billing-visibility problem rather than a verdict on the pricing model itself. It is also the problem spend controls and real-time usage visibility are built to solve.
McKinsey reports that AI token usage can vary by as much as 30 times for the same task. That single fact explains most of the forecasting difficulty in AI: the unit of work is not the unit of cost, so accurate budgeting depends on accurate metering of what was actually consumed.
Charging for consumption is only possible if consumption can be measured, attributed, and forecast. These statistics show how far most organizations still are from that baseline, and what it is worth to close the gap.
Just 41% of Revenera respondents say they can gather product usage data very well. Usage data quality is the hard prerequisite for usage monetization: a company that cannot measure consumption reliably cannot bill for it defensibly.
McKinsey finds that only 20% to 25% of surveyed organizations have mature AI FinOps practices. The gap between AI adoption and AI cost discipline is where most unexpected charges and stalled renewals originate.
74% of software suppliers surveyed by Revenera have adopted usage-based pricing at least moderately. Adoption at that level makes usage metering a core revenue system rather than a reporting nicety.
Organizations with high AI cost forecasting maturity save about 10% more on AI spend on average than their peers, according to McKinsey. Granular consumption data is what makes that forecasting possible on both sides of the invoice.
AI unit economics move faster than almost any other cost base in software, and much of the spend is still invisible to the teams accountable for it. Both facts push toward metering that is granular enough to reprice against.
The cost to perform inference at GPT-3.5 performance levels dropped over 280-fold between November 2022 and October 2024. Read this as evidence that AI unit economics change extraordinarily quickly, not as evidence in itself that usage-based pricing adoption increased. The pricing implication is that any model priced against today's cost curve will be wrong within a few quarters unless it can be revised quickly.
Beyond model-level improvements, Stanford's AI Index reports that AI hardware costs declined about 30% annually. This consistent improvement allows AI companies to pass savings to customers or expand margins.
43% of Revenera respondents identify limited visibility into entitlement and usage data as a renewal management challenge. Poor usage visibility does not just create billing disputes; it undermines the renewal conversation, because neither side can evidence the value delivered.
McKinsey reports that 20% to 30% of AI spend is often unaccounted for. Unattributed spend at that scale is the strongest possible argument for granular, dimensioned consumption data on every event.
The statistics in this section come from Flexprice's June 2026 analysis of 50 AI and SaaS pricing pages, so they describe that sample rather than the whole market. What they show is which billing mechanics AI products actually implement, and each one is a distinct metering, rating, and invoicing requirement.
30 of the 50 analyzed pricing pages used prepaid usage plans. Prepaid models require credit balances, drawdown logic, and expiry rules that traditional subscription billing systems were never designed to track.
23 of the 50 pricing pages used pay-per-use pricing. Pure pay-per-use is the most demanding case for event ingestion accuracy, because every billable event maps directly to a line on an invoice.
22 of the 50 pricing pages included overage fees. Overages sit at the boundary between subscription and consumption, and they are where customers most often encounter charges they did not anticipate.
17 of the 50 pricing pages used volume-based pricing. Volume and tiered structures require the billing engine to evaluate cumulative usage across a period before it can rate a single event correctly.
Every one of the 8 AI infrastructure tools in Flexprice's sample billed on consumption. The closer a product sits to raw compute, the more completely consumption pricing takes over.
7 of the 9 voice AI products analyzed used per-minute pricing. Voice is a useful reminder that the billable unit is category-specific: minutes here, tokens elsewhere, and something different again for image or agent workloads.
All 8 video AI tools in the sample used prepaid pricing. Prepaid dominance in compute-heavy categories aligns cash collection with the cost of delivery.
Knowing that usage-based pricing is spreading is only half the picture. The other half is how hard it currently is to change a pricing model, and where the operational friction actually sits.
36% of Revenera respondents cite uncertainty around how to price AI features as a blocker to price and value alignment. Uncertainty of that kind is best resolved empirically, by testing candidate models against real usage data before launch.
57% of respondents report that introducing a new monetization model commonly takes three to nine months. Against an AI cost curve that moves every quarter, a three to nine month change cycle is itself a competitive disadvantage.
54% of respondents identify complex or manual quote-to-cash workflows as their top quote-to-cash challenge. Manual steps that were tolerable at monthly seat counts break down entirely at the event volumes AI products generate.
McKinsey reports that thoughtful AI consumption management can reduce AI costs by 20% to 30%. Vendors that give customers this visibility natively turn a cost anxiety into a retention argument.
Metronome reports that 77% of the largest software companies have some form of usage-based pricing. Adoption at the top of the market sets the buyer expectations that everyone else is measured against.
Metronome also reports that 64% of the Forbes Next Billion-Dollar Startups offer usage-based pricing. Adoption among both incumbents and the fastest-growing challengers is what makes this a structural shift rather than a segment preference.
The final statistics answer two questions directly: how recent this shift is, and how much of the value gap it has actually closed so far.
Metronome found that 78% of companies with usage-based pricing adopted it within the preceding five years. This is a recent migration, not a long-settled norm, which is why so many billing stacks were originally built for seats.
Nearly half of usage-based pricing adopters adopted the model within the preceding two years. The acceleration tracks the arrival of AI products whose costs simply do not fit a per-seat container.
Just 36% of Revenera respondents report strong alignment between price and value. Adoption of usage-based pricing is running ahead of the metering, packaging, and visibility work required to make it feel fair to customers.
56% of respondents expect revenue from usage-based pricing to grow by 2027. That expectation, held by the people who set pricing, is the clearest available signal of where software monetization is heading.
The statistics above reveal a clear pattern: AI companies need billing infrastructure that can handle variable consumption, support frequent pricing changes, and provide the transparency that prevents unexpected charges. Several capabilities prove essential:
AI applications generate high volumes of billing events from tokens to API calls to compute minutes. Billing systems must ingest and process these events accurately without creating bottlenecks. Orb's metering infrastructure is built for exactly this shape of workload, with raw usage event storage, high-volume ingestion, granular metering, and aggregation surfaced fast enough for real-time invoicing.
With 87% of AI sellers planning a pricing change in the next 12 to 18 months, and 57% of Revenera respondents reporting that a new monetization model commonly takes three to nine months to introduce, the ability to iterate quickly is a direct competitive advantage. Orb's price modeling covers seat-based, hybrid, prepaid credit, tiered, bulk, and dimensional structures, while price evolution adds simulation on real usage data and auditable, customer-level rollout plans, so teams can change pricing without rebuilding billing infrastructure for every pricing change.
With 78% of IT leaders reporting unexpected SaaS charges due to consumption-based or AI pricing models, and 93% of organizations in McKinsey's survey exceeding their AI budgets, transparency is the difference between a renewal and a churn event. Orb's Experience Kit powers advanced real-time dashboards that help customers plan, monitor, and optimize usage across dimensions like teams, locations, and metrics, and spend controls add the thresholds, alerts, and workflow triggers that stop costs before they become surprises.
Before deploying pricing changes, AI companies benefit from testing against historical data. Orb's pricing simulations let teams test pricing with real usage data and forecast the impact across customers and revenue before launch, which is the practical answer to the 36% of vendors who say uncertainty about pricing AI features is blocking price and value alignment.
AI billing involves high volumes and small unit prices where errors compound quickly. Orb's accuracy foundation continually stores raw usage events as a single source of truth and supports backfills, backdated price changes, and automatic invoice recalculation when corrections are needed, so a late-arriving event or a retroactive contract change does not turn into a manual reconciliation project.
LLM pricing typically uses token-based billing, where providers charge separately for input tokens (the prompt sent to the model) and output tokens (the model's response). Prices vary significantly based on model capability, context window size, and volume commitments. Many providers also offer hybrid models that combine base subscriptions with per-token usage charges for overages.
AI costs stem from several factors: compute infrastructure for model training and inference, data storage and processing, model complexity and capability tier, API rate limits and context window sizes, and integration or customization requirements. For enterprises, hidden costs often include unexpected usage overages, integration engineering time, and ongoing model fine-tuning expenses.
It helps to separate two things. Usage-based pricing aligns charges with measured consumption; transparency and predictability come from the controls built around it. Usage pricing on its own can still produce bill shock, which is why 78% of IT leaders report unexpected SaaS charges due to consumption-based or AI pricing models. What actually delivers optimization is the surrounding layer: granular metering, customer-facing dashboards, threshold alerts, budgets, and spend controls. When customers see exactly what they consume, from tokens to API calls to compute hours, they can identify inefficiencies and adjust usage patterns.
Effective AI spend management requires real-time usage dashboards, configurable spending alerts, budget caps by project or team, detailed usage breakdowns by dimension (model, region, application), and historical usage trends for forecasting. Enterprise billing platforms typically provide these capabilities through customer portals and API integrations.
API rate limits cap the throughput available to an application, so organizations that exceed their allocated rates may face throttling that impacts application performance. How each provider handles demand above those limits is provider-specific and changes frequently, which is why rate tiers, quotas, and overage terms differ widely across vendors rather than following a single common model. For teams selling AI services, the practical implication is that tiers, quotas, and overage behavior all have to be metered and rated consistently against the same stream of raw usage events, which is what Orb's metering infrastructure and spend controls are built to do.
Token-based pricing charges based on the text processed, measuring input and output tokens regardless of processing time. Compute-based pricing charges for actual infrastructure resources consumed, typically measured in GPU-seconds or similar units. Token pricing works well for text-focused LLM applications, while compute pricing better suits training workloads, image generation, or custom model inference where processing time varies significantly.



See how AI companies are removing the friction from invoicing, billing and revenue.