AI and Money23 articles
AI and Money / Article 21

Token demand could grow 24-fold Why lower unit prices may not shrink the bill

Cheaper units do not guarantee lower totals. Forecasts and reported consumption explain the pattern

Published: Updated:

Key takeaway

A Goldman Sachs report was described as forecasting that agentic AI could increase token demand by more than 24 times over the coming years. Reporting also said AT&T's consumption grew about 27-fold in 18 months. Why can bills keep rising while unit prices fall? This article draws together the cost patterns in the cases beginning with article 12.

THE RESULT FIRST

Total cost is unit price × volume. Volume growth can outrun price cuts.

More than 24×Reported Goldman Sachs forecast for token-demand growth associated with agentic AI

Even if unit prices halve, a 27-fold rise in volume makes the total 13.5 times larger. Waiting for cheaper units does not by itself protect a budget.

AT&T: about 27-fold growth in 18 months, according to reporting
A US insurer: about 50-fold growth in under a year, according to reporting
Greater efficiency can encourage greater consumption

Reported figures: consumption growth is already visible

OrganizationReported figureGrowth
Goldman Sachs (forecast)Agentic AI could increase token demand by more than 24 timesAbout 24× over the coming years
AT&T (US telecom)From about 1 billion to 27 billion tokens per dayAbout 27× in 18 months
Major US health insurer (not named)A large increase from several million tokens per monthSeveral tens of times in less than a year
Meta (research-firm analysis)Over 60 trillion tokens company-wide over one 30-day period (article 13)─

Reported figure

AT&T (US telecom)
From about 1 billion to 27 billion tokens per day
Major US health insurer (not named)
A large increase from several million tokens per month
Meta (research-firm analysis)
Over 60 trillion tokens company-wide over one 30-day period (article 13)

Growth

Goldman Sachs (forecast)
About 24× over the coming years
AT&T (US telecom)
About 27× in 18 months
Major US health insurer (not named)
Several tens of times in less than a year
Meta (research-firm analysis)
─

Much of this is described as normal expansion of useful work, rather than waste. A simple chat question may use thousands of tokens; an agent researching, trying alternatives, and checking the result can use millions. More capable workflows can consume orders of magnitude more tokens per task.

But aren't unit prices falling? The Jevons paradox

A common management question is: if AI unit prices keep falling, should total spending fall too?

Prices per million tokens often fall as models improve. But the total remains unit price multiplied by volume (article 3). If the price halves while volume rises 27-fold, spending becomes 13.5 times larger.

Economics describes a related effect as the Jevons paradox: improved resource efficiency can encourage enough extra use to increase total consumption. It is known from the history of coal use during industrialization. In AI, cheaper work can invite more tasks, and more tasks can increase the total bill. A falling unit price alone does not guarantee lower spending.

Another structural issue: the provider controls key variables

A Forbes commentary argues that users have limited control over provider-set unit prices, plan structures, and model retirement or replacement. Contracts and negotiation can help, but a dashboard does not control those variables. Reconsider the cases from that perspective:

  • Unit prices change (article 9).
  • Entire plan structures change (article 19).
  • Consumption can grow rapidly, as in this article's demand scenarios.
  • Managing it creates additional work (article 20).

Users can limit consumption, optimize workloads, negotiate, or switch providers, but reducing useful adoption simply to contain a bill can undermine the purpose of AI. The shared goal is to expand useful work while retaining more control over the cost structure.

Reviewing articles 12–21

  1. 12

    Case or theme: Amazon's leaderboard closure

    Key point: Competing over consumption can grow the bill

  2. 13

    Case or theme: Meta's Claudeonomics

    Key point: Enthusiasm and anxiety can increase metered usage

  3. 14

    Case or theme: Microsoft's tool transition

    Key point: Popular tools can still be reviewed for cost and strategy

  4. 15

    Case or theme: Uber's exhausted budget

    Key point: Successful adoption can coincide with a budget overrun

  5. 16

    Case or theme: One employee in Japan, JPY 10 million/month

    Key point: The same cost pattern was reported in Japan

  6. 17

    Case or theme: Unattended agents and leaked keys

    Key point: Individuals can also face an overnight bill

  7. 18

    Case or theme: AI spending versus labor costs

    Key point: AI bills can reach staffing-scale amounts

  8. 19

    Case or theme: Changes to flat-rate plans

    Key point: The assumption of predictable subscription terms can change

  9. 20

    Case or theme: Token management

    Key point: Cost control creates its own labor requirements

  10. 21

    Case or theme: The 24-fold demand forecast

    Key point: Volume growth can outweigh falling unit prices

If demand grows dramatically, a structure that does not bill each extra local token becomes more valuable for suitable workloads. With upfront-purchase AI such as Sovereign GaiXer, more useful work spreads the hardware cost across more tokens. Electricity and operational costs remain, and growing beyond the machine's capacity requires more hardware. Within that capacity, the effect can become better utilization of a fixed investment rather than a larger token bill.

Think of it this way

Imagine fuel prices halving while your driving distance increases thirtyfold because more jobs become economical to reach. The cost per kilometer falls, but the total fuel bill rises. Cheaper operation can encourage more use; it does not automatically mean a smaller total.

Frequently asked questions

How reliable is the 24-fold forecast?

It is a forecast and should be treated as a scenario with uncertainty. Reported measurements such as AT&T's 27-fold growth in 18 months show that order-of-magnitude changes are possible, without proving the forecast will occur everywhere.

Does the Jevons paradox always happen?

No. If demand saturates, total use can level off. Many forecasts expect further growth in AI because there are still many tasks people would delegate if the cost fell.

Is waiting for lower prices a good adoption strategy?

Prices may fall, but delaying also postpones experience with useful AI. And lower unit prices do not guarantee a lower future total if usage expands. Consider both learning value and realistic demand.

Are techniques such as prompt compression pointless?

No. Doing the same work with fewer tokens helps both cloud and local inference. Efficiency improvements may still be outweighed by much larger demand growth, so model both effects rather than assuming one cancels the other.

Can customers gain negotiating power?

Volume agreements and credible access to multiple providers can help. Ownership of local inference offers another form of control, although hardware suppliers, licenses, and operating costs still create dependencies. No option eliminates every external constraint.

Can one on-premises system handle 24 times the demand?

Not necessarily. A machine has finite capacity, so growth may require additional systems. The distinction is planned equipment investment instead of an unexpectedly multiplied usage bill. Size the system for actual workload and concurrency.

Summary

  • Reporting described a Goldman Sachs forecast of over 24-fold demand growth, alongside measured examples such as AT&T's reported 27-fold growth in 18 months.
  • Total cost is unit price × volume. Rising demand can outweigh price reductions, an effect related to the Jevons paradox.
  • Providers control important service variables such as pricing and model availability, while users have varying contractual and switching options.
  • The recurring lesson in articles 12–21 is that linking every unit of consumption to a charge creates budget exposure.
  • Purchased local inference can turn more useful usage into better asset utilization, within capacity and with operating costs included.

This article is based on public reporting, without independent interviews. Forecasts are not guarantees; reported measurements reflect their reporting dates. Please contact us with factual corrections. The media, researchers, and companies discussed do not endorse Sovereign GaiXer.
Company, product, and service names mentioned are trademarks or registered trademarks of their respective owners.