Blogs -AI Token Demand is Shattering Forecasts 

AI Token Demand is Shattering Forecasts 


July 30, 2026

author

Beth Kindig

Lead Tech Analyst

  • Token processing, a key indicator of Inference demand, has surpassed early estimates by leaps and bounds even after large upward revisions.
  • Statements made by one of the market’s leaders in AI server sales illustrates this point clearly.
  • Even as one biggest players in AI inference have seen token processing soar by more than 300X in two years, a top Wall Street bank is calling massive growth ahead.

While Wall Street is busy debating whether AI can monetize, inference demand is exploding, with token processing offering the clearest evidence. Total annual token processing is no longer measured in billions or trillions of tokens, but in the quadrillions and beyond. 

As annual token processing is now tracked in units with 15 trailing zeros, it is becoming more evident that management teams in the AI supply chain and researchers have drastically underestimated the pace of growth.  

Forecast revisions that once seemed dramatic have since been eclipsed years ahead of schedule, with one of the strongest pieces of evidence coming from a top Nvidia partner.  

Below, we outline how token processing growth has been greatly underestimated and why expectations may need to move higher once again. 

Dell's Token Processing Forecast Miss Underscores the Scale of AI Inference Demand 

Dell is one of Nvidia’s key partners in deploying its GPUs through the company’s PowerEdge servers. Dell’s ability to forecast future demand is critical for supply chain readiness, considering its $16.1 billion in server revenue in Q1 was nearly 3X higher than both HPE and Lenovo and 1.6X higher than Super Micro.  

In this context, when we consider the source, a quote from Dell COO Jeffery Clarke is striking. In October 2025, Clarke said: “We thought as we model this, that inference would drive by 2028, 1 quadrillion, that's 15 zeros, 1 quadrillion tokens. Now it's 57 quadrillion, and I'm sure we're wrong.”  

mid

In other words, Dell has upped its past estimate for token processing for 2028 by 57X, and still thinks that number is conservative. The company also noted that its expectations for inference demand increased by a minimum of 100X in less than a year. 

Based on current levels of token processing, calling Dell’s 57 quadrillion token forecast for 2028 an underestimation is putting it kindly. Tokens processed per day is currently tracking near 370 trillion, or roughly 135 quadrillion per year. This means current token consumption is already 2.4X higher than Dell’s 2028 estimate - with years to spare.   

Even at the floor estimate of 300 trillion tokens per day, or 109 quadrillion per year, token consumption today would be 1.9X higher than Dell’s 2028 forecast.  

Updated forecasts by industry analysts shed further light on just how far off Dell’s estimate could be. 

Analysts and AI Leaders Highlight Token Processing Underestimation 

Goldman Sachs’ forecasts from May 2026 estimate that token processing will hit 47 quadrillion per month in 2028. That is approximately 565Q tokens per year, or about 10X Dell’s forecast. The firm sees monthly token processing rising from 1.7Q in mid-2025 to nearly 120Q in mid-2030 (1,440 quadrillion annually) a more than 70X increase in five years. Of this, Goldman estimates that around 101Q tokens will come from agentic workloads, or over 80% of the total. 

At a current monthly run rate of 11Q tokens compared to Goldman’s May 2026 estimate of 5.6Q monthly tokens, the bank’s forecast may be conservative. 

Line chart forecasting global AI token processing to rise from 1.7Q monthly tokens in 2025 to 120Q by 2030.

This chart forecasts rapid growth in global AI token processing, increasing from 1.7 quadrillion monthly tokens in mid-2025 to 47 quadrillion in 2028 and 120 quadrillion by mid-2030. The forecast implies more than 70X growth over five years. The chart also notes a current run rate of 11 quadrillion monthly tokens, roughly double Goldman Sachs' May 2026 estimate of 5.6 quadrillion, and projects that agentic AI will account for approximately 84% of AI workloads by 2030.

Meanwhile, tech consulting firm Tirias Research says when it first forecasted global AI demand in 2023, it expected annual token output would hit 20T by year-end 2024. It estimates that actual token usage hit 667T, more than 33X higher than its original forecast. The company updated its outlook in mid-2025 to 76.9Q tokens annually by 2030, and current token consumption is already around 1.75X higher than this figure. 

Another notable data point comes from Anthropic. CEO Dario Amodei said the company planned to achieve 10X growth in 2026. However, in Q1, revenue and usage climbed 80X on an annualized basis, leaving the firm struggling to keep up with its compute needs. 

AI Token Processing Is Surging Across Big Tech and Inference Platforms 

While many companies developing AI models have not previously provided token processing expectations to compare against, the raw explosion in token processing would have been hard for anyone to predict.  

Google Reports 330X Token Growth in Two Years  

If you think that 70X or 33X growth through 2030 is quite a lot, Alphabet’s growth in monthly tokens processed dwarfs that – coming in at 330X over the last two years.  

Google said that it was processing 3.2 quadrillion tokens per month in May, or 38.4Q a year, which alone would account for 67% of Dell’s 2028 forecast. The growth in the company’s token processing is massive, increasing by 7X since May 2025, and 330X since May 2024.  

While these figures measure token processing across all of Google’s surfaces, the company has also seen a huge increase in usage for its Gemini models. In its Q2 2026 earnings call, Google said Gemini models were processing 22 billion tokens per minute, up 120% in just six months, or equal to 1 quadrillion tokens per month. 

Google's monthly token processing grew from 9.7T in May 2024 to over 3.2Q in May 2026, with 7X YoY growth.

Google's monthly AI token processing increased from 9.7 trillion in May 2024 to over 3.2 quadrillion in May 2026, representing 7X year-over-year growth and accelerating AI inference demand.

Microsoft, OpenRouter, and Fireworks AI Showcase Rapid Token Processing Growth  

Microsoft CEO Satya Nadella noted in the company’s April earnings call that it processed over 100 trillion tokens during the quarter, a 5X increase YoY, and processed a record 50 trillion in March. That is a far cry from Google at just 0.1Q tokens per quarter, but shows rapid processing growth nonetheless.  

In May, LLM interface provider OpenRouter notes that its weekly token volume increased by 5X in six months from 5 trillion to 25 trillion, and that it was on pace to process more than 1Q tokens in 2026. Additionally, Fireworks AI, which provides an inference serving platform, processes 40 trillion tokens per day as of mid-July, more than doubling in three months from 15T per day in April and up 4X from October 2025 when it hit 10T per day.  

Why AI Token Processing Is Growing Faster Than Expected  

Looking at the numbers, the extent to which token processing is far exceeding previous expectations is somewhat staggering. One of the most prevalent reasons for the dramatic increase in token processing compared to what was originally modeled is the rise of agentic AI. As Goldman Sachs puts it plainly; “We weren’t talking about agents a year ago, now we are”. 

Anthropic estimates that multi-agent systems use up to 15X more tokens than chatbot requests. Meanwhile, third-party researchers like those at Stanford say that coding agents consume 1,000X more tokens than code reasoning and code chats. Along with this increase in token consumption needs, agent usage is on the rise. Microsoft said in April that its first-party agent usage had increased by 6X year-to-date, or 6X in just four months. 

Reasoning Models Take Over Token Consumption 

A key enabler of agentic AI is the rise of reasoning models, or models that employ a thought process for how they should respond and refine outputs as they go, introducing significant complexity compared to non-reasoning models. The uptick in reasoning model usage helps explain the vast increase in token processing.  

A study that analyzed 100 trillion tokens on OpenRouter found that tokens served by reasoning models increased from 0% at the start of 2025 to around 60% near the end of 2025. The beginning of this trend aligns closely to when OpenAI released the full version of its o1 reasoning model in December 2024.  

Researchers also found that prompt tokens per request increased by 4X compared to early 2024, and completion tokens per request tripled. This indicates that users are asking models to complete more complex tasks, increasing the amount of input and output tokens for each request. 

OpenRouter chart showing reasoning models growing from 0% to over 60% of token traffic during 2025.

The chart tracks reasoning versus non-reasoning token usage on OpenRouter from late 2024 through late 2025. The share of tokens served by reasoning models steadily increases from near zero to more than 60%, crossing 50% by the end of 2025 and highlighting the rapid adoption of reasoning AI models.

Furthermore, Google has shown evidence that queries are rising far faster than raw user count. The company said that from Q2 2025 to Q3 2025, monthly active users on the Gemini app increased by 44% from 450 million to 650 million. However, during the same period, queries increased by 3X, indicating that it not only added many users, but that each user also increased their engagement. 

Conclusion 

It is clear that token processing has risen faster than early estimates by multiple orders of magnitude. Even as this has taken place, analysts forecast that explosive token growth will continue for years to come, driven significantly by agentic AI inference.  

The continued rise in inference demand may be the strongest driving force behind the AI infrastructure trade going forward. Data center operators will not only require more compute, but also more powerful and efficient computing systems, such as Nvidia’s latest Rubin generation and its inference specific variants.  

The I/O Fund recently released our new 90-page Top 20 AI Stocks for Q3 2026 report, where we identify the lesser-known companies best positioned across AI accelerators, memory, networking, energy infrastructure and other critical layers of the AI stack.  

Prior Top AI Stock reports have identified five positions up more than 100% year to date and ten positions up more than 50% for the I/O Fund, with many held at high allocations. By comparison, the Nasdaq-100 is up just 13% YTD.  

Don’t miss out on the AI trade. Learn more here.

Please note: The I/O Fund conducts research and draws conclusions for the company’s portfolio. We then share that information with our readers and offer real-time trade notifications. This is not a guarantee of a stock’s performance and it is not financial advice. Please consult your personal financial advisor before buying any stock in the companies mentioned in this analysis.    

Leo Miller, AI and Semiconductor Investment Writer at I/O Fund, contributed to this analysis.

👉🏻 Share with a Fellow Investor
Help someone else benefit from this insight.

Recommended Reading:

head bg

More To Explore

Newsletter

Abstract visualization of a flowing stream of digital tokens, numbers, and symbols representing AI token processing and inference demand growth.

AI Token Demand is Shattering Forecasts 

Total annual token processing is no longer measured in billions or trillions of tokens, but in the quadrillions and beyond. As annual token processing is now tracked in units with 15 trailing zeros, i

July 30, 2026
TSMC N3 wafer technology connected to Intel EMIB advanced packaging in an AI semiconductor manufacturing graphic.

Nvidia and Google Are Crowding TSMC’s N3 Node - Can Intel Fill the Gap?

Nvidia is moving its next-generation Rubin GPUs from 4nm to 3nm, yet Google’s latest TPUs are already on N3 and are expected to remain there. Meanwhile, a growing number of AI CPUs from Nvidia, Amazon

July 26, 2026
Illustration of Intel EMIB-T advanced packaging connecting AI compute dies and HBM memory as an alternative to TSMC CoWoS.

Intel vs TSMC: How CoWoS Packaging Constraints Could Create an Opportunity for Intel Foundry 

Taiwan Semiconductor (TSMC) is the single, most important company to the AI industry. However, to compete with the incumbent, Intel does not need to beat TSMC at leading-edge manufacturing. It only ne

July 24, 2026
Amazon, Meta, Microsoft, and Google displayed with financial charts, illustrating rising AI capex and growing free cash flow pressure across Big Tech.

Big Tech’s Free Cash Flow is Turning Negative – Who's Next? 

Big Tech’s AI revenue is accelerating, but free cash flow is moving sharply in the opposite direction. Across Google, Microsoft, Meta and Amazon, capex is rising much faster than operating cash flow a

July 19, 2026
Illustration of Google, Microsoft, Meta, and Amazon stock dashboards against a digital circuit-board background, symbolizing Big Tech earnings, AI growth, and investor performance.

Big Tech Earnings Preview: Is AI Monetization Finally Catching Up to Capex?

The most pronounced difference between 2026’s tech rally compared to rallies in the past is which companies have been left out of it. The names most associated with the AI trade have hardly participat

July 17, 2026
Side-by-side image of an NVIDIA CMX server and a CXL memory expansion card against a data center background.

Nvidia, CXL, and the Battle to Improve AI Inference Economics

This is Part 2 of our two-part series on AI inference economics. In Part 1 — Why Nvidia's Next AI Battle Is About Tokens per Watt, we laid out why tokens per watt has become the defining metric for in

July 12, 2026
NVIDIA BlueField networking platform card shown on a green digital network background.

Why Nvidia’s Next AI Battle Is About Tokens per Watt 

As hyperscalers move from building AI infrastructure to monetizing it, tokens per watt helps to reflect if revenue is scaling and if profitability is improving. Offload engines can increase tokens per

July 10, 2026
Micron HBM3E chip with glowing data streams representing AI memory demand and high-bandwidth memory technology

Micron Is Up 900%. Here’s Why the AI Memory Trade May Still Have Room to Run

Over the past 10 months, memory chip stocks have gone from being solid beneficiaries of the AI boom to capturing a massively outsized piece of the return pie. The inflection in Micron’s performance de

June 26, 2026
Fighter jets flying over a city with smoke rising, overlaid by a rising S&P 500 chart, illustrating markets climbing despite the Iran war.

Why the S&P 500 Shrugged Off the Iran War — and What Could Finally Break the Rally 

On February 28th, the U.S. went to war with Iran, and the market was handed the kind of shock it hasn't contended with for years. The conflict set off a chain reaction across the region: an ongoing su

June 19, 2026
AI cloud and GPU infrastructure visualization representing NVIDIA, CoreWeave, and Nebius in the circular financing of the GPU boom

Nvidia, CoreWeave, and Nebius: Inside the Circular Financing of the GPU Boom

Neoclouds are one of the more hotly debated AI business models, with CoreWeave and Nebius being the two most widely recognized names. These companies have seen their sales, backlog, and share prices s

June 12, 2026
newsletter

Sign up for Analysis on
the Best Tech Stocks


Copyright © 2010 - 2026