- As token processing has jumped far ahead of industry expectations, compute, networking and power companies are poised to benefit.
- Nvidia Rubin, optical networking, and readily available power are specific solutions that can help data center operators tackle in the huge increase in token processing and inference demand.
- Industry analysts are now forecasting exponential growth in a technology we detailed many months ago.
AI token processing, one of the clearest indicators of inference demand, is blowing past initial expectations. We highlighted in “AI Token Demand is Shattering Forecasts” that Dell raised its 2028 token-processing estimate by 57X, yet actual token processing has already moved far beyond that sharply revised forecast. Google, for example, saw surface-wide token processing rise by 330X from May 2024 to May 2026, while several other players have reported similarly dramatic growth.
Exploding token growth is only one part of the equation. Future GPU generations and model optimizations will continue to improve efficiency and lower cost per token. However, these improvements are driving token consumption higher, creating a cycle where demand growth continues to outpace efficiency gains.
In the second installment of our token processing series, we examine what exploding inference demand means for the AI infrastructure market. While many parts of the AI stack are positioned to benefit, we focus on three specific areas where the implications are significant; compute, networking, and power; highlighting notable companies along the way.
Compute Efficiency Is Becoming Critical to AI Inference Economics
Inference demand is rising faster than the power available to support it, making token throughput per unit of energy one of the most important metrics in the next phase of the AI infrastructure buildout. As a result, compute systems are being designed to maximize token throughput and token-processing efficiency.
As we detailed in our recent article “Why Nvidia’s Next AI Battle Is About Tokens per Watt” increasing tokens per watt is key to hyperscalers growing inference revenue despite power constraints, while also expanding inference margins. Nvidia says Vera Rubin racks will deliver dramatically higher token throughput per GW.
As analysts expect agentic AI to drive the majority of token processing over the coming years, Nvidia says that Vera Rubin is designed to deliver up to 10X more agentic throughput per unit of energy compared to Blackwell. Nvidia reaches this metric by not only increasing raw tokens per second per GW, but also allowing for much higher agent interactivity, resulting in 10X more agents and 2X more tool calls per GW.
Nvidia looks to take inference optimization further through its Vera Rubin + Groq 3 LPX deployments, which it says can deliver up to 35X higher token throughput per MW compared to Blackwell.
For Rubin Ultra, Nvidia has yet to release stats on how it compares to Rubin on token throughput and token per watt, but it will pack 2X more HBM per chip compared to Rubin at 576 GB, and will be available in an NVL576 configuration, connecting eight Rubin Ultra 72 GPU racks. The additional memory and larger NVLink domain should allow models to keep more working memory on high-bandwidth tiers while improving communication efficiency between GPUs. This could increase tokens per unit of energy by reducing the amount of time GPUs spend idle.
Nvidia’s Latest Systems Could Expand Data Center Margins
Data from Morgan Stanley Research supports the idea that moving toward Nvidia’s more advanced systems can deliver significant margin benefits to data center operators. The firm estimates that Blackwell-based data centers process tokens at an approximately 58% net margin. It estimates that this will rise to 78% for Rubin-based data centers, and that Nvidia’s further out Feynman generation will push net margin to 90%.
This comes as processing more tokens within the same power envelope means greater revenue generation at the same level of energy expense. Margin improvements of this, or even close to this magnitude, give data center operators a strong incentive to adopt the latest AI compute systems as token processing soars.

AI Networking Demand Is Accelerating With Agentic AI
Agentic AI and the shift from query-based responses to autonomous agentic workflows is expected to create substantial tailwinds for the networking stack. Increasing tokens consumed per user and per workflow means more data must be exchanged between AI accelerators, CPUs and memory, and between networking fabrics.
Here’s what that means if we look at stats: Arm estimates agentic AI will drive up to a 15X increase in tokens per user and Nvidia concurs AgencyBench has forecast that agentic tasks will eventually require 90 tool calls, 1 million tokens and will result in hours of execution time.
For a more extreme example, and perhaps one of the biggest case studies yet on the token consumption of agentic AI, OpenClaw consumed 603 billion tokens from 7.6 million API calls in one month for an estimated $1.3 million in spend. Related to this, Nvidia CEO Jensen Huang estimated in March that combining reasoning models with agents can increase token consumption by roughly 1 million times versus early non-reasoning workloads.
In another supporting metric, Cisco estimates that performing tasks using AI agents increases wide area network (WAN) traffic by 450% compared to humans. This measures data that flows between end users and data centers where inference takes place, rather than directly looking at networking demands within data centers. However, much of this traffic is ultimately still tied to tokens that data center compute must process and output, increasing traffic that flows through networking equipment within data centers. Specifically, Cisco estimates that 70% of this increased traffic comes from AI inference. Agentic AI results in significantly more data entering and exiting data centers, compounding the networking challenge because each task requires more orchestration between GPUs, CPUs and memory.
Optical Networking Is Emerging as a Critical AI Infrastructure Layer
Within data centers, the requirements are shifting both in scale-out and scale-up domains as data transfer speeds rise and the physical size of AI clusters and pods increase, with optical components becoming a necessity in larger domains as copper hits its physical limits.
Optical transceivers are being used to tackle scale-out requirements as data center operators move to clusters of up to 1 million accelerators, with transceivers and components being a huge growth driver for Lumentum, which saw its sales rise by 90.1% YOY to $808.4 million in its latest quarter.
Co-packaged optics (CPO) is an emerging and longer-term opportunity for Lumentum and other networking players, which will come through both scale-out and scale-up content. One of the key drivers of the scale-up CPO opportunity are pods like the NVL576, where optics helps keep latency at ~320 ns, a 5X improvement to how a similar 576-GPU node could be constructed today, per Corning. Notably, TrendForce is forecasting explosive growth in the CPO and near-packaged optics (NPO) market. Overall, it expects the CPO and NPO market to grow from $100 million in 2025 to $39 billion by 2030. This is equal to an astonishing 230% CAGR, or a 390X increase in five years.




