As hyperscalers move from building AI infrastructure to monetizing it, tokens per watt helps to reflect if revenue is scaling and if profitability is improving. Tokens are the revenue side of the inference equation, while power consumption (watts) represents the associated costs. The more tokens a hyperscaler can generate from each unit of power, the more revenue it can drive without a proportional increase in infrastructure costs.
Nvidia’s offload-engine approach is designed to address the problem of accelerators sitting idle, waiting for KV cache, memory or data movement. In other words, offload engines can increase tokens per watt by improving utilization and by increasing the amount of inference revenue generated from the same power footprint. The result is improved margins as revenue generation scales faster than costs.
It’s become quite clear that power is no longer an abundant resource. As new capacity becomes harder to secure, inference growth will be constrained by how efficiently hyperscalers use the power and accelerators they already have.
When it comes to improving tokens per watt, memory is the biggest constraint, rather than the raw power offered by GPUs or XPUs. If Nvidia’s GPUs or Big Tech’s XPUs are not supplied with sufficient memory to constantly perform inference tasks, tokens per watt falls as these accelerators are underutilized.
To solve this issue, offload engines are being actively designed into the rack-scale architectures of AI infrastructure leaders like Nvidia. Meanwhile, other players are building open-standard offload engine solutions through Compute Express Link (CXL).
In this analysis, we break down why XPU underutilization matters, how offload engines improve tokens per watt, and which architectures may create the strongest opportunity for investors.
Nvidia’s Offload Engines and the KV Cache: Solving AI’s Inference Memory Bottleneck
While Nvidia has already been embedding ‘offload engines’ such as its SmartNICs and DPUs to offload data processing from host CPUs, the latest offload engines target GPU memory to accelerate inference. These solutions look to address one of the main causes of XPU underutilization: the growing size of model KV caches, which consume the majority of their working memory during inference.
In an ideal world, the KV cache would be fully stored in the memory tier that XPUs can access most quickly, which is high-bandwidth memory (HBM). This minimizes latency, optimizes XPU utilization, and maximizes tokens per watt.
However, as models and workflows increase in complexity, such as for multi-step reasoning or agentic deployments, KV cache can exceed HBM memory capacity. In this case, the KV cache is either offloaded to slower tiers of the memory stack, introducing latency, or recomputed at each request, wasting resources on work that has already been performed – both key factors in reducing XPU utilization and reducing tokens per watt.
Generally speaking, offload engines aim to keep XPUs fed with KV cache more efficiently, improving XPU utilization and tokens per watt.
The Memory Wall: AI Compute Gains Are Outpacing Memory
A key reason why offload engines can be important solutions is due to the ‘memory wall’. Over the 20 years ending in 2023, accelerator processing power (Peak FLOPs) increased by 3X every two years. However, memory bandwidth has increased by only 1.6X every two years, and memory capacity has increased by only 2X every two years.
With this, the theoretical speed at which XPUs could produce tokens is growing faster than the rate they can in practice, because memory is not feeding them data fast enough. Thus, memory-driven XPU underutilization is increasing.

We can continue to see this play out with Rubin. Nvidia notes that Rubin’s inference processing power is 5X higher than that of Blackwell. However, as Nvidia moves from HBM3E in Blackwell to HBM4 in Rubin, overall HBM bandwidth only increases by 2.8X while capacity rises just 1.5X. Although HBM capacity and bandwidth are increasing, the memory wall is firmly intact.
Amid this, one of the core issues in addressing the KV cache bottleneck is that HBM cannot be added independently of accelerators, as they are packaged together. While adding more XPUs increases HBM capacity and allows operators to serve more concurrent requests, doing so also means that more XPUs go underutilized as the ratio of processing power to memory does not change.
Additionally, due to the intense memory shortage, HBM has become an increasingly large chunk of AI accelerator bill of materials (BOM). In the B200, HBM represents approximately 52% of total BOM. In Rubin, this is expected to increase to 62%. As this percentage rises, hyperscalers are paying more for memory, even though the pace of computing power growth continues to exceed memory capacity and bandwidth growth. In other words, hyperscalers are forced to spend a higher percentage of their capex on memory even as memory-driven XPU underutilization further deteriorates.
In turn, hyperscalers need to find a way to use memory more efficiently and increase memory capacity independent of HBM. This helps alleviate KV cache bottlenecks, sidestep the increasing costs of scaling memory through XPU purchases, and expand revenue and margins amid power constraints. This is the exact value that offload engine solutions like Nvidia’s CMX intend to provide.
I/O Fund Premium members recently received an 18-page deep dive on the investable shifts in the networking stack and what companies are positioned to benefit. Sign up here.
Tokens Per Watt: Scaling Revenue and Margins
For hyperscalers, improving tokens per watt is important from three key perspectives; generating higher inference revenue, increasing inference margins, and better allocating capex spend.
Why Offload Engines Can Increase Revenue Amid Power Constraints
While power is an inference cost, it is also a required input for token and revenue generation. However, power availability is a key AI bottleneck. Notably, ERCOT is tracking more than 438 GW of large load interconnection requests, with nearly 89%, or approximately 390 GW coming from data centers alone. Meanwhile, only 23 GW of capacity were added between 2024-2025, demonstrating the massive gap between data center energy demand and the ability of the grid to meet that demand quickly.
Additionally, GPU power consumption is increasing with each new generation of chips. H100 power consumption was approximately 700W per chip. In Blackwell and Blackwell Ultra, this approximately doubles to 1,200W-1,400W per chip. Rubin will come with another substantial increase to around 2,300W per chip. As each newer chip consumes more power and has a higher upfront cost, the penalty of underutilization rises.
With such a massive backlog of power capacity demand and rising power consumption per chip, hyperscalers that generate more tokens within the fixed power envelope they have already secured pull one of the most immediate levers to increasing inference revenue. This creates a strong incentive for hyperscalers to increase token throughput using offload engine solutions.




