This is Part 2 of our two-part series on AI inference economics. In Part 1 — Why Nvidia’s Next AI Battle Is About Tokens per Watt, we laid out why tokens per watt has become the defining metric for inference profitability: as memory and power grow scarce, the hyperscalers that generate the most tokens from a fixed power footprint scale revenue faster than cost. We also discussed how the “memory wall,” defined as compute outpacing memory bandwidth and capacity, leaves expensive accelerators underutilized. One solution is offload engines, like those offered by Nvidia, that lift tokens per watt by keeping XPUs fed with KV cache instead of sitting idle or recomputing.
In the article below, we turn from the why to the how and the who. Two architectural paths are competing to solve the KV cache bottleneck: Nvidia’s proprietary CMX platform and the open, vendor-agnostic CXL standard. Although they tackle the same problem, each approach points to different sets of beneficiaries.
With CMX, Nvidia is designing a tightly co-designed platform leveraging its BlueField DPUs and software optimizations to control KV cache allocation and movement, aiming to increase KV cache capacity and sharing across the pod.
CXL offers a vendor-agnostic route to share memory resources coherently across the pod, including for KV cache tasks, enabling similar gains in inference throughput.
In the analysis below, we explain how both CMX and CXL work, touch on the broader CXL ecosystem, and provide our take on which of the two architectures we believe creates the strongest opportunity for investors.
Nvidia CMX: Adding Shared Memory for KV Cache Offloading
To solve for KV cache bottlenecks, Nvidia has developed its CMX Context Memory Storage Platform. CMX uses an SSD-based enclosure that is connected to Rubin GPUs and Vera CPUs over Spectrum-X Ethernet.
The hardware offload engine in CMX is the BlueField-4 DPUs. CMX incorporates 64 BlueField-4 DPUs, which connect to approximately 9,600 TB of SSD storage capacity per rack. Through this, CMX extends effective GPU KV cache capacity to a massive degree with relatively inexpensive SSDs. For reference, one GB200 NVL72 rack comes with 13.4 TB of HBM capacity.
BlueField-4 is dedicated to controlling how the KV cache moves between the high-capacity SSD tier and GPU memory. It is enabled by a mixture of software to manage KV caches, including DOCA Memos, Dynamo, and Nvidia Inference Transfer Library (NIXL). Together, they determine the proper memory tier that certain parts of the KV cache should reside on at a given time, as well as when the KV cache should be sent to GPUs to improve utilization.
By employing offload engines in the form of BlueField and specialized software, GPU KV cache capacity greatly increases, and GPUs can retrieve it quickly, avoiding idle and recompute time and improving utilization. Additionally, CMX allows for pod-wide KV cache sharing. Nvidia notes that this improves XPU utilization by reducing KV cache duplication and stranded memory capacity between nodes.
Nvidia says that CMX allows for up to 5X higher token throughput and up to 5X better power efficiency for KV cache operations compared to traditional storage methods—directly targeting the tokens per watt improvements that are critical to hyperscaler inference economics.
Nvidia also created the STX reference architecture to provide a standardized blueprint for how storage vendors should build CMX systems to connect to Vera Rubin racks. This allows an ecosystem of storage providers to proliferate and provide CMX products that can be seamlessly integrated with computing resources, accelerating adoption.
STX is modular, which is key to CMX being a new revenue driver for Nvidia. Because data center operators can add CMX independently of computing resources, capex dollars can be incrementally allocated to directly improving GPU utilization in KV-cache intensive inference and agentic workflows.
We also covered Nvidia’s other offload solution, its new Groq 3 LPX racks, that aim to significantly accelerate inference throughput in our March analysis, Nvidia Stock to See New Growth Catalyst; 35X Faster AI with Groq 3 LPX.
CXL Memory: The Open Standard Alternative to CMX
Compute Express Link (CXL) is another pathway for building offload engine solutions. Unlike Nvidia’s proprietary CMX, it is vendor-agnostic, though it shares the same ideas as CMX: adding memory where the KV cache can be stored and allowing multiple devices to access it.
One of the key benefits of CXL is that it is built on Peripheral Component Interconnect Express (PCIe). PCIe is a universally adopted interconnect standard, making CXL-based systems relatively easy to implement and able to benefit from PCIe’s 16X increase in data rates since 2019, rising from 32 GB/s to 512 GB/s.
As a vendor agnostic solution, CXL is advantaged in the fact that it can allow data center operators to scale memory independently without being further locked into the Nvidia ecosystem. Notably, the Vera CPU supports CXL 3.1, meaning that data center operators using Nvidia hardware have an option to scale memory capacity through this open standard rather than only through its proprietary offerings.
Yole Research estimated that in Q1 2025, two-thirds of servers were CXL-capable, and projects this figure will rise to more than 90% by the end of 2026. Meanwhile, Yole estimates that near 0% of servers are CXL enabled, and expects that percentage to rise to 13% by 2030.
Networking Stocks Show Large Improvements in Inference Throughput With CXL
CXL 2.0 allows for memory pooling and switching, and is the standard currently being deployed or soon to be deployed in data centers. Through this, CPUs can dynamically allocate memory from the pool between processors.
According to data from Marvell, by adding a 16TB DRAM memory pool through its Structera S CXL switches, CXL pooling enabled a 4.8X improvement in inference throughput, greatly increasing GPU utilization. Additionally, this enabled an 82.7% drop in time to first token.
CXL 3.0 will extend beyond memory pooling into memory sharing—where any device, including CPUs, GPUs, and others, can access the same parts of the memory pool simultaneously. Marvell expects to begin sampling its Structera S CXL 3.0 switch in calendar Q3 2026.



