Blogs -NVIDIA’s GB200s for up to 27 Trillion Parameter Models: Scaling Next-Gen AI Superclusters

NVIDIA’s GB200s for up to 27 Trillion Parameter Models: Scaling Next-Gen AI Superclusters


March 21, 2025

author

I/O Fund

Team

Supercomputers and cutting-edge AI data centers are fueling the artificial intelligence (AI) revolution. Large-scale systems need comprehensive builds that are increasingly integrated to meet the evolving demands of complex workloads. As AI applications become more sophisticated, the need for infrastructure that's not only incredibly powerful but also energy-efficient is growing exponentially. Innovations like NVIDIA’s GB200 are designed to deliver the scalability needed for next-generation AI superclusters.  

At the 2025 NVIDIA GPU Technology Conference (GTC), VP and Chief Architect of Systems, Mike Houston, and Senior Director of Applied Systems Engineering, Julie Bernauer, discussed large-scale systems design principles in their May 18 presentation, “Next-generation at Scale Compute in the Data Center.”

NVIDIA’s First Rack-Scale Product is the GB200 Superchip

The NVIDIA Grace Blackwell 200 (GB200) Superchip combines two Blackwell GPUs and one Grace CPU. It’s NVIDIA’s first rack-scale product. The NVIDIA GB200 NVL72 is a configuration and rack-scale, liquid-cooled AI computing platform, which is purpose-built for AI training and inferencing, handling up to 27 trillion parameters for generative AI models. The GB200 includes base components like Grace Hopper compute trays, NVLink switches (a connector in the middle of the rack linking all GPUs) and cable cartridges (literally miles of cables in the back to tie everything together). The design includes quantum switches for InfiniBand (a high-speed network for linking clusters) and spectrum switches for Ethernet.

AI 101: What are Clusters and Superclusters?

Clusters 101: are a network of independent computers (called nodes) connected by a high-speed network. A cluster serves as a unified resource, as they are separate machines configured to work together to act as a single powerful computing system. They are often used for parallel processing, which breaks down a large task into smaller parts distributed across the nodes, enabling faster processing than just a single computer could do. A key benefit of a node is high availability, meaning if one node (computer) fails, the other nodes can take over its workload, ensuring that the system remains operational. High-performance compute (HPC) clusters are used for tasks like research, scientific simulations and AI training.

Join thousands of investors who trust I/O Fund’s expert stock analysis on AI, semiconductors, cryptocurrency, and adtech — sign up for free! Click here!

Superclusters 101: are very large clusters that may be comprised of hundreds to thousands of GPUs through many data centers. For example, Elon Musk’s xAI supercomputer Colossus, powered by 100,000 NVIDIA GPUs, is definitely a supercluster.

DGX started as single machines for AI but evolved into clusters for AI training. Pre-training can involve superclusters, but post-training can still involve 16,000 GPUs with smaller setups for fine-tuning and inference using trained AI to answer questions.

NVIDIA AI and HPC platform architecture diagram featuring the GB200 NVL72 SuperPOD

NVIDIA AI & HPC Platform architecture diagram, featuring GB200 NVL72 SuperPOD.
Source: NVIDIA

Optimizing the Benefits of Rack-Scale Architecture with GB200

NVIDIA’s GB200 NVL72 is a rack-scale system. Rack-scale designs a whole rack as one big, coordinated unit, not just random machines stuck together. Rack scale refers to integrating and compressing systems that may span across multiple servers, storage and networking devices onto a single server rack. GB200 can replace or consolidate a large number of GPU compute servers. This provides many benefits, including:

  • Improved GPU Density: The GB200 NVL72 contains 72 Blackwell GPUs, and 36 Grace CPUs interconnected with NVLink, NVIDIA’s proprietary high-speed (130 TB/s) signaling interconnect that enables all 72 GPUs and 36 CPUs to act as a single massive GPU. It's designed to offer exceptional performance in AI training and inference for large language models (LLMs).
  • Performance: The GB200 delivers up to 720 petaFLOPs for AI training and 1.4 exaFLOPs for inference. Since all components are within proximity in a single rack, communication between components has much lower latency, which is especially beneficial in data-intensive tasks, reducing bottlenecks and improving data throughput.
  • Increased Efficiency: Rack-scale architecture allows for better utilization of hardware by pooling resources to optimize performance. Consolidating resources within a single rack reduces the need for separate units, saving space and power in the data center.
  • Easier Management: Centralized management of the entire rack's resources simplifies setup and maintenance, also enabling automation tools for scaling, provisioning and monitoring to reduce manual interventions.
  • Cost Efficient: Fewer servers, storage, networking equipment, physical space, cooling, and energy usage save money. As IO Fund discussed in its article “AI Power Consumption: Rapidly Becoming Mission-Critical," the GB200 is “expected to consume 2,700W”, which can add dramatically to operating expenses, especially without rack-scale architecture.
  • Future Proofing: Rack-scale architecture enables the integration of evolving technologies as components can be switched out, repaired and upgraded, enabling more adaptability for future growth.
  • Unified Power and Cooling: Housing multiple components within a single rack reduces the complexity of cooling systems and improves energy efficiency to lower operational costs.

Scaling Up AI Factories with DGX SuperPOD, Reference Architecture and Fabric

At the 2025 NVIDIA GPU Technology Conference (GTC), NVIDIA unveiled its next-generation DGX SuperPOD AI infrastructure. In the “Next-generation at Scale Compute in the Data Center” presentation, VP and Chief Architect of Systems, Mike Houston, and Senior Director of Applied Systems Engineering, Julie Bernauer, spoke about

The SuperPOD is NVIDIA’s all-in-one HPC solution designed to handle the massive computational needs of AI models and simulations. Grace Blackwell nodes are the building blocks of the SuperPOD. When scaling up clusters and superclusters, there are three factors to consider. Reference architecture is comprised of pre-tested system designs that serve as a blueprint for new data center deployments to ensure optimal installation and performance, accelerating time to the first token.

Fabric refers to the data center’s network infrastructure that connects all the servers and devices enabling them to seamlessly communicate with each other to reduce latency between components, especially GPUs. Cooling is critical in large data centers. Liquid cooling is preferred to manage the heat produced by thousands of GPUs as it is much more efficient for high-density platforms. Future GPU architectures aim for higher density and more efficient connectivity to push the limits of AI computation.

The I/O Fund recently entered five new small and mid-cap positions that we believe will be beneficiaries of this AI spending war. We discuss entries, exits, and what to expect from the broad market every Thursday at 4:30 p.m. in our 1-hour webinar. For a limited time, get $110 off an Annual Pro plan with code PRO110OFF [Learn more here.]

Please note: The I/O Fund conducts research and draws conclusions for the Fund’s positions. We then share that information with our readers. This is not a guarantee of a stock’s performance. Please consult your personal financial advisor before buying any stock in the companies mentioned in this analysis.

Recommended Reading:

head bg

More To Explore

Newsletter

Futuristic AI data center filled with server racks and glowing network pathways, representing large-scale AI infrastructure investment and rising capital expenditures to meet growing demand.

AI Capex to Hit $1 Trillion – And Estimates Are Still Too Low

Big Tech capex is the driving force behind the AI infrastructure trade, yet Wall Street has repeatedly underestimated the sheer scale of the buildout. Currently, in 2026, the guidance for $732.5 billi

August 05, 2026
Abstract illustration of layered AI computing hardware processing digital tokens, symbols, and data streams representing AI infrastructure and inference demand growth.

Token Growth is Surging - Here Are the Beneficiaries 

The reality of AI demand growth has shattered early estimates for token processing, yet expectations continue moving up and to the right. In the second installment of our token processing series, we e

July 31, 2026
Abstract visualization of a flowing stream of digital tokens, numbers, and symbols representing AI token processing and inference demand growth.

AI Token Demand is Shattering Forecasts 

Total annual token processing is no longer measured in billions or trillions of tokens, but in the quadrillions and beyond. As annual token processing is now tracked in units with 15 trailing zeros, i

July 30, 2026
TSMC N3 wafer technology connected to Intel EMIB advanced packaging in an AI semiconductor manufacturing graphic.

Nvidia and Google Are Crowding TSMC’s N3 Node - Can Intel Fill the Gap?

Nvidia is moving its next-generation Rubin GPUs from 4nm to 3nm, yet Google’s latest TPUs are already on N3 and are expected to remain there. Meanwhile, a growing number of AI CPUs from Nvidia, Amazon

July 26, 2026
Illustration of Intel EMIB-T advanced packaging connecting AI compute dies and HBM memory as an alternative to TSMC CoWoS.

Intel vs TSMC: How CoWoS Packaging Constraints Could Create an Opportunity for Intel Foundry 

Taiwan Semiconductor (TSMC) is the single, most important company to the AI industry. However, to compete with the incumbent, Intel does not need to beat TSMC at leading-edge manufacturing. It only ne

July 24, 2026
Amazon, Meta, Microsoft, and Google displayed with financial charts, illustrating rising AI capex and growing free cash flow pressure across Big Tech.

Big Tech’s Free Cash Flow is Turning Negative – Who's Next? 

Big Tech’s AI revenue is accelerating, but free cash flow is moving sharply in the opposite direction. Across Google, Microsoft, Meta and Amazon, capex is rising much faster than operating cash flow a

July 19, 2026
Illustration of Google, Microsoft, Meta, and Amazon stock dashboards against a digital circuit-board background, symbolizing Big Tech earnings, AI growth, and investor performance.

Big Tech Earnings Preview: Is AI Monetization Finally Catching Up to Capex?

The most pronounced difference between 2026’s tech rally compared to rallies in the past is which companies have been left out of it. The names most associated with the AI trade have hardly participat

July 17, 2026
Side-by-side image of an NVIDIA CMX server and a CXL memory expansion card against a data center background.

Nvidia, CXL, and the Battle to Improve AI Inference Economics

This is Part 2 of our two-part series on AI inference economics. In Part 1 — Why Nvidia's Next AI Battle Is About Tokens per Watt, we laid out why tokens per watt has become the defining metric for in

July 12, 2026
NVIDIA BlueField networking platform card shown on a green digital network background.

Why Nvidia’s Next AI Battle Is About Tokens per Watt 

As hyperscalers move from building AI infrastructure to monetizing it, tokens per watt helps to reflect if revenue is scaling and if profitability is improving. Offload engines can increase tokens per

July 10, 2026
Micron HBM3E chip with glowing data streams representing AI memory demand and high-bandwidth memory technology

Micron Is Up 900%. Here’s Why the AI Memory Trade May Still Have Room to Run

Over the past 10 months, memory chip stocks have gone from being solid beneficiaries of the AI boom to capturing a massively outsized piece of the return pie. The inflection in Micron’s performance de

June 26, 2026
newsletter

Sign up for Analysis on
the Best Tech Stocks


Copyright © 2010 - 2026