Recently, I’ve reiterated my $20 trillion market cap thesis for Nvidia, which implies upside of roughly 310% over the next four years. However, my thesis does not hinge on Nvidia reaching that milestone through hardware alone. Instead, the thesis hinges on software advancements and the recurring revenue that will inevitably come from Nvidia’s lead in robotics and simulation. I have emphasized the growing importance of Nvidia’s software business relative to hardware since 2023.
The distinction cuts both ways. By arguing that software is central to the $20 trillion thesis, I am also implying that Nvidia’s hardware moat becomes less effective over time. Seven years ago, when Nvidia was still a roughly $100 billion company trading near $3.15 on a split-adjusted basis, my original thesis for why it could become the world’s most valuable company centered on the CUDA moat. At the time, I wrote: “Developers will self-regulate the number of competitors for processing units due to a need for a universal platform that supports all frameworks.”
My firm, the I/O Fund, has held Nvidia through the full seven-year journey, sometimes at an allocation as high as 20%, through both remarkable upside and equally remarkable downside (may be hard to remember, but the stock was down 60% in 2022 when I publicly defended the stock).
The thesis on Nvidia’s hardware moat has played out exceptionally well, but that also highlights one of the biggest risks investors face, which is becoming emotionally attached to a winning stock. While I still believe Nvidia will reach $20 trillion by 2030, I believe much of that 310% return is likely to be back-half weighted in the years of 2028-2030. This is what separates investors from AI enthusiasts. While an AI enthusiast can sit back, relax and discuss specifications and other fandom, an investor must always answer — is my capital better deployed elsewhere?
Is Nvidia Stock Still the Best AI Stock in 2026?
So far in 2026, answering the question of where to deploy capital has not been easy to answer. Nvidia’s stock only recently turned positive; the QQQs are barely positive this year, as is the same for many tech-related ETFs such as IVES, GRNY and ARKK.
In sharp contrast, the I/O Fund is up roughly 33% year-to-date, reflecting a willingness to follow the opportunities as they shift across the AI landscape. We count recent winners such as Bloom Energy, which is up 1100% since our initial entry, an optical networking name we highlighted ahead of its 2026 surge, now up nearly 300% YTD and 650% since our lowest entry in November. Plus, a photonics position we doubled down on in January with a 10% allocation that has since gained more than 130% year-to-date.
The same framework that surfaced those opportunities is what tells me Nvidia’s 2026 setup may no longer be as rewarding as what I can find elsewhere. The analytical case comes down to three things: the CUDA moat matters less with inference, custom silicon is gaining market share, and the delay in Rubin creates uncertainty at exactly the wrong moment.
On the flip side, the valuation is lower than its historic average, and in a volatile market, Nvidia could still stand out simply by continuing to post stronger earnings growth than most of large-cap tech. The company will remain the dominant system-level player in AI, and the CUDA moat will certainly not vanish overnight.
The debate, in my view, is not about whether Nvidia stays important. It is about whether the return profile is still as compelling as what can be found elsewhere in the AI trade.
CUDA Matters Less as AI Inference Takes Over
In 2018, my original thesis on why Nvidia can become the world’s most valuable company was centered on the moat the CUDA platform provides when I stated: “Nvidia is already the universal platform for development, but this won’t become obvious until innovation in artificial intelligence matures. Developers are programming the future of artificial intelligence applications on Nvidia because GPUs are easier and more flexible than customized TPU chips from Google or FGPA chips used by Microsoft […] When artificial intelligence matures, you can expect data center revenue to be Nvidia’s top revenue segment. Despite the corrections we’ve seen in the technology sector, and with Nvidia stock specifically, investors who remain patient will have a sizeable return in the future.”
At the time, Nvidia’s data center revenue was 1/6th of Intel’s — whereas today, the AI juggernaut reported $194 billion in data center revenue compared to Intel’s $17 billion. Although you could pontificate on the many defensible design elements of Nvidia’s AI systems, one way to simply describe this historic ascent is that the mature libraries and frameworks from CUDA makes it hard for an engineer to go anywhere else. Notably, that’s not a regurgitated thesis, but rather my thesis on the stock implications of the CUDA moat pre-dated the Street and AI experts by many years. by many years. That’s important because I am shelving that thesis as the inference market approaches.
There is an incoming shift to my original investment thesis from 2018.
Programming GPUs with the CUDA platform is primarily a training exercise as this is the phase where engineers are experimenting and need the developer ecosystem, including extensive tools like cuDNN, NCCL, debugging, custom kernel support, and CUDA’s massive libraries. The ecosystem has been built for over 20 years, has over 6 million developers contributing and every ML framework is first optimized for CUDA. The switching costs today remain extraordinarily high for engineers.
To contrast, inference is repetitive to where once a model is trained, the model is running millions of times per day. Serving platforms and inference frameworks like vLLM and TensorRT-LLM reduce dependency to develop on a specific software platform, like Nvidia’s CUDA.
Training a frontier model is a one-time, multi-month event. Inference, by contrast, is the revenue-generating phase. Every ChatGPT query, every Copilot suggestion, every Waymo autonomy decision is inference. As frontier labs reach the limits of practical model size and enterprise AI adoption scales, inference workloads are projected to grow several times faster than training workloads through the rest of the decade. The segment where CUDA’s moat is strongest is becoming a smaller share of total compute, while the segment where it is weakest is becoming the larger share.
There is also more of a push toward open standards for the inference phase to reduce dependency on hardware specific code for serving paths, as tools like ONNX runtime, vLLM and the compiler Triton help to export models (or compile them) to be run agnostically on any AI accelerator.
In response to CUDA’s moat weakening in the inference phase, Nvidia has pushed for their inference stack to remain proprietary by offering inference optimization software called TensorRT-LLM. TensorRT-LLM analyzes and optimizes LLMs to improve performance by fusing multiple operations into a single GPU kernel, selecting the optimal precision and optimizing memory usage for the key-value cache. Overall, Nvidia states this leads to 5X faster model performance for inference.
However, something to consider is that Nvidia is needing to make this new attempt at preserving its ecosystem as the CUDA empire will not neatly hold as the inference market plays out. The open-source market is growing to become a serious contender to proprietary optimization software like TensorRT-LLM, as alternatives that are more community driven are available and accomplish something similar, such as vLLM and SGLang. Both have moved from research-project status to production deployment at major AI operators, with vLLM in particular now powering inference at some of the largest LLM serving workloads outside of the hyperscalers themselves. Furthermore, large inference players like Cloudflare can build their own custom engines.
The point is not to be an alarmist, but rather to note when the piece most central to my original thesis is shifting. CUDA will remain the most popular software development platform in AI by a wide margin, however, the freedom to go elsewhere is something Nvidia has not contended with at this level.
Custom Silicon is Undeniably Increasing in Market Share
A few months back, the market had a brief scare around what Google’s TPU v7 Ironwood might mean for Nvidia’s grip on AI compute. The concern was not simply that Google had built another custom chip, but that Ironwood was introduced as the first TPU designed specifically for inference.
At the time, Google emphasized better power efficiency and stronger “intelligence per dollar” for serving workloads. Ironwood scales up to 9,216 chips, delivers 42.5 exaflops in its largest pod, and Google has paired it with software support such as vLLM on TPU, reinforcing the idea that inference is becoming a more open and cost-sensitive market than training. We covered this more in the write-up: “This AI Stock is Set to Surge from Inference Demand.”
Although Ironwood v7 offers major headway in narrowing the performance gap with Nvidia on inference workloads, the reality is that custom silicon programs require long development cycles. Designing the chip is only the initial stages, and from there, hyperscalers need to optimize the compiler stack, optimize frameworks and also validate performance at scale. The result is a far slower product road map that typically lags Nvidia’s current generation of GPUs. This lag puts additional emphasis on Nvidia delivering on time.
Why the Advantages of Custom Silicon Outweigh Development Timelines
Nvidia’s data center GPUs carry gross margins above 70%. For companies spending $50-100 billion annually on AI infrastructure, the savings from moving even 20-30% of inference workloads to in-house silicon compounds into tens of billions of dollars per year. That math is driving Google and Amazon to accept slower product cycles in exchange for architectural independence. It is also the math incentivizing Meta and Microsoft to follow suit. Perhaps most importantly, the inference market will offer a catalyst for custom silicon compared to training because workloads are more specific, and cost savings can be achieved at massive volumes.
Below is what a few industry analysts are predicting. Although I believe these are aggressive, they help to illustrate the challenges in front of Nvidia.
Counterpoint Research believes that by 2028, custom silicon will cross the 15-million mark to surpass GPU shipments as the top 10 hyperscalers will have deployed 40 million AI server compute ASIC chips cumulatively during 2024-2028, stating:
“What is also supporting this unprecedented demand is AI hyperscalers building significant rack-scale AI infrastructure based on their in-house stacks, such as Google TPU Pods and AWS Trainium UltraClusters, enabling them to operate as one supercomputer.”
TrendForce is the most aggressive forecast, stating GPU-based AI servers will account for 69.7% of shipments in 2026 with ASIC-based servers rising to 27.8%. This doesn’t account for GPU market share from AMD, which if you put that at 10%, would result in Nvidia’s market share being 59.7%.
With the information that I have today, these forecasts could be too aggressive.
According to Broadcom, they’ll see $100B in AI chip revenue in 2027 and we’ve modeled another $50B in networking. If we allocate $45B base case to AMD and go with what we know of Nvidia’s stated trajectory to $1 trillion in revenue, then the split looks something more like this for 2027:
- NVDA $500B
- AVGO $150B to $200B (assuming mgmt team was being conservative we will use the $200B number)
- AMD $45B
- Total among top 3 silicon providers: $745B with NVDA at 67% market share versus the 59.7% implied above
However, one data point that complicates things is MediaTek could see 150,000 CoWoS wafers in capacity in 2027, compared to 20,000 in 2026. Thus, the landscape is evolving in terms of the number of competitors.
Notably, the level of erosion may be up for debate, but the most probable outcomes do not favor Nvidia continuing to dominate AI accelerator sales at the level it has in the past. In training, Nvidia represented 90% of workloads.





