Last week, AMD offered more details on the release of their groundbreaking GPUs with little fanfare in the markets – which is par for the course as AMD has a history of being forgotten about until the company can no longer be ignored.
Five years ago, I dubbed AMD the “Dark Horse” for my premium research members as the company had a mere 4% share in the CPU-data center and was up against the near-monopoly of Intel. The term “Dark Horse” refers to a competitor that unexpectedly achieves victory as I was predicting AMD would eventually overtake Intel.
Two quarters ago, AMD posted CPU server market share of 39.4% — officially surpassing Intel.
In the technology industry, the probability of an underdog successfully taking on a first-place contender with a formidable lead is incredibly rare. Yet, there is an element of catching the market off guard that helps to compound the returns. The opposite of this is known as a crowded trade.
Does AMD have what it takes to overtake Nvidia on stock performance in the next few years? Most investors assume Nvidia will continue to dominate — and AMD will remain a distant second. In this piece, I’ll walk you through why AMD’s positioning in the AI cycle could lead to an outcome few are prepared for.
Background on what AMD Achieved
When Lisa Su became CEO of AMD in 2014, the company was on the brink of bankruptcy, operating at a loss from 2012 to 2017. The huge bets the company made with the Zen architecture were bold, and saved the company from going under.

Examining how AMD was able to stage the comeback through architectural changes in CPU architecture, process technology, and chiplets is key for investors as not only did it result in over 3,600% returns in 10 years, but the company is now setting up to become a strong contender in the GPU server market.

Source: YCharts
AMD Released the Zen 2 Architecture in 2019:
Five years after Lisa Su became CEO, AMD was preparing to not merely survive but rather to rival Intel. The Zen 2 architecture was an important release that allowed AMD to leapfrog Intel with a 7nm chip while Intel was still producing 14nm and 10nm chips. Because 7nm are twice as dense as 14nm, AMD was able to release a 64-core server chip and 128 threads rather than AMD’s previous 32-core server chip. Up until early 2019, Intel’s offering has been a 28-core server chip and 64 threads. The result of being first to the 7nm is that AMD was able to produce a more power efficient chip that allowed more cores.
The Zen-2 architecture also introduced a multi-chip module that used the most advanced technology where it’s needed most by combining 7nm chiplets with a 14nm die. This was quite a competitive leap as Intel was still using a monolithic design.
In this case, the 14nm was leveraged for memory controllers because the central hub runs input/output (I/O) and memory better. This helped AMD beat Intel on memory bandwidth. The design also greatly improved performance by putting the L2 cache on the core and the L3 cache across the core. Overall, these design improvements lower the power required while increasing the performance as it requires fewer NUMA hops, which in turn, increases instructions per clock, and this ultimately reduces latency.
AMD’s second-generation EPYC server processors sparked the company’s comeback with 1.8 to 2 times the performance advantage of Intel’s Xeon processors, but perhaps most importantly, EPYC 2nd Gen was at half the cost as Intel in some instances. Undercutting Intel on price became a virtuous cycle as driving down costs means more chips will be bought from AMD.
In a 2021 webinar on AMD’s stock that I held for Premium Members, I noted at the time that a third-party analyst named Michael Larabel benchmarked AMD as being 14% faster than Intel while costing about 30% less. The result is that for every $1.00 Rome chip sale, Intel lost $2.25 in Xeon SP sales. The savings can then be deployed to buy more Rome chips to further depress Intel’s revenue.
Since the Rome Series, AMD has been able to take more market share with the Milan Series and Bergamo Series with improvements such as 3D stacking in Zen3, tripling the L3 cache size while only adding four clock cycles of latency, and further customizing CPUs for cloud native workloads with less cache and more performance per watt. Genoa was the 4th generation, and provided more cache for general purpose workloads.
AMD versus Nvidia: Why Memory Gives AMD an Inference Edge
The word “inference” will come up a lot in the coming years for AI investors, and thus, it makes sense to have a brief discussion on how it differs from training.
- Training:
Training is the process of a model learning patterns from labeled data through internal parameters (called weights). There is forward and backward pass or propagation for updating the parameters. This phase is computationally intensive, requiring significant memory and parallel processing power.
Training is where Nvidia’s strengths are nearly insurmountable as the leader in combining parallel processing (CUDA) cores with matrix computations (Tensor Cores). Over the past few years, Nvidia has increased compute power by an order of magnitude to the point of defying Moore’s Law with architectural changes such as tensor cores and lower precision floating points.



