Earlier this month, Arm unveiled an AGI CPU to address one of AI’s biggest bottlenecks, which is orchestration. During the chatbot craze of 2023-2025, GPUs did most of the heavy lifting while CPUs had become an afterthought. Yet with agentic workloads, which is perhaps the single largest catalyst on the horizon for the AI trade in 2026 and beyond, the importance of CPUs is set to increase.
In agentic workflows, the GPU still handles inference, but between each inference call, the CPU is doing the orchestration – which are best described as handling tool calls, API requests and memory tasks. AI agents are surfacing this new constraint, which is how to prevent latency and underutilized GPUs following the exponential growth of orchestration needs.
For investors, what matters is that CPUs account for 50% to 90% of total latency in workflows, which means the CPU-to-GPU ratio in AI clusters will need to increase. Earlier this year, both AMD and Intel saw analyst upgrades based on the outstripped supply of CPUs leading to higher average sales prices of roughly 10% to 15%. Reuters also reported that Intel’s unfulfilled orders are reaching longer than six months while AMD delivery times are believed to be eight to 10 weeks.
Regarding how Arm fits in, the company’s expertise in lowering power requirements could matter more than the market expects. After years of supplying the architecture IP behind other companies’ CPUs, Arm is preparing to directly compete with its customers and x86 CPU competitors by transitioning to a chip designer themselves. This comes during a time when CPU cores are expected to go up 4X from 30 million CPU cores per gigawatt to 120 million CPU cores per GW.
Brief History of Arm
We covered Arm two years ago in our free newsletter, Arm Stock: AI Chip Favorite Is Overpriced, and how its mobile background and power-efficient focus positioned it well for increased AI adoption.
Arm offers the most popular CPU architecture in the world with more than 350 billion chips shipped since its inception, with the company essentially building a monopoly in the mobile CPU market with 99% market share stemming from its ‘heterogenous compute’ design and RISC architecture.
This design helped facilitate lower power requirements as the architecture allows different CPU parts to work together for improved efficiency, with workloads to be allocated across both high-performance and low-performance CPU cores to lower energy by balancing performance.
Arm is translating this expertise in power efficient chips to the data center, powering both Nvidia’s Grace and Vera CPUs, as well as custom CPUs from Amazon, Google and Microsoft. For example, Google says its Axion CPUs can offer up to 65% better price-performance and 60% better energy efficiency versus x86 alternatives, while Microsoft’s Cobalt 100 and Amazon’s Graviton4 CPU both have shown significant performance and price-performance advantages over competing x86 products.
The Role of CPUs in Agentic AI and the Coming 4X Increase in CPU Cores
Agentic AI represents a natural evolution from the query-and-response based nature of chatbots, to a more complex system capable of running dozens of different tasks and tools autonomously to reason through a problem and provide a response.
As a result, CPUs will play a more integral role in agentic and multi-agent systems to help solve how the system will schedule dozens of concurrent API requests and tool calls across independent agents, as quickly as possible with minimal delay. This is the orchestration constraint: how dozens of agents can make hundreds of concurrent requests needed to complete their independent workflows without causing significant latency or GPU underutilization.
Multi-agent systems are also expected to drive an exponential increase in token generation, which Arm estimated at up to a 15X increase in tokens per user, due to the increase in tool calls and API requests associated with each agent. This is expected to drive CPU core demand much higher, at a time where key x86 suppliers AMD and Intel battle growing supply constraints.
Arm CEO Rene Haas detailed at the company’s latest ‘Arm Everywhere’ event that a typical AI data center of today will feature around 30 million CPU cores per GW. Solving the flow bottlenecks of agentic inference will require “CPUs near the head node, CPUs next to the accelerator rack, more CPU racks inside the data center,” driving CPU core demand as much as 4X higher to 120 million cores per GW, per Haas.
CPU-Centric AI Systems
We are seeing evidence of how agentic AI and the importance of CPUs are beginning to affect incoming architecture designs. If you did not catch last week’s free stock newsletter, Nvidia Stock Prediction: The Path to a $20 Trillion Market Cap is Strengthening, I want to relay the importance of one sentence: “The goal is no longer to simply sell faster and more powerful chips, but to deliver superior economic value at the system level relative to custom silicon (in other words, let the battle begin).”
Think of the Groq LPX racks that cost Nvidia a record-breaking $20 billion that it is soon deploying, addressing the memory-intensive decode phase of inference to significantly accelerate token throughput. This system-level focus is now moving to CPUs, with Nvidia pivoting to deploy its Grace and Vera CPUs as standalone racks. Nvidia says that when “paired with Rubin GPUs as a host CPU, or deployed as a standalone platform for agentic processing, Vera enables higher sustained utilization by removing CPU-side bottlenecks that emerge in training and inferencing environments.”
Arm’s big announcement was not just that it was launching its first in-house silicon and CPU rack platform, but that it will now be directly competing on the standalone CPU system side. Arm’s foray into the standalone rack market offers a new choice for hyperscalers and data center operators to seamlessly deploy CPU-centric racks alongside GPUs, customize the CPU-to-GPU ratio to optimize for agentic orchestration and enable power-efficient AI inference at scale from start to finish. It will also provide another outlet to avoid vendor-lock in to Nvidia’s ecosystem with both air-cooled and liquid-cooled CPU racks.
Key Advantages of Arm’s ‘AGI CPU’ for Agentic AI Workloads
Arm also marked its long-awaited foray into physical chip development with its ‘AGI CPU’, launched at its Arm Everywhere event last week. The company’s pivot into physical CPU and rack development is one the AI industry will watch with great anticipation given Arm’s history of owning significant IP in the mobile space combined with the company setting out to solve agentic AI’s orchestration challenges.
Leveraging Arm’s history of delivering high performance with low power requirements for mobile devices, the new AGI CPU is designed to offer a similar balance between high performance and low power consumption.
The AGI CPU was co-developed with key partner Meta, the chip’s first customer, who revealed they turned to Arm almost two-and-a-half years ago to see if there was a CPU option that fit Meta’s needs: “put in a lot more cores per watt, but we do not want to compromise on the performance piece.” Meta had only been finding options satisfying one of the two criteria: meeting the performance but with too much power, or meeting the power but with too little performance.
2X Performance of x86 and Record CPU Rack Density
The new CPU features up to 136 of Arm’s highest-performing Neoverse V3 cores, drawing just 300W of power in a single-unit (1U) dual-node server (blade) featuring two chips. This compares to x86 chips, such as AMD’s fifth-gen EPYC CPUs, which deliver 128 to 192 cores per chip but at 390W to 500W in a two-unit (2U) rack.
In an air-cooled rack, Arm can pack 30 blades (or 60 CPUs) for a total of 8,160 cores in a 36kW power envelope, saying this configuration can deliver up to 2X the performance per rack versus x86 chips based on its internal estimates. Arm says this 30-blade design is “setting records for air cooled” racks that is not feasible with other systems, as power consumption is too high.

Arm is taking this a step further with a fully-liquid cooled, 200kW open-standard rack in partnership with Super Micro, packing 168 blades, or 336 CPUs, delivering a total of up to 45,696 cores. Arm EVP of Cloud AI Mohamed Awad stated that while it is a “200-kilowatt rack. We actually will consume about half that much power. We ran out of space. That’s why we couldn’t put more cores in there.”
This is one of the key advantages – it is not just about offering 2X the performance of x86 chips, but providing that performance boost while freeing up power for more compute or for more networking:
“So if you have a CPU that can draw less power, it could be just as performant, but use less power, it means you have more leftover for everything else that you want to do. That means more inference and more compute. That means more intelligence.”
One Petabyte of Memory and Low-Latency Chip Design
However, Arm architected the new chip with other key optimizations in mind, notably on the memory side. The AGI CPU features 96 PCI Gen6 lanes with CXL 3.0 connectivity, which Arm says allows the new CPU to be attached to any accelerator of choice, allowing flexibility of deployments with Nvidia or AMD GPUs or custom chips.
The chip also features up to 6TB of memory, providing more than 1 petabyte (PB) of low-latency memory in the liquid-cooled 200kW rack. Arm is achieving this extreme low latency of <100 nanoseconds from memory via a dual chiplet design, with each chiplet having both the memory and I/O directly onboard to avoid multiple links across the silicon.







