Nvidia releases Vera processor: a "new CPU species" specially built for AI agents. OpenAI and SpaceX have entered the game first

6.1 In the CompuTEX 2026 keynote speech at the Taipei International Computer Show, NVIDIA CEO Huang Renxun officially released the Vera processor. This is NVIDIA's first CPU specially designed for the AI agent AgenticAI. It is also NVIDIA's next major bet in the data center CPU field after Grace.

On June 1, during a keynote speech at CompuTEX 2026, NVIDIA CEO Huang Renxun officially released the Vera processor-this is NVIDIA's first CPU specially designed for AI agents. It is also Nvidia's second major bet in the data center CPU field after Grace. Huang Renxun positioned it as "the company's next multi-billion dollar business," marking Nvidia's strategic transition from a GPU overlord to a full-stack AI infrastructure provider.


 

2026-06-01_173016_420


1. Product positioning: from "GPU supporting role" to "agent leading role"

 

with the previous generation Grace CPU is mainly used as a GPU, while Vera is an independently positioned dedicated processor for the AI agent era. Ian Buck, Vice President of NVIDIA, bluntly stated when delivering the first products: "Proxy AI is creating a new CPU moment in the AI factory-as the model shifts from simply 'answering questions' to proactive 'taking action', Vera is precisely designed to ensure that this workload works efficiently on a large scale. "

 

traditional CPU design is centered on core density and has never made agent load a priority goal. The operating logic of AI agents places unprecedented requirements on the CPU: each agent sandbox environment, every tool call, every orchestration system, and every long-context retrieval operation are all the work of the CPU. Actual measurements by semiconductor organizations show that CPU processing delays account for as much as 50%-90% of agent tasks. If CPU performance is insufficient, no matter how powerful the GPU is, it will be idle waiting for scheduling, forming system-level congestion.


2026-06-01_173303_452


2026-06-01_173316_891

 


2. Core specifications: 88-core Olympus architecture, performance surpasses x86 doubles

 

Vera's core competitiveness lies in its self-developed architecture and ultimate collaborative engineering design:


core indicatorsVera Specificationscomparative superiority
CPU core88 NVIDIA self-developed Olympus cores (Armv9.2-A)22% higher than Grace's 72-core
thread configuration176 threads, supporting spatial multi-threading technologyTwo threads are truly running simultaneously, non-traditional SMT time slicing
memory bandwidth1.2 TB/s(LPDDR5X)About 4 times that of Intel Xeon and 2.4 times that of AMD EPYC
memory capacityUp to 1.5 TBThree times as much as Grace
CPU-GPU interconnectionSecond-generation NVLink-C2C, bandwidth 1.8 TB/s7 times faster than PCIe 6.0
cache configuration2 MB L2 per core, total 164 MB L3Communication between cores is 50% faster than traditional CPUs
power consumption range250W - 450W TDPTwice the energy efficiency of traditional infrastructure


 Spatial Multithreading is a major technical highlight of Vera. Unlike traditional synchronous multi-threading (SMT) time-slice rotation, Vera physically isolates various components of the pipeline, allowing two threads to truly run simultaneously on a single core, greatly improving instruction-level parallelism (ILP) and performance predictability, which is particularly critical for multi-tenant AI factory environments.

 

In In the first benchmark tests announced by Phoronix, Vera performed amazingly: its overall performance (Geomean) was approximately 63% higher than that of its predecessor Grace, defeated the AMD EPYC9575F (64-core Zen 5) by 10%, and crushed the Intel Xeon 6980P (128-core Granite Rapids) by 55%, even surpassing the dual-way configuration. Phoronix calls it "the strongest ARM Linux server processor ever tested."

 


3. Extreme collaborative design: Vera Rubin NVL 72 rack-level supercomputer

 

Vera is not alone, but the core hub of NVIDIA's "Six-Core Collaborative" AI factory blueprint. In the ** Vera Rubin NVL72 ** rack-level system, Vera is deeply integrated with Rubin GPUs, NVLink 6 switches, ConnectX-9 SuperNIC, BlueField-4 DPU and Spectrum-X Ethernet switches to form a complete AI infrastructure unit.

 

Vera Rubin NVL 72 Key Performance Indicators:

  • 72 Rubin GPUs +36 Vera CPUs

  • Inference performance: 3600 PFLOPS (NVFP4)

  • Training performance: 2520 PFLOPS (NVFP 4)

  • Memory configuration: 20.7 TB HBM4, bandwidth 1580 TB/s

  • Rack interconnection bandwidth: 260 TB/s (NVLink 6)

  • Inference throughput per watt is 10 times higher than Blackwell, and cost per token is reduced to 1/10 

 

the adoption of the second generation In NVLink-C2C, Vera and Rubin GPUs share a unified memory architecture to achieve zero-copy data access. This means that the GPU can directly access the CPU's LPDDR5X memory without the need for complex data handling of traditional PCIe, completely eliminating the "data transfer bottleneck."

 


4. The first batch of customers: OpenAI, Anthropic, SpaceX, and Oracle signed for "rushing"

 

Vera's commercialization speed exceeded market expectations. On May 18, Ian Buck personally "drove delivery" and started a Silicon Valley delivery tour:

 

  • Anthropic (SoMa, San Francisco): Received by James Bradbury, head of computing, to expand the agent workload

  • OpenAI (Mission Bay, San Francisco): Buck disassembled the machine on the spot to show the internal structure, received by Sachin Katti, head of OpenAI infrastructure

  • SpaceX AI (Palo Alto): Musk personally signed for it and asked in detail about the number of cores, memory layout and cooling solutions, which are planned to strengthen the learning workload and agent simulation pipeline

  • Oracle Cloud (OCI)(South Bay): Product leader Karan Batta received that Oracle promised ** to deploy hundreds of thousands of Vera CPUs starting in 2026 **, becoming the first cloud service provider to deploy on a large scale 

 

In addition, Dell,System manufacturers such as HPE, Lenovo, and Ultramicro have also integrated Vera into their AI infrastructure product lines. Cloud service providers such as ByteDance, CoreWeave, Lambda, Nebius, and Nscale also plan to adopt it.

 


5. Business prospects: US$20 billion in annual revenue, pointing to the US$200 billion market

 

In At the May earnings conference, Huang Renxun confirmed that Vera CPU is expected to achieve sales of US$20 billion this year, and this figure only counts independent CPU sales and does not include the Superchip portion sold in combination with Blackwell/Rubin. This means Vera will become Nvidia's largest source of new business beyond its US$1 trillion GPU revenue forecast.

 

Analysts pointed out thatThe demand for CPU in the AI agent era is growing exponentially: traditional AI data centers require about 30 million CPU cores per GW of power, while in the agent era, this demand will soar to 120 million cores, a four-fold increase. Both Intel and AMD admit that the ratio of CPU to GPU in data centers will gradually become balanced in the future, or even reverse the ratio.

 

In terms of shipping rhythm,Vera has entered full mass production, with shipments expected to reach 1.2 million in fiscal year 2027 and jump to 4.2 million in fiscal year 2028. Citigroup, Morgan Stanley and other institutions believe that the release of Vera marks Nvidia's shift from "selling chips" to "selling AI factories." Cabinet architecture is replacing single servers as a new system boundary.

 


6. Industry significance: Redefining the "CPU moment" in the AI era

 

The release of Vera is not only an expansion of Nvidia's product line, but also a sign of a paradigm shift in AI computing architecture:

 

  • From "GPU-centrism" to "heterogeneous collaboration": AI workloads are moving from pure model training to complex multi-step reasoning, tool invocations and reinforcement learning, tasks that require the CPU to assume key roles in orchestration, control and data handling. The emergence of Vera proves that the efficiency of an AI factory depends not only on GPU computing power, but also on whether the CPU can "feed" the GPU fast enough. 

  • From "general purpose computing" to "scene-specific": Vera is optimized for agent scenarios. Its 88-core design, 1.2 TB/s memory bandwidth and FP8 native support are all tailored to the concurrent features of AI reasoning and reinforcement learning. This "AI-first" design concept is in sharp contrast to the path of traditional x86 CPUs pursuing versatility. 

  • From "chip competition" to "system competition": Vera Rubin's NVL72's cabinet-level delivery model shows that the future AI infrastructure competition is no longer a competition for computing power on a single chip, but a full-stack collaboration of CPU, GPU, network, storage, and cooling. Nvidia completed the last piece of the puzzle of "six-core collaboration" through Vera, further consolidating the closed-loop advantages of its AI ecosystem.

 


Conclusion

 

When Huang Renxun was When he lifted the Vera processor on the stage of COMPUTER 2026, he showed not only an 88-core ARM chip, but also Nvidia's strategic prediction for AI for the next decade: in the era of agents, the CPU will no longer be a "supporting role" for the GPU, but a "dual engine" that keeps pace with the GPU. With the early entry of top laboratories such as OpenAI, Anthropic, and SpaceX, and Oracle's commitment to deploy hundreds of thousands of tablets, Vera is rapidly moving from "release" to "implementation." For Intel and AMD, a powerful new competitor has emerged; for the entire AI industry, Vera may be ushering in a new era of "CPU redefinition."