On June 1, in a keynote speech at CompuTEX 2026, NVIDIA CEO Huang Renxun officially announced the full production of the Vera Rubin platform. This next-generation AI super chip, positioned by Nvidia as "the largest POD-level platform to date", marks a key node in the transition of AI computing power infrastructure from "kilocal-level" to "ten-thousand-level", and also indicates that the computing power competition in the age of Agents has entered a white-hot stage.

1. Core performance: tenfold cost reduction and threefold speed increase
Vera Rubin's most eye-catching performance promise is a reduction in single-token inference costs to one-tenth of Blackwell's, while simultaneously increasing throughput for large-scale agents by a factor of 10. This combination of "reducing costs and increasing efficiency" directly hits the biggest pain point of current AI deployments-reasoning costs.
flagship configuration The hard-core parameters of Vera Rubin NVL72 include:
72 Rubin GPUs + 36 Vera CPUs, realizing all-to-all non-blocking interconnection through NVLink 6
Total video memory is 20.7 TB HBM4, bandwidth reaches 28.8 TB/s
NVFP4 inference performance 3,600 PFLOPS, FP8 training performance 2,520 PFLOPS
NVLink's total throughput is 260 TB/s, double that of the previous generation
10 times higher inference throughput per watt than Blackwell
A single Rubin GPU can provide 50 PFLOPS of NVFP4 inference power, which is 5 times that of Blackwell; it is equipped with 288 GB HBM4 memory and a bandwidth of 13 TB/s, which is more than 60% higher than Blackwell's HBM3E. The Vera CPU is equipped with 88 custom Olympus cores (ARM v9.2-A architecture), supports 176 threads and spatial multithreading technology, and has a memory bandwidth of 1.2 TB/s, which doubles the performance of the previous generation Grace CPU.
2. POD-level platform: AI supercomputer composed of five racks
Vera Rubin is not a single chip or server, but Nvidia's largest POD-level platform to date-a huge AI supercomputer made up of five dedicated racks, designed for intelligent workloads.
rack assembly | functional orientation |
Vera Rubin NVL72 | GPU computing core, equipped with 72 Rubin GPUs |
Vera CPU | Agent orchestration and data preprocessing, equipped with 88-core Olympus CPU |
Groq 3 LPX | Dedicated inference acceleration to optimize agent response speed |
Vera BlueField-4 STX | Storage and data offload, supporting contextual memory sharing |
Spectrum-6 SPX | Ethernet network interconnection, using CPO Silicon Photonics Technology |

These five racks are deeply coupled through Nvidia's self-developed high-speed interconnection technology and integrated into a fully integrated system. Among them,A consistent bandwidth of up to 1.8TB/s is achieved between the Vera CPU and the Rubin GPU through the second-generation NVLink-C2C. The CPU can directly access the GPU memory, eliminating the data handling bottleneck of traditional PCIe. The new "contextual memory storage" technology in the BlueField-4 STX storage rack allows multiple GPUs to share KV Cache at high speed, improving inference speed and energy efficiency by 5 times.
thisThe "five-rack in one" POD-level design makes Vera Rubin no longer a simple stack of scattered components, but a complete AI factory unit deeply optimized for the concurrent workload of the agent-from computing, storage to network, full stack collaboration, plug and play.
3. Cooling revolution: 100% liquid cooling, fanless design
supported For the stable operation of the 2.3kW ultra-high power consumption chip, Vera Rubin adopts the third-generation completely cable-less design and 100% fully liquid-cooled coverage solution. The flagship NVL72 rack achieves cable-free and fan-free modular design, building a complete cooling ecosystem of "precise temperature control + efficient heat exchange."
Analysts at Morgan Stanley pointed out thatThe cooling upgrade of the Rubin platform will drive an explosion in demand for liquid-cooled core components. The total value of cooling components per cabinet for the next generation Vera Rubin NVL144 platform is expected to reach US$55,700, a 17% increase compared to the previous generation GB300 platform. Liquid cooling suppliers with precision processing capabilities and full-stack solution capabilities will usher in significant market opportunities.
4. Mass production rhythm: The size of the supply chain has doubled, and shipments will begin in autumn
Huang Renxun is Computex emphasized that Vera Rubin's supply chain is twice the size of Grace Blackwell. With proven open source MGX design, hundreds of partners in the NVIDIA supply chain ecosystem are accelerating production in more than 350 factories in more than 30 countries.
The production timeline is clear:
June 2026: Fully put into production and start trial production
July 2026: Start supplying North American CSPs (Microsoft, Google, Amazon, Meta, Oracle)
Autumn 2026: The first batch of products will be officially shipped
2026 Q4: Production capacity will increase and large-scale shipments will be carried out
TSMC adopted it earlier this year Vera Rubin chips are mass-produced in 3nm process. Major AI server manufacturers such as Hon Hai, Quanta, Wistron and Supermicro have received actual chip samples in preparation for mass production. Among them, Hon Hai has started the development of Vera Rubin NVL144 MGX server in advance.
It is worth noting that Nvidia's Vera Rubin adopts a stricter supply chain management strategy: the system is first assembled to the L10 level by Wistron, Quanta and Hon Hai, and then Nvidia completes the integration and delivery of racks in a unified manner to ensure consistent performance of each rack. This "whole-cabinet delivery" model marks the transformation of AI infrastructure from a business model of "selling chips" to "selling AI factories."
5. Customer Ecosystem: Cloud giants are betting on the entire board, and AI laboratories are the first to take the lead in blocking.
Vera Rubin's customer lineup can be called an "All-Star":
Cloud service providers: Microsoft (Fairwater AI Super Factory will deploy hundreds of thousands of units), AWS, Google Cloud, Oracle Cloud
AI laboratories: OpenAI, Anthropic (received the first batch of samples)
Commercial aerospace: SpaceX AI (Musk personally signed for it for enhanced learning and agent simulation)
System manufacturers: Dell, HPE, Lenovo, Ultramicro
Microsoft has taken the lead in launching Vera Rubin NVL72 engineering machine cabinet became the first cloud service provider to deploy on a large scale. Oracle promises to deploy hundreds of thousands of Vera CPUs starting in 2026. This full coverage pattern of "cloud factory +AI factory + system factory" ensures that Vera Rubin has mature ecological implementation capabilities from the beginning of its release.

6. Strategic significance: from "computing power supplier" to "AI factory architect"
The full production of Vera Rubin is not only an iteration of Nvidia's product line, but also an upgrade of its strategic positioning:
Business model transition: From selling single GPUs/CPUs to selling "rack-level AI factories." Vera Rubin NVL72 is delivered as a complete computing unit, and customers buy "plug and play" AI productivity rather than discrete components.
Technical route locking: Through NVLink-C2C to realize the CPU-GPU unified memory architecture, and through CPO silicon photonics technology to reconstruct the network layer, NVIDIA is building a full-stack closed loop from chip to cabinet, from computing to network. Even if competitors break through at a single point, it will be difficult for them to shake their system-level advantages.
Occupancy in the agent era: Vera Rubin is specially designed for Agentic AI. Its CPU is responsible for model scheduling, memory management and tool call orchestration, the GPU is responsible for parallel computing, and the DPU is responsible for data offloading-this "clear division of labor, collaborative and efficient" architecture, It is the infrastructure needed for large-scale concurrent operation of future AI agents.
7. Future roadmap: Rubin Ultra and Feynman are on the road
Nvidia'sThe pace of "generation of year" is still accelerating:
Second half of 2026: Vera Rubin NVL144 platform, performance is 3.3 times higher than GB300 NVL72
Second half of 2027: Rubin Ultra NVL576 platform, performance improved by 14 times
2028: The next generation architecture Feynman (named after physicist Richard Feynman) is unveiled
When Huang Renxun was When Vera Rubin was fully put into production on the stage of COMPUTER 2026, his core message was not only "We have built faster chips," but "We have redefined the standard form of AI factories." In the era of agents, the competition for computing power has shifted from "single-card computing power" to "system efficiency" and from "training speed" to "reasoning cost." Vera Rubin is pushing the economic feasibility of AI deployment to new heights with a combination of "six-core collaboration, ten-fold cost reduction, and full-stack liquid cooling." As autumn shipments approach, the infrastructure landscape of the global AI industry may undergo another profound restructuring.