Nvidia Pushes CPU Strategy as AI Agents Drive Data Center Shift
The GPU giant is positioning its Vera Rubin system as a complete solution for agentic AI workloads, challenging AMD and Intel in the CPU market.

Nvidia's Expanding Ambitions Beyond GPUs
Nvidia is making an aggressive push into the CPU market as artificial intelligence workloads evolve beyond traditional model training. During a technical workshop at its Santa Clara headquarters last week, company executives unveiled new performance benchmarks for the Vera Rubin chip system, positioning the hardware as a complete solution for next-generation AI data centers rather than just a GPU supplier.
The timing is strategic. Nvidia held its briefings just days before rival AMD's annual product event in San Francisco, where the competitor is expected to showcase its own data center offerings, including the Helios AI chip rack designed to compete directly with Vera Rubin.
Why it matters
The industry's shift toward agentic AI systems—autonomous software that can plan, reason, and execute complex tasks—requires different hardware capabilities than traditional large language models. While GPUs remain essential for training and inference, CPUs are increasingly critical for orchestrating data flows, networking, and coordinating multiple AI agents. Nvidia's emphasis on CPU performance signals that the company recognizes this architectural shift and is positioning itself to capture revenue across the entire data center stack, not just accelerator chips.
Vera Rubin's Technical Approach
The Vera Rubin system pairs one CPU with every two GPUs. A complete NVL72 configuration includes 36 Vera CPUs alongside 72 Rubin GPUs, all housed in a single liquid-cooled rack. According to Ian Buck, Nvidia's vice president of accelerated computing who led the briefings, the company claims the system will process ten times as many tokens per watt compared to the previous-generation Grace Blackwell architecture.
Nvidia is also selling the Vera CPU as a standalone product, with reports indicating Chinese customers could receive shipments as early as August.
The chip design represents a departure from industry trends. While AMD and others have embraced chiplet architectures—stitching together multiple smaller chips—Nvidia built Vera on a monolithic design. Hannah Coutand, who runs product marketing for Nvidia DGX Cloud, argued that chiplets impose "a heavy tax on memory bandwidth and data movement," whereas the single integrated circuit allows faster data transfer.
Installation and Efficiency Claims
Nvidia executives emphasized practical deployment advantages. The company has significantly reduced cabling requirements in multi-rack configurations, marketing Vera Rubin as "cable-free compute" with hot-swappable components. Andrew Bell, senior vice president of hardware engineering, said this could reduce rack installation time from hours to minutes.
The system's 100 percent liquid-cooling approach also addresses energy concerns, as air-cooling requires more power. Localized memory subsystems offer nearly three times the memory bandwidth of Blackwell, addressing ongoing high-bandwidth memory shortages that have constrained AI development.
Production Timeline and Early Customers
CEO Jensen Huang has repeatedly stated that Vera Rubin will reach full production in the second half of 2025. During a brief tour of an Nvidia data center lab, executives revealed that OpenAI already has one Vera Rubin rack in operation. Microsoft and Oracle are also listed as early customers.
The company remains sensitive to delivery concerns after its Blackwell chips reportedly experienced overheating issues in custom server racks, forcing design changes and shipment delays.
These details were first reported by WIRED, which attended the technical workshop in Santa Clara.
This is an original analysis by the Omega editorial team. Source reporting: WIRED.
Want systems like this working for your business?
Book a Call
