Nvidia Rubin GPU Targets 10x Agentic AI Throughput Gains
New architecture combines HBM4 memory, enhanced Tensor Cores, and rack-scale power management to handle multi-step reasoning workloads more efficiently than Blackwell.

Nvidia has detailed the architecture behind its Rubin GPU, positioning the chip as purpose-built for agentic AI workloads that demand sustained inference across multiple reasoning steps rather than single prompt-response exchanges.
The company claims Rubin delivers up to 10x more agentic throughput per unit of energy compared to its Blackwell architecture. That efficiency gain stems from coordinated improvements across compute, memory, and system design rather than raw performance increases alone.
Core architectural changes
Rubin packs 336 billion transistors across two reticle-limited compute dies unified through Nvidia's High-Bandwidth Interface. The GPU features 224 streaming multiprocessors and 896 Tensor Cores, with a third-generation Transformer Engine that delivers up to 50 petaflops of NVFP4 inference performance.
The memory subsystem represents a substantial upgrade. Rubin integrates up to 288 GB of HBM4 memory delivering up to 22 TB/s peak bandwidth—a 2.8x increase over Blackwell. The expanded capacity and bandwidth address the decode phase bottleneck in agentic workloads, where models must rapidly move weights and key-value cache state to generate tokens sequentially.
Nvidia enhanced the Tensor Memory Accelerator to reduce data movement overhead in mixture-of-experts models, which dynamically route tokens across many expert networks. The updated design supports inline descriptor updates, allowing kernels to modify memory pointers and strides directly in instructions rather than rewriting descriptors in memory.
Optimizations for long-context attention
Rubin tackles attention performance through activation sparsity and improved softmax throughput. The architecture can compress intermediate attention scores into a structured 2:4 sparse format, reducing both compute and data movement in long-context scenarios without changing the model's input-output interface.
Exponential math throughput increased 2x for FP32 and 4x for BF16/FP16 operations compared to Blackwell, helping softmax operations keep pace with faster matrix computations as context windows expand.
The GPU also enables finer-grained coordination between dependent kernels. Consumer kernels can begin work as soon as required input data becomes available from producer kernels, rather than waiting for broader dependencies to resolve. This tighter kernel-to-kernel execution reduces idle gaps in the GPU timeline.
Rack-scale efficiency engineering
Nvidia positions the Vera Rubin NVL72 rack system as an integrated execution domain rather than a collection of individual GPUs. The rack architecture incorporates state-of-charge Intelligent Power Smoothing, which the company says reduces average power consumption by approximately 10 percent and peak power by roughly 20 percent compared to previous power-smoothing techniques.
At the data center level, Nvidia's DSX MaxLPS software extends power optimization across GPUs, racks, and workloads. The company claims operators can provision up to 40 percent more GPUs within the same power budget at energy-efficient operating points, with minimal performance impact.
The third-generation MGX rack design combines cable-free compute and switch trays, 45°C liquid cooling, dynamic power steering, and hot-swappable NVLink switch trays to support resilient operation at scale.
Why it matters
Agentic AI systems that reason, plan, and execute multi-step tasks represent a fundamentally different workload profile than training or simple inference. These applications spend more time in memory-bound decode phases, require sustained low-latency execution across many reasoning steps, and must scale across tightly coupled GPU clusters. Rubin's architectural choices—particularly the HBM4 memory subsystem and rack-level power management—directly address these operational constraints. For enterprises deploying agentic AI at scale, the efficiency gains translate to meaningful reductions in infrastructure cost and energy consumption per unit of useful output.
These architectural details were first reported by Nvidia in a technical blog post on its developer site.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
