AI

Nvidia ships Groq 3 LPX inference chip to accelerate AI agents

The new accelerator, built with licensed Groq technology, delivers 3,400 tokens per second and pairs with Vera Rubin GPUs in rack-scale deployments.

Omega Editorial· August 24, 2026· 3 min read

Nvidia launches dedicated inference accelerator

Nvidia has begun full production of the Groq 3 LPX, a specialized inference accelerator designed to eliminate latency bottlenecks in AI agent workloads. The chip was unveiled at Hot Chips 2026 and represents Nvidia's first product built with technology licensed from Groq Inc., the inference-focused chipmaker Nvidia acquired access to for $20 billion in December 2025.

The Groq 3 LPX is positioned as a purpose-built extension to Nvidia's Vera Rubin data center platform. While Vera Rubin GPUs handle large-scale context processing, the new LPX accelerators offload token generation — the step that determines how quickly an AI system responds to users. According to benchmark data from Artificial Analysis, the chip outputs 3,400 tokens per second when running the open-source Gemma 4 31B model with a 100,000-token context window, making it four times more responsive than competing platforms for latency-sensitive tasks.

Neocloud provider Nebius Group has signed on as the first customer, planning to integrate Groq 3 LPX into its Nebius Token Factory inference platform. A full Vera Rubin NVL72 rack can now incorporate up to 256 LP30 accelerators linked by Nvidia's high-bandwidth interconnects, creating what the company describes as a unified inference engine for enterprise-scale AI agents.

Why it matters

AI agents represent a shift from single-query models to systems that reason, plan, execute code, and call external tools in continuous loops. These workflows can generate thousands of tokens across complex chains, and decode latency — the delay between processing context and generating output — becomes a critical performance barrier. By disaggregating context ingestion from token generation, Nvidia is addressing a fundamental architectural challenge as enterprises move from experimental chatbots to production agents that must feel instantaneous. The move also signals Nvidia's strategy to maintain dominance across the full AI compute stack, not just training.

Groq acquisition pays off quickly

Nvidia's December deal with Groq Inc. included licensing the startup's inference-focused processor technology and hiring founder Jonathan Ross and President Sunny Madra. The Groq 3 LPX marks the first commercial product to emerge from that arrangement. Nvidia CEO Jensen Huang said the chip "transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness."

Nebius CTO Danila Shtan emphasized the chip's focus on the generation phase: "Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what Groq 3 LPX is built to accelerate."

Broader platform announcements

Nvidia also announced SpaceX as a flagship customer for the Vera Rubin platform, with plans to deploy Vera CPUs for orchestration and simulation tasks spanning terrestrial data centers and orbital satellites. Additional infrastructure technologies unveiled include Spectrum-X Multiplane, an Ethernet architecture for scaling clusters up to 512,000 GPUs; Nvidia Scale-In, software that offloads security and networking from compute nodes; and NVLink Fusion, which connects custom CPUs and data processing units to sixth-generation NVLink systems.

These details were first reported by SiliconANGLE.

#nvidia#ai inference#groq#ai agents#data center hardware#vera rubin

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in AI

AI· 3 min read

SLAC's neural network compresses scientific data 100x without loss

New AI method preserves fine-grained details like X-ray speckles that traditional compression erases, addressing storage crisis at next-gen facilities.

Via AI Watch · Aug 24, 2026
AI· 3 min read

LinkedIn Reports 40% Drop in AI Slop Post Reach

The platform's crowdsourced feedback system has logged over one million reports since launching in late July.

Via AI Watch · Aug 24, 2026
AI· 3 min read

Intel Unveils Diamond Rapids, Crescent Island, Wildcat Lake for Agentic AI

Three new architectures span datacenter orchestration, inference acceleration, and edge computing using Intel 18A process and UCIe interconnects.

Via AI Watch · Aug 24, 2026