StorageReview.com

NVIDIA Groq 3 LPX Enters Full Production: 3,400 Tokens per Second at 100K Context, 256 LP30s per Rack

AI  ◇  Enterprise

NVIDIA announced the production release of the NVIDIA Groq 3 LPX, a dedicated interactive inference accelerator built to complement the Vera Rubin NVL72 platform. As enterprise AI workloads shift from raw model pre-training toward real-time reasoning and autonomous agent orchestration, infrastructure requirements are changing. Agentic workflows require iterative, multi-step execution where latency during token generation can compound across automated reasoning paths, external tool queries, and multi-agent communications.

NVIDIA Groq 3 LPX compute tray diagram: a 1U liquid-cooled tray with eight Groq 3 LPUs, host CPU, BlueField-4 DPU or ConnectX-9 NIC, and 400Gb/s Ethernet

The Groq 3 LPX architecture specifically excels at this decode phase. Each 2U liquid-cooled compute tray carries sixteen Groq 3 LPUs alongside a host CPU, fabric expansion logic and DRAM, and either a BlueField-4 DPU or a ConnectX-9 NIC, with 32 LPU chip-to-chip optical links and 400Gb/s Ethernet on the front. In a heterogeneous deployment alongside Vera Rubin NVL72 systems, NVIDIA Rubin GPUs handle context ingestion and prefill calculations across massive input datasets, while the Groq 3 LPX units process latency-sensitive token output generation.

Architecture and Benchmark Performance

In standard Artificial Analysis performance evaluations running the open-source Gemma 4 31B model with an active 100,000-token context window, the Groq 3 LPX reached an output rate of 3,400 tokens per second. NVIDIA notes that this throughput level is 4x faster than the nearest alternative platform on identical long-context parameters, and Artificial Analysis records it as the fastest result measured for the model. The comparison is worth reading carefully: NVIDIA’s figure came from a private pre-release endpoint measured on August 21, while the competing numbers are live serverless production endpoints. Maintaining extreme generation velocity across 100k-token windows is essential for enterprise agent architectures that continuously parse extensive code repositories, complex API schemas, and dense technical documentation.

NVIDIA Groq 3 LPU package render showing the die used in the Groq 3 LPX inference accelerator Artificial Analysis chart showing the NVIDIA Groq 3 LPX at 3,401 output tokens per second on Gemma 4 31B at 100K context, about 4x the nearest platform

From a physical system architecture perspective, a rack-scale Groq 3 LPX deployment incorporates up to 256 LP30 accelerators linked through high-speed, direct chip-to-chip interconnects.

Cloud Deployments and Systems Ecosystem

Early commercial adoption is underway across specialized cloud service providers. Nebius has been named the initial AI cloud partner to deploy Groq 3 LPX, with plans to bring the hardware into its Nebius Token Factory production inference service. This architecture is designed to allow developers to access real-time inference acceleration via existing API interfaces without migrating codebases to new framework abstractions. Inference provider Groq expects to be among the platform’s earliest adopters.

NVIDIA chart claiming up to 30x token throughput per megawatt for Vera Rubin NVL72 versus GB300 NVL72 on a 140K-context agentic coding workload NVIDIA chart plotting tokens per second per megawatt against user interactivity, showing Groq 3 LPX extending Vera Rubin NVL72 into the lowest latency region

NVIDIA pairs the launch with updated platform math for the rack it plugs into, claiming up to 30x the token throughput per megawatt for Vera Rubin NVL72 against GB300 NVL72 on a 140,000-token agentic coding workload, with the advantage widening as per-user interactivity targets rise. The hardware implementation fits within NVIDIA’s full-stack Rubin AI factory infrastructure, which spans seven discrete chips across five purpose-built rack configurations.

NVIDIA Vera Rubin rack family render showing the five purpose-built rack configurations in the AI factory platform

System elements operating alongside Rubin and LPX compute clusters include NVIDIA Vera CPU racks, Vera BlueField-4 STX storage arrays, and NVIDIA Spectrum-6 SPX networking fabrics.

What NVIDIA and Nebius Say

“Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” said Jensen Huang, founder and CEO of NVIDIA. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”

Nebius framed the value in terms of where latency actually lands for users. “Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate,” said Danila Shtan, chief technology officer of Nebius. “As the first AI cloud bringing it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant, through the same API developers are already using, with no migration to a new stack.”

Spectrum-X Multiplane and Scale-In Infrastructure

To support the massive networking demands of these inference factories, NVIDIA highlighted its Spectrum-X Multiplane Ethernet architecture. Traditional cluster expansion beyond modern density limits often requires adding a third switching tier, which introduces latency penalties, optical cabling costs, and power overhead. Spectrum-X Multiplane circumvents this by splitting host network interfaces into multiple independent planes, each running a flat, two-tier network that scales to 512,000 GPUs.

NVIDIA Spectrum-X Multiplane topology slide showing multi-rail SuperNIC planes scaling to 512,000 GPUs with 11x faster recovery

Traffic management and failure mitigation are handled in silicon by the NVIDIA ConnectX-9 SuperNIC, which supports network speeds up to 1,600 Gb/s per GPU when paired with 102.4 Tb/s Spectrum-6 Ethernet ASICs. Spectrum-XGS Ethernet extends the same codesign between facilities, which NVIDIA says accelerates multi-site NCCL collectives by 1.9x. In an eight-plane topology, hardware failover reroutes traffic instantly around failed paths, maintaining roughly 90 percent of aggregate network bandwidth. NVIDIA puts hardware recovery at 11x faster than software-based multiplane load balancing, which it says translates to 1.6x higher AI factory output.

NVIDIA Scale-In networking slide showing DOCA services, BlueField-4, and Spectrum-X Ethernet for AI factory security and data acceleration

In parallel, NVIDIA introduced its Scale-In networking pillar, combining BlueField-4 data processing units with the DOCA software stack. Scale-In targets the north-south infrastructure fabric, offloading multi-tenant isolation, line-rate in-silicon encryption, telemetry, and high-performance storage data paths from host processors.

NVLink Fusion for Custom Silicon Integration

NVIDIA also introduced NVLink Fusion for hyperscale operators requiring hybrid architectures. The framework allows third-party custom XPUs and CPUs to integrate directly with NVIDIA’s sixth-generation NVLink and NVLink-C2C scale-up networking fabrics. By leveraging modular NVIDIA MGX chassis blueprints and management software, data center operators can standardize on shared power, thermal, and networking rack designs while adjusting the specific ratio of GPUs, LPUs, and custom silicon as workloads evolve.

NVIDIA made a second Vera Rubin announcement at Hot Chips: SpaceXAI is adopting Vera CPUs for the agentic workloads behind Grok, and plans to put an optimized Vera Rubin NVL72 in orbit aboard its first-generation Starmind satellite.

Engage with StorageReview

Newsletter | YouTube | Podcast iTunes/Spotify | Instagram | Twitter | TikTok | RSS Feed

Harold Fritts

I have been in the tech industry since IBM created Selectric. My background, though, is writing. So I decided to get out of the pre-sales biz and return to my roots, doing a bit of writing but still being involved in technology.