NVIDIA announced the production release of the NVIDIA Groq 3 LPX, a dedicated interactive inference accelerator built to complement the Vera Rubin NVL72 platform. As enterprise AI workloads shift from raw model pre-training toward real-time reasoning and autonomous agent orchestration, infrastructure requirements are changing. Agentic workflows require iterative, multi-step execution where latency during token generation can compound across automated reasoning paths, external tool queries, and multi-agent communications.
The Groq 3 LPX architecture specifically excels at this decode phase. Each 2U liquid-cooled compute tray carries sixteen Groq 3 LPUs alongside a host CPU, fabric expansion logic and DRAM, and either a BlueField-4 DPU or a ConnectX-9 NIC, with 32 LPU chip-to-chip optical links and 400Gb/s Ethernet on the front. In a heterogeneous deployment alongside Vera Rubin NVL72 systems, NVIDIA Rubin GPUs handle context ingestion and prefill calculations across massive input datasets, while the Groq 3 LPX units process latency-sensitive token output generation.
Architecture and Benchmark Performance
In standard Artificial Analysis performance evaluations running the open-source Gemma 4 31B model with an active 100,000-token context window, the Groq 3 LPX reached an output rate of 3,400 tokens per second. NVIDIA notes that this throughput level is 4x faster than the nearest alternative platform on identical long-context parameters, and Artificial Analysis records it as the fastest result measured for the model. The comparison is worth reading carefully: NVIDIA’s figure came from a private pre-release endpoint measured on August 21, while the competing numbers are live serverless production endpoints. Maintaining extreme generation velocity across 100k-token windows is essential for enterprise agent architectures that continuously parse extensive code repositories, complex API schemas, and dense technical documentation.
From a physical system architecture perspective, a rack-scale Groq 3 LPX deployment incorporates up to 256 LP30 accelerators linked through high-speed, direct chip-to-chip interconnects.
Cloud Deployments and Systems Ecosystem
Early commercial adoption is underway across specialized cloud service providers. Nebius has been named the initial AI cloud partner to deploy Groq 3 LPX, with plans to bring the hardware into its Nebius Token Factory production inference service. This architecture is designed to allow developers to access real-time inference acceleration via existing API interfaces without migrating codebases to new framework abstractions. Inference provider Groq expects to be among the platform’s earliest adopters.
NVIDIA pairs the launch with updated platform math for the rack it plugs into, claiming up to 30x the token throughput per megawatt for Vera Rubin NVL72 against GB300 NVL72 on a 140,000-token agentic coding workload, with the advantage widening as per-user interactivity targets rise. The hardware implementation fits within NVIDIA’s full-stack Rubin AI factory infrastructure, which spans seven discrete chips across five purpose-built rack configurations.
System elements operating alongside Rubin and LPX compute clusters include NVIDIA Vera CPU racks, Vera BlueField-4 STX storage arrays, and NVIDIA Spectrum-6 SPX networking fabrics.
What NVIDIA and Nebius Say
“Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” said Jensen Huang, founder and CEO of NVIDIA. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”
Nebius framed the value in terms of where latency actually lands for users. “Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate,” said Danila Shtan, chief technology officer of Nebius. “As the first AI cloud bringing it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant, through the same API developers are already using, with no migration to a new stack.”
Spectrum-X Multiplane and Scale-In Infrastructure
To support the massive networking demands of these inference factories, NVIDIA highlighted its Spectrum-X Multiplane Ethernet architecture. Traditional cluster expansion beyond modern density limits often requires adding a third switching tier, which introduces latency penalties, optical cabling costs, and power overhead. Spectrum-X Multiplane circumvents this by splitting host network interfaces into multiple independent planes, each running a flat, two-tier network that scales to 512,000 GPUs.
Traffic management and failure mitigation are handled in silicon by the NVIDIA ConnectX-9 SuperNIC, which supports network speeds up to 1,600 Gb/s per GPU when paired with 102.4 Tb/s Spectrum-6 Ethernet ASICs. Spectrum-XGS Ethernet extends the same codesign between facilities, which NVIDIA says accelerates multi-site NCCL collectives by 1.9x. In an eight-plane topology, hardware failover reroutes traffic instantly around failed paths, maintaining roughly 90 percent of aggregate network bandwidth. NVIDIA puts hardware recovery at 11x faster than software-based multiplane load balancing, which it says translates to 1.6x higher AI factory output.
In parallel, NVIDIA introduced its Scale-In networking pillar, combining BlueField-4 data processing units with the DOCA software stack. Scale-In targets the north-south infrastructure fabric, offloading multi-tenant isolation, line-rate in-silicon encryption, telemetry, and high-performance storage data paths from host processors.
NVLink Fusion for Custom Silicon Integration
NVIDIA also introduced NVLink Fusion for hyperscale operators requiring hybrid architectures. The framework allows third-party custom XPUs and CPUs to integrate directly with NVIDIA’s sixth-generation NVLink and NVLink-C2C scale-up networking fabrics. By leveraging modular NVIDIA MGX chassis blueprints and management software, data center operators can standardize on shared power, thermal, and networking rack designs while adjusting the specific ratio of GPUs, LPUs, and custom silicon as workloads evolve.
NVIDIA made a second Vera Rubin announcement at Hot Chips: SpaceXAI is adopting Vera CPUs for the agentic workloads behind Grok, and plans to put an optimized Vera Rubin NVL72 in orbit aboard its first-generation Starmind satellite.




Amazon