Thursday, August 27, 2026
Tech Beat
Aug 24, 2026, 3:00 PMAI Infrastructure

NVIDIA Groq 3 LPX Enters Production, Accelerating Vera Rubin Agentic AI

NVIDIA's Groq 3 LPX enters production, delivering 3,400 tokens per second as Nebius, CoreWeave and SpaceXAI adopt the Vera Rubin agentic AI factory platform.

Listen to this briefingAudio briefing

Summary

NVIDIA put Groq 3 LPX into full production on August 24, 2026, extending Vera Rubin NVL72 with low latency token generation for agentic AI. An Artificial Analysis benchmark using open source Gemma 4 31B reached 3,400 output tokens per second across a 100,000 token context, four times the nearest alternative. Rubin GPUs process context while LPUs accelerate decoding and jointly compute every model layer. One rack can connect 256 LP30 accelerators through direct chip links.

Nebius is the first AI cloud adopting Groq 3 LPX, adding it to Nebius Token Factory for interactive agents, coding systems and real time AI. CoreWeave has deployed Spectrum X Multiplane in production. SpaceXAI plans Vera CPUs for orchestration, tool use, code execution, processing and simulation across terrestrial data centers and orbital satellites.

Spectrum X Multiplane divides connections among parallel two tier networks, scaling to 512,000 GPUs without a third tier. NVIDIA claims 1.6 times better networking performance and AI factory output, about 90% bandwidth after one of eight planes fails, and hardware recovery 11 times faster than software balancing. Spectrum 6 switches deliver 102.4 Tb/s, ConnectX 9 SuperNICs support 1,600 Gb/s per GPU, and Spectrum XGS accelerates multisite NCCL collectives 1.9 times. At Hot Chips in Palo Alto, NVIDIA also introduced BlueField 4 and DOCA powered Scale In infrastructure for networking, storage, security and operations, plus NVLink Fusion, using sixth generation NVLink, NVLink Switch, NVLink C2C and MGX to integrate custom XPUs and CPUs. Further Groq 3 LPX model and performance optimizations are planned, without a disclosed schedule.

Positives

  • Groq 3 LPX produced 3,400 output tokens per second on Gemma 4 31B with a 100,000 token context.
  • Nebius became the first AI cloud adopting Groq 3 LPX through its Token Factory service.
  • Spectrum X Multiplane can scale a flat two tier network to 512,000 GPUs.
  • An eight plane network retains about 90% of bandwidth after one plane fails.
  • SpaceXAI plans to use Vera CPUs in terrestrial data centers and orbital satellites.
  • NVLink Fusion lets operators combine custom XPUs and CPUs with NVIDIA networking and MGX infrastructure.

Risks & concerns

  • The claimed fourfold performance lead is based on one disclosed Gemma 4 31B benchmark.
  • NVIDIA disclosed no Groq 3 LPX pricing or schedule for promised model and performance optimizations.
  • Agentic workflows multiply small decoding delays as agents reason, call tools and communicate with other systems.
  • Conventional third network tiers increase latency, jitter, cabling, optics and power costs at large cluster scales.
Primary sourceNVIDIA Bloghttps://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

AI InfrastructureAug 24

NVIDIA NVLink Fusion Brings Custom XPUs Into Unified AI Factories

AI InfrastructureAug 24

NVIDIA Vera Rubin NVL72 Claims 30x More Agentic AI Throughput Per Megawatt

AI InfrastructureAug 21

Nvidia Takes Cloverleaf Stake to Expand AI Data Center Financing