NVIDIA Vera Rubin and DSX Boost AI Tokens per Megawatt
NVIDIA unveils Vera Rubin and DSX gains, including 30x throughput per megawatt, 45x lower token costs and flexible grid load controls for AI factories.
Summary
At 9:55 a.m. PT on Sept. 15, 2026, NVIDIA hyperscale and high-performance computing vice president Ian Buck told 8,000-plus AI Infra Summit attendees, versus 3,500 last year, that AI factory economics are shifting from peak performance to agentic tokens per megawatt. NVIDIA’s stack spans Vera Rubin NVL72, Dynamo, NeMo, NVLink, Spectrum-X, ConnectX SuperNICs, and BlueField context-memory storage and security DPUs. Amazon’s Annapurna Labs is developing NVHBM custom memory, d-Matrix is integrating NVLink Fusion with Raptor XPUs, and Pinterest is combining Blackwell and Dynamo for conversational visual discovery.
Lambda’s first Blackwell validation of DSX MaxLPS ran 19 nodes within a 16-node power budget, raising throughput 24%, from roughly 4 million to 5 million tokens per second, and performance per watt 23%. MaxLPS promises up to 1.4x more tokens per megawatt and, with Vera Rubin, 40% more GPUs and 35% more throughput without new power lines. Emerald AI and Silicon Valley Power handled hundreds of demand signals without harming workloads; planned DSX Flex integration in Conductor will pause lower-priority jobs while preserving critical work. Vera Rubin plus Groq 3 LPX claims 35x GB200 NVL72 throughput per megawatt for long-context models above 2 trillion parameters; Groq reached 2,529 output tokens per second per user on 100K-context Qwen 3.8 27B.
SemiAnalysis AgentX found agentic sessions use about 15x a simple chat’s token volume and measured Vera Rubin at up to 30x GB300 NVL72 throughput per megawatt and 45x lower cost per million tokens on DeepSeek V4 Pro. Its codesign includes NVLink 6, NVFP4 fifth-generation Tensor Cores, TensorRT LLM and Dynamo; production is scaling. Vera CPU tests found Perplexity SPACE starts 1.9x faster, DeepInfra orchestration 2.2x faster, Redpanda latency 5.5x lower and throughput 73% higher, Starburst queries 3x faster, and Kinetica analytics 2.7x faster. Daytona, ClickHouse and Prime Intellect also reported gains. NVLink 6 adds error correction, retry, recovery, flow control, dynamic routing and link rebalancing for hundred-thousand-GPU systems.
Positives
- Lambda increased Blackwell cluster token throughput 24% and performance per watt 23% while fitting 19 nodes into a 16-node power budget.
- Vera Rubin NVL72 can support up to 40% more GPUs and 35% higher throughput within the same site-power envelope.
- SemiAnalysis AgentX measured up to 30x higher throughput per megawatt and 45x lower token costs on DeepSeek V4 Pro.
- Emerald AI responded to hundreds of Silicon Valley Power demand signals while protecting AI workload performance.
- Redpanda measured 5.5x lower latency and 73% higher throughput with the NVIDIA Vera CPU.
Risks & concerns
- Power remains the defining constraint for AI factories, making new capacity dependent on extracting more work from each megawatt.
- Agentic sessions can consume roughly 15x the token volume of simple chat requests as context, tool calls and sub-agents accumulate.
- DSX MaxLPS can deliver 40% more GPU capacity only in suitable deployment environments.
- DSX Flex preserves critical workloads during grid events by temporarily pausing lower-priority jobs.
- Hundred-thousand-GPU systems face inevitable transient errors, signal degradation and hardware failures that NVLink 6 must contain.