NVIDIA Vera Rubin NVL72 Claims 30x More Agentic AI Throughput Per Megawatt
NVIDIA says Vera Rubin NVL72 delivers up to 30x GB300 throughput per megawatt and 35x lower token costs for power-hungry agentic AI workloads at scale.
Summary
On August 24, 2026, NVIDIA said Vera Rubin NVL72 delivered up to 30x more agentic workload throughput per megawatt than GB300 NVL72 and up to 35x lower cost per million tokens. OpenRouter data indicates agents consume 15x more tokens than chat requests because tools, sub-agents and repeated reasoning expand context from the 1K to 8K tokens typical of chat or summarization into hundreds of thousands. NVIDIA used SemiAnalysis AgentX recordings of real coding sessions, preserving context growth, tool calls and sub-agent spawning. Blackwell led Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro, with GB300 reaching up to 15x Hopper throughput per megawatt on DeepSeek V4 Pro. Vera Rubin’s results remain pending SemiAnalysis review and exclude Vera CPU tool-calling performance. DSX MaxLPS can provision up to 40% more GPUs within one megawatt.
Vera Rubin combines disaggregated prefill and decode serving, rate matching, expert parallelism, distributed KV caching and offloading, KV-aware routing and fused MegaMoE CUDA kernels. Rubin GPUs add enhanced fifth-generation Tensor Cores, a third-generation Transformer Engine and NVFP4 4-bit quantization. The NVL72 scale-up domain, also used by Grace Blackwell, employs sixth-generation NVLink and NVLink Switches with 10x Ethernet packet rates and 3x lower latency. NVIDIA TensorRT LLM and Dynamo complete the codesigned software stack. The seven-chip platform comprises Rubin GPU, Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC. Vera Rubin is in full production and scaling through partners, while software optimization continues for Vera Rubin and GB300.
Positives
- Vera Rubin NVL72 delivered up to 30x GB300 NVL72 throughput per megawatt on DeepSeek V4 Pro in NVIDIA’s early AgentX measurements.
- Vera Rubin NVL72 reduced cost per million tokens by up to 35x compared with GB300 NVL72.
- DSX MaxLPS can provision up to 40% more GPUs within the same megawatt budget.
- Sixth-generation NVLink provides 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives.
- Vera Rubin is in full production and scaling across NVIDIA’s partner ecosystem.
Risks & concerns
- OpenRouter data indicates agentic workloads consume 15x more tokens than simple chat requests.
- Agentic sessions can accumulate hundreds of thousands of input tokens, increasing memory, computation and power requirements.
- Vera Rubin’s AgentX performance results are still pending SemiAnalysis review.
- Current measurements exclude Vera CPU performance for tool calling, leaving part of the seven-chip platform untested in the published comparison.