NVIDIA Vera Rubin NVL72 Hits 3.7x GB300 Throughput in MLPerf v6.1 Debut
NVIDIA’s Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x GB300 throughput, while four GB300 racks achieve 99% scaling efficiency in tests.
Summary
MLCommons released MLPerf Inference v6.1 Closed Division results on September 16, 2026. NVIDIA's first Vera Rubin NVL72 preview, entries 6.1-0106 and 6.1-0074, delivered up to 3.7x GB300 NVL72 throughput on Qwen3-VL across offline, server and interactive tests using vLLM and open source NVIDIA Dynamo, and up to 2.5x on DeepSeek-R1 using TensorRT-LLM. Enhanced Tensor Cores, Transformer Engine and NVFP4 accelerate prefill and decode while reducing model-weight, attention and KV-cache memory with minimal claimed output-quality loss. Disaggregated serving, expert parallelism, sixth-generation NVLink and NVLink Switch support rack-scale execution, with 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet. Nebius also submitted Vera previews. SemiAnalysis AgentX preview testing showed 30x GB300 performance, while the forthcoming MLPerf Endpoints benchmark will standardize agentic inference.
A DeepSeek-R1 submission, entries 6.1-0073 and 6.1-0074, scaled GB300 NVL72 from one 72-GPU rack to four racks and 288 GPUs with 99% offline efficiency. GB300 produced 0.65 720p WAN 2.2 videos per second at 5.7 seconds each, delivering 9x single-node throughput and 7.5x lower latency. Qwen3-VL performance rose up to 1.6x from v6.0 through software changes; later GPT-OSS-120B and DLRMv3 gains remain unverified by MLCommons. Jetson AGX Thor ran Qwen3.6-27B on the new Edge-Agentic benchmark with TensorRT Edge-LLM. Nineteen partners participated, eight on multi-node Blackwell NVL72: ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro and Wiwynn.
Positives
- Vera Rubin NVL72 achieved up to 3.7x GB300 NVL72 throughput on Qwen3-VL and up to 2.5x on DeepSeek-R1.
- Four GB300 NVL72 racks containing 288 GPUs retained 99% scaling efficiency on the offline DeepSeek-R1 test.
- GB300 NVL72 delivered 9x single-node WAN 2.2 throughput while reducing latency 7.5x.
- Software improvements raised GB300 NVL72 Qwen3-VL performance by up to 1.6x over MLPerf Inference v6.0.
- Nineteen NVIDIA partners participated, including eight running multi-node Blackwell NVL72 systems.
Risks & concerns
- Vera Rubin's MLPerf submissions and its 30x SemiAnalysis AgentX result remain preview measurements rather than final production evidence.
- Post-submission GPT-OSS-120B and DLRMv3 performance gains have not been verified by MLCommons.
- NVFP4 output-quality loss is described only as minimal, without a quantified quality measurement.
- No power, acquisition-cost or operating-cost figures substantiate NVIDIA's claims about lower token costs and higher revenue.
- MLPerf Endpoints is still forthcoming, leaving agentic inference without the cited standardized measurement framework.