Tuesday, September 15, 2026
Tech Beat
Sep 15, 2026, 4:55 PMAI Infrastructure

NVIDIA DSX Boosts AI Token Output 24% Without More Power

NVIDIA DSX raises Lambda token throughput 24% within a fixed power budget, while Emerald AI automatically cuts Eos demand 25% for Silicon Valley Power.

Listen to this briefingAudio briefing

Summary

At the September 15, 2026 AI Infra Summit, NVIDIA hyperscale and high-performance computing vice president Ian Buck detailed DSX, introduced at GTC Taipei in May. Lambda validated DSX MaxLPS on five racks containing 19 NVIDIA HGX B200 nodes. Running all 19 at 85% power within the budget for 16 full-power nodes raised throughput 24%, from roughly 4 million to 5 million tokens per second, and improved performance per watt 23%. Lambda serves more than 10,000 customers. NVIDIA projects MaxLPS could provide up to 40% more Vera Rubin NVL72 GPU capacity under suitable conditions.

In Santa Clara, Emerald AI Conductor automatically reduced NVIDIA’s Eos AI factory from four megawatts to three in under a minute while high-priority inference continued. Silicon Valley Power has issued more than 200 successful signals through its Flexible Load Interconnect Program, the first commercial utility program treating AI factories as dispatchable resources. The trial informs DSX Flex, with Conductor planned for integration. The first dedicated commercial deployment is a 96-megawatt Vera Rubin factory in Manassas, Virginia, following five demonstrations across two continents.

DSX combines MaxLPS power allocation, Flex grid response, open-source DSX OS, DSX Sim, DSX Exchange and validated reference designs spanning computing, networking, storage and facilities. NVIDIA is adding 800 VDC distribution, projecting 3% to 5% end-to-end efficiency gains over 54V systems with Vera Rubin NVL72 in 2027. Whole-factory optimization also addresses cooling, including roughly 120 kW of heat from each direct-liquid-cooled GB200 NVL72 rack.

Positives

  • Lambda increased cluster token throughput 24%, from roughly 4 million to 5 million tokens per second, without expanding its fixed power budget.
  • Performance per watt improved 23% when 19 HGX B200 nodes ran at 85% power instead of 16 nodes at full power.
  • More than 200 Silicon Valley Power demand signals triggered successful automated responses from NVIDIA’s Eos AI factory.
  • Four megawatts fell to three in under a minute while high-priority inference continued without operator intervention.
  • A 96-megawatt Vera Rubin AI factory in Manassas will become the first dedicated commercial DSX Flex deployment.

Risks & concerns

  • Fixed grid capacity remains the central constraint because a one-gigawatt AI factory cannot become a two-gigawatt facility through software alone.
  • Up to 40% more Vera Rubin NVL72 GPU capacity is an NVIDIA projection limited to suitable deployment environments.
  • The 3% to 5% gain from 800 VDC remains projected, with Vera Rubin NVL72 availability planned for 2027.
  • Each direct-liquid-cooled GB200 NVL72 rack produces roughly 120 kW of heat that facility systems must remove.
  • The stated four-megawatt to three-megawatt event equals a 25% reduction, while NVIDIA’s metrics table separately lists a 40% demand reduction.
Primary sourceNVIDIA Bloghttps://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

AI InfrastructureSep 15

NVIDIA Vera Rubin and DSX Boost AI Tokens per Megawatt

AI InfrastructureSep 14

Cornelis Raises $205 Million to Challenge Nvidia With Open AI Networking

AI InfrastructureSep 7

Inside the $3.2 Billion Lake Mariner AI Data Center’s Accountability Maze