Nvidia Says NeMo Switchyard Cuts AI Agent Costs to One Third
Nvidia pairs Nemotron 3.5 Lightning with open source NeMo Switchyard, claiming frontier level agent results at one third of Opus 4.8's benchmark cost.
Summary
On August 11, 2026, Nvidia released Nemotron 3.5 Lightning, a 30 billion parameter open mixture of experts model for specialized agents, and NeMo Switchyard, an open source library routing each workflow step by agent state, classifiers, predicted verbosity and cost. Nvidia says Lightning is up to 4 times faster than class peers and 30% faster than Qwen3.6 35B at equal accuracy. With Switchyard, it maintained frontier level completion at about one third of Opus 4.8's benchmark cost. The pairing competes with Not Diamond, powering OpenRouter Auto, and UC Berkeley and LMSYS's RouteLLM as Alibaba, Moonshot, Zhipu, DeepSeek and Meta's 30 billion parameter Muse Glimmer expand open models.
Cognition, LangChain and Nous Research call Switchyard directly; Kong, LiteLLM and OpenRouter embed it, including Kong AI Gateway. Among nine testers, LangChain cut cost 74% across 145 multi turn Deep Agents tasks by routing 7% of calls to a frontier model, with 6% lower accuracy. Ramp matched frontier performance on Ramp SWE Bench, cutting cost 58% and runtime 33%. Cognition's staged router in Devin Desktop neared frontier FrontierCode Main performance at 28% lower mean cost than using one frontier model.
Lightning runs Nvidia's hybrid Mamba Transformer latent mixture of experts architecture, introduced with Nemotron 3 in December 2025. Its 24 on Artificial Analysis's nine evaluation Intelligence Index tied gpt oss 120b, below Nemotron 3 Super, Gemma 4 31B, Claude 4.5 Haiku and Mistral Medium 3.5 at 30. Nvidia's PinchBench data says Lightning matched Qwen3.6 35B accuracy 30% faster and beat Gemma 4 26B accuracy at similar speed. Early access post training covered CrowdStrike, CodeRabbit, Harvey, Trajectory and Lila Sciences against named baselines. CodeRabbit used NeMo Auto's one epoch recipe to build a router agent for $85 in about two hours. Nvidia gave no broader Chinese model comparison, citing openness and customizability.
Positives
- Switchyard and Lightning maintained frontier level task completion at roughly one third of Opus 4.8's benchmark cost in Nvidia's tests.
- LangChain reduced costs 74% across 145 Deep Agents tasks by sending only 7% of calls to a frontier model.
- Ramp matched frontier performance on Ramp SWE Bench while cutting costs 58% and runtime 33%.
- Cognition approached frontier performance on FrontierCode Main while reducing mean cost 28% in Devin Desktop.
- Lightning matched Qwen3.6 35B accuracy about 30% faster and exceeded Gemma 4 26B accuracy at similar completion time.
- CodeRabbit built a working router agent with NeMo Auto's one epoch recipe for $85 in about two hours.
Risks & concerns
- LangChain's 74% cost reduction came with a 6% accuracy tradeoff.
- Lightning scored 24 on Artificial Analysis's Intelligence Index, while Nemotron 3 Super and four named competitors scored 30.
- Nvidia supplied the PinchBench data, and the headline cost reduction also came from Nvidia's benchmark tests.
- Nvidia offered no broad head to head benchmark against Chinese models, instead emphasizing openness and customizability.
- Dynamic routing shifts competition toward system performance, which the article says is harder to benchmark and market.