VentureBeat
Follow
Nvidia's Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests
Enterprises face a significant tradeoff with always-on AI agents: using frontier models is expensive, while custom routing logic requires extensive engineering and maintenance. Nvidia proposes a solution by releasing Nemotron 3.5 Lightning, an open mixture-of-experts model for specialized tasks, and NeMo Switchyard, an open-source library for intelligent model routing. Lightning offers faster output compared to similar models, and when combined with Switchyard, it promises frontier-level task completion at a fraction of the cost of premium models. This release enters a competitive market with numerous open-weight models emerging from China and Meta. Nvidia’s strategy emphasizes an integrated system of both an optimized model and a smart router, rather than individual components. Switchyard competes with existing open-source routers by offering dynamic routing strategies that adapt to an agent's changing state and cost considerations. The library integrates with popular agent frameworks and LLM gateways, facilitating its adoption. Early testers have reported substantial cost reductions and maintained accuracy using Switchyard. Nemotron 3.5 Lightning itself is designed for high-volume, specialized agent tasks, extending Nvidia's hybrid Mamba-Transformer architecture. While not a general intelligence leader, Lightning demonstrates strong performance in speed-to-accuracy benchmarks. For enterprises, this signifies a shift towards dynamic, per-step routing, and the strategic advantage of open source at both model and routing layers. The competitive landscape is evolving from solely focusing on the best model to optimizing the entire system of model-to-task matching.