NVIDIA Launches Nemotron 3.5 Lightning and NeMo Switchyard
New 30B-parameter model and routing library aim to optimize multi-agent AI systems across deployment environments.
NVIDIA has released two new tools designed to give enterprises more control over how they deploy and manage AI agent systems: Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, and NeMo Switchyard, an open-source routing library.
The releases reflect a shift in how organizations are building AI applications. Rather than relying on a single large model for all tasks, companies are increasingly deploying systems of specialized models that work together—what NVIDIA calls "model ensembles."
Why it matters
As AI moves beyond chatbots to autonomous agents that run continuously, enterprises need architectural flexibility. The ability to route tasks to appropriate models—balancing cost, latency, and accuracy—becomes critical at scale. Organizations also want control over where sensitive workloads run, whether on-premises, at the edge, or in the cloud.
Built for high-volume specialized tasks
Nemotron 3.5 Lightning targets specific use cases within larger agent workflows: code review, security monitoring, tool invocation, and domain-specific queries. According to NVIDIA's blog post, the model delivers up to 4x faster output speed and 30% faster agentic task completion compared to other models in its class.
The model was developed with input from the Nemotron Coalition, whose members contributed evaluation methods, inference software, and datasets. NVIDIA says organizations can customize Lightning using NeMo on their own data to improve accuracy for specialized tasks.
Early adopters include CrowdStrike for cybersecurity workloads, Harvey with Trajectory for legal services, and CodeRabbit with Baseten for code review. Lila Sciences is working on reasoning capabilities for physical and life sciences applications, while Fastino Labs reports strong results in software development, finance, and healthcare.
Intelligent routing across model portfolios
NeMo Switchyard addresses a practical challenge: how to automatically direct each request to the most suitable model without manual integration work. The open-source library plugs into existing agent frameworks and lets developers build custom routers based on their priorities—whether optimizing for quality, speed, or cost.
NVIDIA's internal benchmarks suggest Switchyard can maintain frontier-level accuracy while reducing task completion costs to roughly one-third of what a single premium model would cost.
Partners integrating the technology include Kong, which will deliver routing through Kong AI Gateway; LiteLLM, adding Switchyard as a plug-in to its proxy layer; and LangChain, which achieved 74% lower costs on multi-turn agent tasks by routing only 7% of calls to a frontier model.
Cognition integrated Switchyard into Devin Desktop for internal NVIDIA use, achieving near-frontier performance while cutting mean cost by 28%. Ramp reported matching frontier model performance while reducing costs by 58% and runtime by 33%.
Deployment flexibility
Nemotron 3.5 Lightning can run on local infrastructure—NVIDIA RTX PCs, DGX systems, and Jetson edge devices—or scale across data centers and cloud environments. The model is available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as an NVIDIA NIM microservice, as well as through NVIDIA Cloud Partners.
NeMo Switchyard is available on GitHub, with partner platform integrations coming soon.
NVIDIA also released Nemotron-RL-Agentic-Terminal-Pivot, a reinforcement learning dataset used to post-train Lightning for coding agent capabilities. As with previous Nemotron releases, NVIDIA is publishing training data and techniques where licensing permits, enabling traceability and auditing.
These details were first reported by NVIDIA on its official blog.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

