AI

NVIDIA Launches Nemotron 3.5 Lightning and NeMo Switchyard

New 30B-parameter model and routing library aim to optimize multi-agent AI systems across deployment environments.

Omega Editorial· August 11, 2026· 3 min read

NVIDIA has released two new tools designed to give enterprises more control over how they deploy and manage AI agent systems: Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, and NeMo Switchyard, an open-source routing library.

The releases reflect a shift in how organizations are building AI applications. Rather than relying on a single large model for all tasks, companies are increasingly deploying systems of specialized models that work together—what NVIDIA calls "model ensembles."

Why it matters

As AI moves beyond chatbots to autonomous agents that run continuously, enterprises need architectural flexibility. The ability to route tasks to appropriate models—balancing cost, latency, and accuracy—becomes critical at scale. Organizations also want control over where sensitive workloads run, whether on-premises, at the edge, or in the cloud.

Built for high-volume specialized tasks

Nemotron 3.5 Lightning targets specific use cases within larger agent workflows: code review, security monitoring, tool invocation, and domain-specific queries. According to NVIDIA's blog post, the model delivers up to 4x faster output speed and 30% faster agentic task completion compared to other models in its class.

The model was developed with input from the Nemotron Coalition, whose members contributed evaluation methods, inference software, and datasets. NVIDIA says organizations can customize Lightning using NeMo on their own data to improve accuracy for specialized tasks.

Early adopters include CrowdStrike for cybersecurity workloads, Harvey with Trajectory for legal services, and CodeRabbit with Baseten for code review. Lila Sciences is working on reasoning capabilities for physical and life sciences applications, while Fastino Labs reports strong results in software development, finance, and healthcare.

Intelligent routing across model portfolios

NeMo Switchyard addresses a practical challenge: how to automatically direct each request to the most suitable model without manual integration work. The open-source library plugs into existing agent frameworks and lets developers build custom routers based on their priorities—whether optimizing for quality, speed, or cost.

NVIDIA's internal benchmarks suggest Switchyard can maintain frontier-level accuracy while reducing task completion costs to roughly one-third of what a single premium model would cost.

Partners integrating the technology include Kong, which will deliver routing through Kong AI Gateway; LiteLLM, adding Switchyard as a plug-in to its proxy layer; and LangChain, which achieved 74% lower costs on multi-turn agent tasks by routing only 7% of calls to a frontier model.

Cognition integrated Switchyard into Devin Desktop for internal NVIDIA use, achieving near-frontier performance while cutting mean cost by 28%. Ramp reported matching frontier model performance while reducing costs by 58% and runtime by 33%.

Deployment flexibility

Nemotron 3.5 Lightning can run on local infrastructure—NVIDIA RTX PCs, DGX systems, and Jetson edge devices—or scale across data centers and cloud environments. The model is available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as an NVIDIA NIM microservice, as well as through NVIDIA Cloud Partners.

NeMo Switchyard is available on GitHub, with partner platform integrations coming soon.

NVIDIA also released Nemotron-RL-Agentic-Terminal-Pivot, a reinforcement learning dataset used to post-train Lightning for coding agent capabilities. As with previous Nemotron releases, NVIDIA is publishing training data and techniques where licensing permits, enabling traceability and auditing.

These details were first reported by NVIDIA on its official blog.

#nvidia#nemotron#agentic ai#model routing#open source ai#enterprise ai

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in AI

AI· 3 min read

Riot Platforms Signs $9B Anthropic Deal, Pivots to AI Hosting

The bitcoin miner will lease 191 megawatts of Texas power capacity to the AI company in a 20-year agreement that marks a broader industry shift.

Via AI Watch · Aug 11, 2026
AI· 3 min read

Memory Chips Now the Primary AI Infrastructure Bottleneck

SpaceX CEO identifies supply constraints that could sustain pricing power for chipmakers through 2027 and beyond.

Via AI Watch · Aug 11, 2026
AI· 3 min read

Blind Runner Completes NYC Half Marathon Using AI Glasses

Thomas Panek relied on Meta's smart eyewear to navigate the 13.1-mile course, marking a breakthrough for assistive technology in sports and daily life.

Via AI Watch · Aug 11, 2026