AI

NVIDIA Launches Nemotron 3.5 Lightning and NeMo Switchyard

New 30B-parameter model and routing library aim to optimize multi-agent AI systems across deployment environments.

Omega Editorial· August 11, 2026· 3 min read

NVIDIA has released two new tools designed to give enterprises more control over how they deploy and manage AI agent systems: Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, and NeMo Switchyard, an open-source routing library.

The releases reflect a shift in how organizations are building AI applications. Rather than relying on a single large model for all tasks, companies are increasingly deploying systems of specialized models that work together—what NVIDIA calls "model ensembles."

Why it matters

As AI moves beyond chatbots to autonomous agents that run continuously, enterprises need architectural flexibility. The ability to route tasks to appropriate models—balancing cost, latency, and accuracy—becomes critical at scale. Organizations also want control over where sensitive workloads run, whether on-premises, at the edge, or in the cloud.

Built for high-volume specialized tasks

Nemotron 3.5 Lightning targets specific use cases within larger agent workflows: code review, security monitoring, tool invocation, and domain-specific queries. According to NVIDIA's blog post, the model delivers up to 4x faster output speed and 30% faster agentic task completion compared to other models in its class.

The model was developed with input from the Nemotron Coalition, whose members contributed evaluation methods, inference software, and datasets. NVIDIA says organizations can customize Lightning using NeMo on their own data to improve accuracy for specialized tasks.

Early adopters include CrowdStrike for cybersecurity workloads, Harvey with Trajectory for legal services, and CodeRabbit with Baseten for code review. Lila Sciences is working on reasoning capabilities for physical and life sciences applications, while Fastino Labs reports strong results in software development, finance, and healthcare.

Intelligent routing across model portfolios

NeMo Switchyard addresses a practical challenge: how to automatically direct each request to the most suitable model without manual integration work. The open-source library plugs into existing agent frameworks and lets developers build custom routers based on their priorities—whether optimizing for quality, speed, or cost.

NVIDIA's internal benchmarks suggest Switchyard can maintain frontier-level accuracy while reducing task completion costs to roughly one-third of what a single premium model would cost.

Partners integrating the technology include Kong, which will deliver routing through Kong AI Gateway; LiteLLM, adding Switchyard as a plug-in to its proxy layer; and LangChain, which achieved 74% lower costs on multi-turn agent tasks by routing only 7% of calls to a frontier model.

Cognition integrated Switchyard into Devin Desktop for internal NVIDIA use, achieving near-frontier performance while cutting mean cost by 28%. Ramp reported matching frontier model performance while reducing costs by 58% and runtime by 33%.

Deployment flexibility

Nemotron 3.5 Lightning can run on local infrastructure—NVIDIA RTX PCs, DGX systems, and Jetson edge devices—or scale across data centers and cloud environments. The model is available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as an NVIDIA NIM microservice, as well as through NVIDIA Cloud Partners.

NeMo Switchyard is available on GitHub, with partner platform integrations coming soon.

NVIDIA also released Nemotron-RL-Agentic-Terminal-Pivot, a reinforcement learning dataset used to post-train Lightning for coding agent capabilities. As with previous Nemotron releases, NVIDIA is publishing training data and techniques where licensing permits, enabling traceability and auditing.

These details were first reported by NVIDIA on its official blog.

#nvidia#nemotron#agentic ai#model routing#open source ai#enterprise ai

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in AI

AI· 3 min read

Google adds encrypted cloud memory to Private AI Compute

New architecture lets AI assistants retain context across devices while keeping data inaccessible to Google itself through device-held encryption keys.

Via AI Watch · Sep 24, 2026
AI· 3 min read

Zuckerberg Says AI Outpaced Metaverse Hardware, Prompting Shift

Meta's CEO acknowledges the company pivoted strategy after artificial intelligence capabilities advanced faster than affordable holographic technology.

Via AI Watch · Sep 24, 2026
AI· 4 min read

Computer Science Grads Pivot to AI Roles as Entry-Level Coding Jobs Vanish

Recent graduates face a 7.1% unemployment rate as tech giants automate development work and smaller firms seek AI implementation help instead.

Via AI Watch · Sep 24, 2026