Nvidia has introduced two new AI services: Nemotron 3.5 Lightning, a customizable open model optimized for high-volume agentic tasks, and NeMo Switchyard, a model routing tool that directs workflows to the most suitable AI model based on task complexity and enterprise needs.

  • Nemotron 3.5 Lightning is a 30 billion-parameter model supporting fast, customizable AI workflows.
  • NeMo Switchyard routes tasks to AI models based on performance, cost, and task-specific needs.
  • Both offerings emphasize ease of customization and integration into existing enterprise AI systems.

What happened

Nvidia unveiled two major AI services aimed at enhancing how enterprises deploy and optimize AI models. The first, Nemotron 3.5 Lightning, is a large-scale, customizable mixture-of-experts model designed to accelerate high-volume, always-on agentic AI tasks. This new model boasts up to four times the output speed and 30% faster agent task completion compared to its peers. The second announcement is NeMo Switchyard, an open-source model routing system that intelligently allocates tasks to the most appropriate AI model available within an enterprise’s ecosystem.

This new suite builds upon Nvidia’s existing Nemotron family, introduced previously with Nano, Super, and Ultra versions, by emphasizing rapid customization and efficient deployment. Enterprises can post-train Nemotron 3.5 Lightning on their hardware using proprietary data, enabling tailored AI capabilities without significant infrastructure overhead. At the same time, NeMo Switchyard accommodates multiple models, routing requests based on criteria such as quality, cost, and speed to align with specific business priorities.

Why it matters

As AI adoption expands, enterprises face challenges not only in choosing powerful models but also in optimizing the right model for each specific task. NVIDIA’s new solutions address this by shifting the focus from raw power to contextual fit and operational efficiency. Nemotron 3.5 Lightning’s ease of customization lowers barriers for companies to quickly refine models to their own domain data, significantly reducing time and cost related to post-training.

Meanwhile, NeMo Switchyard introduces dynamic management of AI workflows by routing tasks intelligently between different models based on their strengths and resource costs. This minimizes wasted computational power, reduces operational expenses, and improves overall AI response accuracy. By combining these two tools, enterprises can implement a more agile, cost-effective, and specialized AI infrastructure that supports complex workflows varying from simple sorting to high-level reasoning.

What to watch next

Enterprise adoption and integration of Nemotron 3.5 Lightning and NeMo Switchyard will be a key indicator of their impact on the AI ecosystem. Success stories like CodeRabbit Inc., which managed a cost-effective, rapid model training process using Lightning, suggest broad operational benefits. Observers should track how industries leverage these tools to solve task-specific challenges, especially in environments requiring diverse AI skills such as coding, contextual inference, and specialized domain knowledge.

Additionally, Nvidia’s release of training recipes and datasets signals a collaborative approach that could foster hybrid models combining Nvidia’s data with proprietary enterprise inputs. The ongoing evolution of AI model routing strategies and their customization within NeMo Switchyard will also be essential to watch, as enterprises refine their priorities around cost, speed, and quality, potentially influencing the broader market for AI model orchestration solutions.

Source assisted: This briefing began from a discovered source item from SiliconANGLE. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings