Nvidia Aims to Cut AI Costs with Model Routing Platform
Nvidia has introduced a new software platform called NeMo Switchyard to help enterprises manage soaring AI costs and uncertain returns on investment.
The platform blurs the line between expensive proprietary models and open weights alternatives by routing requests to different models in order to optimize for cost, latency, or output quality.
Nvidia claims that Switchyard can cut job completion costs by 74% relative to using Claude Opus 4.8 alone, albeit with an approximately six-point accuracy tradeoff.
The company has also developed several application-specific models, such as Nemotron Parse, which is a small model that excels at one task: explaining the context inside PDFs.