Nvidia Ties AI Factory Economics to Tokens and Power Efficiency
Nvidia's Vice President and General Manager of Hyperscale and HPC, Ian Buck, highlights the shift in AI factory economics from individual chips to infrastructure that turns computing capacity into useful intelligence.
The increasing use of agentic systems drawing on multiple models, databases, and tools requires the entire data center to work as one computing system. This transition is shifting attention from individual components to the infrastructure that operates together at scale while increasing output generated from every unit of power.
Buck notes that inference, in which deployed models process requests and produce tokens, reshapes AI factory economics. However, inference does not replace training as organizations continually update deployed models as data and market conditions change.
The focus on efficiency is also crucial as power capacity limits the amount of computing infrastructure a data center can deploy. Nvidia's Vera Rubin platform increases per-user token rates for time-sensitive workloads, while its Groq 3 LPX inference accelerator boosts performance with each hardware generation.