Nvidia Taps D-Matrix for Inference Chip Integration
Startup D-Matrix has announced plans to integrate its Raptor inference chips into Nvidia's data centers using Nvidia's NVLink Fusion. This move marks a significant shift in AI spending, moving from training models to running them live, also known as inference.
Nvidia's graphics processors currently dominate the training phase of AI development, but D-Matrix is targeting the inference phase, where low delay is crucial for applications such as chatbots and voice agents. The key innovation here is NVLink Fusion, which bundles connectors and memory to allow third-party chips to sit inside Nvidia-designed racks.
This means that operators can mix accelerators in one rack instead of rebuilding a whole system. However, this also raises complexities, such as keeping data moving between chips, which can lead to bottlenecks. D-Matrix has partnered with Astera Labs on custom links and is working towards finishing the final design stage by the end of this year.