Nvidia Partners with d-Matrix for Disaggregated Inference Solution
Nvidia has partnered with d-Matrix to develop an alternative disaggregated inference solution. The two companies announced a multi-year roadmap for integrating the d-Matrix 2nd generation platform, code-named Raptor.
The partnership involves pairing d-Matrix's SRAM-based AI technology with Nvidia GPUs for encoding and tokenizing input queries. This approach is seen as a customer-retention and time-to-value strategy by Nvidia, which gives customers a deployable solution to low-latency token generation now.
Nvidia's competitor, Groq, has developed the LPU3, which solves the same problem using fast SRAM for disaggregated inference processing. However, the partnership with d-Matrix allows customers to extract more value from existing Nvidia GPUs and deploy the d-Matrix Corsair accelerator today alongside existing or new GPUs.
The longer-term risk for Nvidia is that decode becomes a large enough portion of inference spending that specialist silicon gains durable account control. However, Nvidia's response is pragmatic: it will partner while the category is forming, remain indispensable in high-value portions of the stack, and be prepared to substitute its own LPU3/LPX solution when customer and workload justify tighter vertical integration.