Together AI IBM and NVIDIA team up to scale enterprise AI inference
Together AI has announced a major collaboration with IBM and NVIDIA to expand its enterprise-grade AI inference capacity. The partnership introduces a dedicated cluster of NVIDIA B300 GPUs running on IBM Cloud, supported by NVIDIA Spectrum-X Ethernet networking. This marks the first large-scale inference cluster of its kind on IBM Cloud, with Together AI as the inaugural customer.
The infrastructure combines NVIDIA's high-performance silicon, IBM's enterprise cloud services, and Together AI's inference platform. The goal is to deliver scalable, reliable, and secure AI inference for open models, catering to enterprises and AI-native companies. The collaboration aims to ensure that open-source AI can run at scale with the same reliability as closed systems.
Together AI highlights the growing demand for open models, emphasizing that they allow enterprises to maintain data sovereignty while achieving frontier-level performance at a lower cost. The company serves hundreds of trillions of tokens monthly to over a million developers, with demand continuing to rise. The new infrastructure is designed to meet this escalating need, providing enterprise-grade inference with robust security and guardrails.
The partnership underscores a commitment to making open-source AI accessible and efficient. By leveraging NVIDIA's advanced hardware and IBM's cloud expertise, Together AI aims to set a new standard for running open models in production environments. This initiative represents a significant step toward democratizing AI technology.