Google's TPUv7 Ironwood Challenges NVIDIA's Dominance in AI Inference Chip Market
Google's Tensor Processing Units (TPUs) are making strides in the AI inference chip market, challenging NVIDIA's dominance. According to SemiAnalysis, Google's TPUv7 Ironwood outperforms NVIDIA's B200 and B300 chips in terms of performance per dollar, with up to 50% higher performance per dollar.
The test results show that at an interaction speed of 100 tokens per user per second, Ironwood's inference cost per million tokens is approximately $0.181, which is lower than the $0.222 for the B200 and $0.276 for the B300, representing cost savings of about 19% and 34%, respectively.
Google is also advancing its next-generation native PyTorch backend, named TorchTPU, which will further lower the barrier for external developers to adopt TPUs. The initiative aims to replace the previously limited TorchAX solution and is expected to be open-sourced in October.