TileRT Breaks AI Inference Performance Record with 469 Tok/s on AMD Instinct MI355X
TileRT has achieved a new milestone in AI inference performance, reaching 469 tokens per second (tok/s) on AMD Instinct MI355X GPUs. This result surpasses NVIDIA's GB300 NVL72 system by over 100 tok/s and was accomplished using GLM-5.3 on 8 AMD Instinct MI355X GPUs.
The achievement is notable not only for its raw speed but also for its ability to maintain efficiency over long contexts. When input context expanded from 1,000 to 1 million tokens, a 1,000x increase, TileRT still retained about two-thirds of its short-context performance, delivering 425 tok/s at the maximum context length.
The TileRT team attributes these results to its optimization strategies tailored for AMD's CDNA 4 architecture. Key innovations include streaming synchronization, persistent engine kernels, and a fused communication-compute model that minimizes synchronization bottlenecks during data transfer.
This advancement underscores the growing importance of optimizing inference speed in AI deployments. The ability to sustain high performance over long-context workloads positions TileRT as a strong contender in latency-sensitive applications like AI-assisted coding, high-frequency trading, and real-time decision-making.