Nvidia Dominates Data Center Workloads in Recent Benchmark
A recent benchmark from SemiAnalysis shows Nvidia's hardware five times more cost-efficient than AMD on certain data center workloads. AgentX, an open-source benchmark, was used to replay real coding-agent sessions against production inference stacks.
The test focused on GLM 5.3 served through SGLang, where Nvidia reached up to five times better cost efficiency than AMD at 150 output tokens per second per user. At this operating point, competing accelerators would still produce a higher cost per token even if they were free to host and power.
The advantage is attributed to Nvidia's TensorRT-LLM adding boundary-aware incremental tokenization, allowing for faster processing times per turn. This benefit cannot be shown in fixed-length benchmarks, as there is no prior turn to reuse.
AMD's ATOM engine does have some wins, particularly on the Kimi K3 curve between 40 and 60 seconds of latency. However, SemiAnalysis notes that this adoption is limited outside one advertising unit at Alibaba, making it less relevant for customers.