Cognition reports 4.8x token throughput boost with Nvidia's Vera Rubin NVL72 rack
Cognition, the team behind the Devin coding agent, has reported significant performance gains using Nvidia's Vera Rubin NVL72 rack. According to early tests, the system delivered up to a 4.8x increase in total token throughput for SWE-2 inference workloads compared to the GB200 NVL72. This figure, sourced from a September 30 post by Nvidia about CoreWeave's deployment, highlights the potential of the new hardware in AI-driven coding tasks.
CoreWeave announced the availability of Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet networking at its Fully Connected event in San Francisco. Cognition was named as the first customer running production workloads on Vera Rubin, having received its first production racks in September. The 4.8x performance boost was measured by Cognition through a subset of tasks from FrontierCode, though the number remains unverified by external parties.
Nvidia also highlighted the CPU capabilities of the Vera Rubin rack, which includes 128 CPUs and 11,264 cores in a single rack. The company claims this setup supports more than 11,000 concurrent environments, though this is based on arithmetic core count rather than measured concurrency. Additional tests showed more than 3x faster agentic sandbox startups and a 1.7x performance gain on Terminal-Bench across all passing tasks.
Cognition emphasized the cost efficiency of agentic coding workloads, noting that long contexts, high concurrency, and token volumes make cost per token a critical factor. The company has scaled to thousands of GPUs on CoreWeave in nine months, running various workloads for Devin. However, key details such as the price of a Vera Rubin NVL72 rack, installation count, and general availability date remain undisclosed.