Nvidia Challenges CPU Dominance with Custom Vera Chip
Nvidia has made significant strides in challenging Intel and AMD's CPU dominance with its Vera CPU. The chip is designed for hyperscalers and cloud providers, with eight major companies already signed up to deploy it. Among these are Alibaba, ByteDance, Meta, Oracle, CoreWeave, Lambda, Nebius, and NScale.
Vera's inner workings have largely remained a mystery until recently, but Nvidia released a whitepaper that shed light on its architecture. The chip features 88 custom Armv9.2 cores, 176 threads, and support for up to 1.5 TB of LPDDR5X memory. It also boasts availability as a standalone platform independent of Nvidia's GPUs.
Nvidia's Olympus core is the heart of Vera, featuring a 10-wide decoder and dispatch, eight integer arithmetic logic units (ALUs), six vector/FP pipelines, four load units, and two store units. The chip has a custom neural branch predictor that can explore two branches simultaneously, reducing the likelihood of misprediction and boosting performance.
Additionally, Nvidia claims to have made several innovations to the mid-core, aimed at mitigating pipeline stalls. Memory renaming speeds up store-to-load dependency chains by allowing dependent instructions to execute before a load completes if the relationship to the data can be divined. Value prediction identifies stable dependency chains and predicts future values before they are produced, allowing dependent instructions to execute speculatively while correctness is verified later.
Vera's Superchip features two Vera CPUs connected over NVLink-C2C at 1.8 TB/s of bidirectional bandwidth, for a total of 176 cores and 352 threads. The two chips are fed by 16 SOCAMM2 LPDDR5X memory modules, delivering 2.4 TB/s of aggregate memory bandwidth.