Nvidia Vera Architecture Signals Shift to Custom Arm Cores for AI Infrastructure
A technical deep dive into Nvidia's upcoming Vera CPU reveals a highly customized Arm architecture designed to eliminate memory and interconnect bottlenecks in massive AI clusters.
Nvidia's upcoming Vera CPU, powered by the custom-designed Olympus Arm cores, represents a fundamental shift in how the company approaches datacenter host processing. Moving away from off-the-shelf Arm Neoverse designs, Nvidia is tailoring its silicon to address the specific, high-throughput bottleneck challenges of modern AI clusters. The architecture prioritizes massive memory bandwidth and interconnect speed over raw single-threaded compute performance, signaling a future where the CPU functions primarily as an orchestrator for GPU-heavy workloads.
At the heart of the Vera silicon are 88 custom Olympus cores supporting two-way simultaneous multithreading, yielding 176 logical threads. While Arm-based server chips have historically favored single-threaded cores to maximize power efficiency and predictable virtualization performance, Nvidia's implementation of SMT suggests a focus on parallel handling of system overhead, network stacks, and storage I/O. This design choice optimizes the CPU to feed data to hungry GPU clusters without stalling, ensuring that the critical accelerator pipelines remain fully utilized.
The memory subsystem of the Vera CPU is equally unconventional, opting for up to 1.5 TB of LPDDR5X memory rather than traditional DDR5 or ultra-expensive High Bandwidth Memory. By leveraging LPDDR5X, Nvidia achieves a compelling sweet spot: significantly higher bandwidth than standard server DDR5, lower power consumption, and a much lower cost profile than HBM. This massive memory pool acts as a high-speed staging area for large language models and dataset caching, directly connected to the processor package to minimize latency.
Crucially, the chip is bound to the rest of the node via a 1.8 TB/s NVLink interconnect. This bandwidth matches the throughput of high-end GPU-to-GPU links, effectively integrating the CPU into the unified memory fabric of the entire server rack. For systems engineers, this means the boundary between host memory and accelerator memory is further blurred. By optimizing Vera for data movement rather than general-purpose compute, Nvidia is reinforcing its vertical integration strategy, making it increasingly difficult for competing CPU vendors to slot their silicon into next-generation AI infrastructure.
Sources
- 01 A deep dive into Nvidia's Vera CPU and the Olympus cores that power it — The Register