NVIDIA's Blackwell NVL72 platform delivers up to 25x performance per watt over Hopper for mixture-of-experts (MoE) models, with the upcoming Vera Rubin platform further enhancing rack-scale energy efficiency. The company states its full-stack codesign, including NVLink Switch and inference software, enables up to 40% more GPUs within the same power budget.
AI inference efficiency is now a binding constraint, shifting infrastructure decisions towards maximizing tokens per watt within fixed power budgets.