NVIDIA's upcoming Rubin GPU, part of the Vera Rubin platform, features 72 Rubin GPUs and 36 Vera CPUs per NVL72 rack-scale system, offering 288GB of HBM4 memory with 22 TB/s bandwidth. The architecture includes an improved Tensor Memory Accelerator for MoE models, doubled K-dimension throughput in Tensor Cores, and up to 4X faster Softmax operations compared to Blackwell. These enhancements aim to increase inference efficiency from GPU to data-center scale, particularly for agentic AI workloads.