The Chinese AI ecosystem has coalesced around Mixture-of-Experts (MoE) architecture and wide Expert Parallelism (EP) deployment to run efficiently on older chips and lower bandwidth HBM. This approach, which increases throughput per GPU by up to 80% when moving from EP=8 to EP=32, reduces pressure on HBM demand and bandwidth. Huawei's UB network in Ascend 950 Pods (1024 NPUs) and SuperPods (8192 NPUs) offers 2TB/s chip-to-chip communication, comparing favorably to NVIDIA's NVLink for GraceBlackwell at 1.8TB/s.
Chinese AI infrastructure is adapting to hardware constraints by optimizing model architectures and leveraging domestic high-bandwidth interconnects, reducing reliance on cutting-edge HBM and GPUs.