Moore Threads Claims 0.06% Deviation from International GPU Training Results
The Chinese GPU maker pre-trained a 236B-parameter Mixture-of-Experts (MoE) model from scratch on its 10,000-GPU-class domestic cluster. This achievement, announced at the World AI Conference in China, marks a significant step for domestic AI compute previously dominated by the Huawei ecosystem.