Moonshot AI's K3 model, a 3T parameter LatentMoE, employs a stable version of Nvidia's LatentMoE technique, optimized for training by fixing load-balancing estimator noise at extreme sparsity. The model also uses Quantile Balancing to ensure experts receive exact token quotas, eliminating host round-trips and compile-time buffer sizing issues. Additionally, K3 incorporates Attention Residuals, a depth trick allowing layers to attend over earlier layers, reducing overhead for significant gains.