The new oMLX 0.5.2 release introduces custom Metal kernels for models like GLM, MiniMax, DeepSeek V4, and Qwen, with prefill improvements up to +99% for GLM-5.2 on M3 Ultra. It also adds Lightning MTP native speculative decoding, accelerating token generation for Qwen3.6-35B-A3B from 89.6 to 140.4 tokens/second.