Microsoft's MAI family of domain-specific models, detailed at its Build conference in June, is now replacing OpenAI's models in Microsoft products, driven by the need for cheaper and more efficient AI inference. This strategy allows Microsoft to deploy the right AI for specific tasks, freeing up memory and improving hardware utilization, while also leveraging its custom Maia 200-series AI accelerators for optimized performance.
Hyperscalers are prioritizing cost-efficient, domain-specific AI inference over large frontier models, shifting compute demand towards custom silicon and away from general-purpose GPU clusters.