The Opus 5 model has been released, consistently outperforming Fable 5 while costing half as much. GPT-5.6 xHigh, running on Cerebras, is expected next week at 750 tokens/s, with GPT-6 and an updated Fable 5 version to follow soon.
New model releases with significant performance and cost advantages will drive shifts in inference compute demand and competitive positioning.