Mythos Preview demonstrated superior performance over both GPT-5.6-Sol and Mythos 5 across the first 30 million tokens, also achieving slightly faster processing speeds. Open-source models are reportedly lagging frontier models by seven months on UK AISI's long-horizon cyber ranges, with GLM-5.2 performing comparably to Opus 4.5 on specific cyber range tasks.
Frontier model capabilities continue to advance, widening the performance gap against open-source alternatives in critical areas like long-horizon reasoning.