The UK AI Safety Institute's Cyber and Autonomous Systems Team conducted early access testing of GPT-5.6 Sol, finding it significantly improved over GPT-5.5. Sol completed "The Last Ones" in 7 out of 10 attempts, slightly outperforming Claude Mythos 5 which solved it in 6 out of 10 attempts. On the "Doing Life" range, Sol reached step 21 of 23 in 3 of 10 attempts, performing comparably to Claude Mythos 5.
A new frontier model demonstrates advanced offensive cyber capabilities, raising the bar for AI safety and security evaluations.