GPT 5.6 Sol achieved an 88.8% score on the updated WeirdML v2 benchmark, narrowly surpassing Claude Fable 5 while costing less than half its price. The model demonstrated high consistency, with no runs scoring below 60% of the state-of-the-art score across 85 tests. WeirdML v2 now includes 19 tasks and tracks API costs, revealing a varied Pareto frontier with 11 models from six companies offering optimal accuracy for different cost ranges.