GPT-5.6 Sol achieved a 72-73% score on the DeepSWE coding-agent benchmark at approximately $8.4 per task, surpassing Claude/Fable-5's ~70% score which costs $13-22 per task. The DeepSWE benchmark evaluates models on complex software engineering tasks within real open-source repositories, including repo navigation, bug diagnosis, and multi-step repair across 91 repositories and five languages.
OpenAI's new model version delivers superior agentic coding performance at a significantly lower inference cost, increasing pressure on competitors like Anthropic.