GPT-5.6 Sol Demonstrates Stronger Adversarial Defense Against Older Models
Older 'defending' models like GPT-Red would outperform prior generations, but newer adversarially trained models exhibit significantly higher resistance to attacks. GPT-5.1 failed 95% of the time against adversarial attacks, while GPT-5.6 Sol fails less than 10% as it learned to defend over time.
So What
Model robustness against adversarial attacks is improving, reducing the risk of prompt injection and other exploits for deployed AI systems.