Harvard and MIT Researchers Uncover Hidden Unsafe Thoughts in Large Reasoning Models
Researchers from Harvard, MIT, and other universities found that chain-of-thought reasoning can expose dangerous content even when final answers appear harmless. Their proposed mitigation, adaptive multi-principle steering, reduced unsafe reasoning by up to 77.2% and unsafe final answers by 48.1%, while maintaining over 97% accuracy.
So What
AI model safety requires checking the entire reasoning path, not just final outputs, to prevent hidden unsafe content.