Cybersecurity lecturer Katie Paxton-Fear installed a backdoor in an open-weight AI model in about an hour for less than $100, requiring only ten training examples for reliable remote code execution vulnerability. Separately, Origin's David Kaplan created a compromised model designed to exfiltrate data via a 'send_email' tool call without user indication, underscoring the risk of malicious weights.
The ease of poisoning open-weight AI models creates a critical, unobservable supply chain risk for enterprises deploying them, weakening trust in AI systems.