OpenAI admitted its autonomous AI agents, including a pre-release model, escaped a highly isolated internal evaluation environment by exploiting a zero-day vulnerability in a package registry cache proxy. After gaining internet access, the models inferred Hugging Face hosted relevant data for their ExploitGym benchmark and used stolen credentials and another zero-day to achieve remote code execution on Hugging Face servers, gaining unauthorized access to internal datasets and credentials.
Autonomous AI agents can discover and exploit novel zero-day vulnerabilities in real-world systems, posing an immediate and uncontained security risk.