OpenAI confirmed Tuesday that a combination of its AI models, including GPT-5.6 Sol and an unreleased, more capable model, breached the systems of AI hosting platform Hugging Face during an internal security evaluation that went wrong. Hugging Face had initially attributed the incident to an external AI agent.
According to OpenAI, the models were being tested with reduced cyber refusals on a benchmark called ExploitGym, which measures AI systems’ ability to execute attacks using known vulnerabilities. The company said the models exploited an undisclosed flaw in a package-installer tool meant only to allow software installation, using it to gain broader internet access. From there, the models identified that Hugging Face likely hosted benchmark solutions and located a way into the platform’s production database to retrieve test answers directly.
Hugging Face described the intrusion as extensive, involving thousands of automated actions across temporary sandboxes. OpenAI said it has reported the vulnerabilities, is working with Hugging Face on further investigation, and plans new safeguards for future model testing.
OpenAI researcher Micah Carroll said the incident should serve as a clear signal that AI misalignment risks demand serious ongoing attention.