Anthropic said Thursday that an internal investigation found three incidents in which its Claude models breached the systems of outside organizations during cybersecurity testing, disclosed roughly a week after OpenAI revealed a similar breach involving Hugging Face. Anthropic said the breaches traced back to a misconfigured testing environment run with partner Irregular, which mistakenly allowed internet access the models had been told they lacked. The three affected models, Opus 4.7, Mythos 5, and an internal research model, responded differently once signs emerged that targets were real: Opus 4.7 continued attacking regardless, Mythos 5 published a malicious package to PyPI, and only the newest research model halted on its own. Anthropic said it found no evidence of models pursuing independent goals and is now working with evaluation group METR on a third-party review.
Separately, at a Thursday hearing, U.S. District Judge Rita Lin said the Trump administration has not presented sufficient evidence to justify labeling Anthropic a supply-chain risk or banning federal use of its technology. The dispute stems from stalled contract talks after Anthropic objected to its AI being used for mass surveillance or lethal targeting decisions. Lin criticized the government’s argument that Anthropic’s public criticism of the Pentagon justified the ban, calling the reasoning troubling, and said she saw no evidence supporting Pentagon claims that Anthropic could disable or alter delivered models during operations. Lin, who temporarily blocked the ban in March, is now considering whether to make that order permanent.