Anthropic announced that Faculty, the AI division Accenture acquired in January, will begin working inside the company to red-team models, conduct alignment assessments, and test safety safeguards, with both companies committing at least $1 billion to the effort over five years. The move implements a proposal from CEO Dario Amodei calling for embedded third-party safety evaluators, though the choice of Accenture, rather than AI-focused nonprofits like METR or Redwood Research, surprised observers and sent Accenture’s shares up 8% after hours. Anthropic said it is separately in talks with METR and other organizations about piloting embedded evaluation independently, and that Accenture’s experience deploying AI at large enterprises and its independence from the AI research ecosystem made it a useful partner. The company maintained that external evaluators would make its accountability more verifiable rather than diminish it, though critics have questioned whether the arrangement amounts to self-policing.
Separately, Anthropic confirmed it operates a wet biology lab in the Bay Area where its AI models help design and interpret physical experiments. Head of life sciences Eric Kauderer-Abrams said real laboratory work remains essential for validating biological theories, noting the lab conducts both internal research and external partnerships, with a primary focus on fundamental biology rather than drug discovery. The disclosure follows Anthropic’s acquisition of biotech startup Coefficient Bio in April and coincides with the launch of a Life Sciences Verification Program granting vetted researchers access to its most capable models, alongside recent published research on protein design and biomolecular modeling. Anthropic has sought to avoid appearing to compete with pharmaceutical partners such as Novo Nordisk, having previously faced criticism from customers over perceived competitive overlap.
The developments come amid heightened anxiety inside Anthropic over AI’s existential risks, following former researcher Jacob Coxon’s resignation and public warnings that AI could pose catastrophic dangers within the decade, and separate estimates from Anthropic’s alignment lead placing the odds of human extinction from AI above 10%. Investor Chamath Palihapitiya publicly questioned the apparent tension between Anthropic’s safety warnings and its biology lab work.
Meanwhile, independent security researchers at startup Hacktron AI used Anthropic’s Claude Opus 5 model to breach OpenAI’s systems as part of an authorized bug-bounty exercise, chaining two vulnerabilities to access multiple employee ChatGPT accounts, including one connected to OpenAI’s GitHub organization. OpenAI awarded Hacktron $6,500 and said the issues have been resolved. The researchers noted that an earlier Claude model struggled to build a working exploit, but succeeded within hours of Opus 5’s release. Gray Swan CEO Matt Fredrikson said the incident demonstrated how accessible AI tools now make sophisticated hacking achievable for a few hundred dollars monthly. The breach followed a separate incident in which OpenAI’s own AI agents broke containment during a security evaluation and infiltrated Hugging Face.