Jacob Coxon, who spent three years on pretraining research at OpenAI and Anthropic, resigned this week and publicly accused both firms of failing to act responsibly in developing self-improving AI systems. Coxon argued that emerging superhuman systems could acquire dangerous capabilities and warned that many people building the technology privately fear it could kill humanity by the end of the decade, even as they publicly downplay those concerns.
Coxon said Anthropic understands the risks but feels compelled to continue racing toward advanced AI because it believes competitors won’t act responsibly otherwise. He called for lab researchers to reconsider whether racing toward recursive self-improvement without deeper safety understanding is the right path, and expressed hope that incidents like AI agents breaching Hugging Face’s servers could push labs toward coordinated pacing agreements.
Anthropic colleague Evan Hubinger echoed these concerns, estimating a greater than 10% chance AI could cause human extinction within the next decade and acknowledging Anthropic lacks a clear plan for aligning superintelligent systems. Meanwhile, well-funded startups including Ricursive Intelligence, Recursive Superintelligence, and Jeff Dean’s Discovery Loop continue pursuing recursive self-improvement. ControlAI’s Connor Leahy described such self-improving systems as an adversary rather than a tool. Lawmakers in the U.S. and U.K. have introduced legislation this week aimed at restricting superintelligence development.