Anthropic and OpenAI Propose Embedding Independent Safety Evaluators Inside AI Companies

Anthropic CEO Dario Amodei proposed in a weekend essay that frontier AI companies embed third-party evaluators with the power to assess model alignment, report safety incidents, and publish findings without editorial control. Amodei committed Anthropic to giving evaluators such as METR and Redwood Research expanded access to its systems, and OpenAI CEO Sam Altman said his company would follow suit.

Evaluators broadly welcomed the proposal but said critical details remain unresolved, including how much access they will actually receive. Adam Gleave of FAR.AI said meaningful oversight would require access to intermediate training checkpoints, post-training environments, and employee interviews, rather than just finished models. Alexander Meinkeof Apollo Research said companies should be able to demonstrate whether models attempted to undermine their own alignment training, something currently unverifiable by outsiders.

Researchers noted past evaluations were often limited by time constraints, citing OpenAI’s week-long review of the Hugging Face incident and a three-day testing window for its GPT-6 Astra model. Henry Papadatos of Safer AI argued voluntary commitments remain fragile without binding regulation, since companies could reverse course during a public crisis.

Meta, Google DeepMind and SpaceXAI have not committed to the practice, though DeepMind’s Demis Hassabis has proposed a separate industry standards body. California’s SB 53 and SB 813, along with the EU AI Act, have begun establishing formal evaluation requirements.

Need Deeper Intelligence on the AI Market?

AI Insider's Market Intelligence platform tracks funding rounds, competitive landscapes, and technology trends across the global AI ecosystem in real time. Get the data and insights your organization needs to make informed decisions.

Related Articles

SK Hynix in Talks With Intel to Manufacture AI Memory Chips in the US

SK Hynix is reportedly in discussions with Intel to manufacture RAM chips in the United States for the first time, according to Reuters, citing anonymous sources.

Amazon Launches Alexa+ AI Assistant in India With Hindi Support

Amazon has launched its conversational AI assistant, Alexa+, in India with Hindi language support, expanding the generative AI-powered assistant’s rollout following earlier launches in the

TypeSafe AI Emerges From Stealth With $40M to Build Machine-Native AI Models

TypeSafe AI, a frontier AI lab focused on building composable, machine-native intelligence, has emerged from stealth with $40 million in seed funding led by DCVC.

Stay Updated with AI Insider

Get the latest AI funding news, market intelligence, and industry insights delivered to your inbox weekly.

$ 0 M

Seed round tracked

Gitar — Code Validation

Get the Weekly Briefing

Funding analysis, market intelligence, and industry trends delivered to your inbox every week.

Need bespoke intelligence?

Our team combines real-time data with decades of sector experience to guide your decisions.

Subscribe today for the latest news about the AI landscape