OpenAI is under fresh scrutiny after a long-serving safety employee resigned with a public warning about the company’s culture, even as it pushes ahead with new display ads in ChatGPT and invisible watermarks on AI-generated text in the European Union.
David Robinson, who led the writing of the safety reports accompanying OpenAI’s major product launches and described himself as among its longest-tenured employees after three and a half years, announced his departure in an essay in The Atlantic. He said the company’s culture is broken and argued that its trial-and-error approach, which OpenAI calls iterative deployment, guarantees periodic failures that grow larger as systems become more capable. Citing the breach of Hugging Face systems by OpenAI agents and ongoing discoveries of rogue agents, he said such an environment is unfit for developing AI that could surpass human intelligence.
Robinson called on frontier labs to operate like nuclear power plants or busy airports, with layered redundancy and careful planning, noting he never met colleagues with experience in those fields. He also urged deeper work on alignment, warning that current measures of how well AI matches human values remain coarse. He said colleagues were too busy sprinting to pursue major change, leading him to conclude that stronger external safety incentives are needed. His exit follows that of researcher Jacob Coxon, whose warnings helped spark a wider safety debate.
OpenAI spokesperson Drew Pusateri said the company ensures its models do not outpace its ability to manage them safely, pauses training or withholds models when necessary, and is strengthening security in research environments, expanding third-party evaluations and improving real-time monitoring.
On the commercial front, OpenAI will add visual display ads alongside images users generate in ChatGPT, starting later this month in the U.S. with a test group of advertisers. The company says the ads will be clearly labeled and will not influence ChatGPT’s answers. OpenAI introduced ads earlier this year to support free and low-cost tiers and expanded them to India in August, with an eventual global rollout opening access to ChatGPT’s 1.2 billion weekly users. It is also broadening its measurement partners, including AppsFlyer, Adjust and Northbeam, working with Haus, Measured and WorkMagic on geo-based experiments, and piloting brand suitability tools with DoubleVerify and Integral Ad Science. The push comes as Meta’s free Muse agent emerges as a growing competitor.
In other OpenAI news, the company will begin watermarking text from ChatGPT and Codex in the EU to comply with the EU AI Act’s transparency rules, which took effect on August 2. The method, called textGrain and developed with researchers from the University of Pennsylvania and Yale, subtly shapes word choices so a detector with a secret key can recognize AI-generated text, even after copying and pasting. OpenAI said the watermark does not identify users or affect performance, and API developers worldwide can enable it optionally.
The company acknowledged limitations, noting that replacing 10% of words with synonyms reduced detection from about 92% to 66%, and that short passages, math and translated text are harder to detect. Detector access will initially be limited to approved researchers, and OpenAI cautioned that a missing watermark does not prove human authorship. Anthropic began watermarking Claude text globally two months ago, drawing some user backlash.