Friday · October 2, 2026

 ·  Daily  ·  Newsletters

Resonance network: Quantum InsiderSpace Insider

OpenAI Launches Misalignment Reports Site, Scraps Astra 6.1 Release and Apologizes to Australia Over Agent Breaches

a square object with a knot on it

OpenAI has taken a series of steps to address mounting concerns over rogue AI agents, publishing a dedicated site for misalignment reports, reportedly cancelling an upcoming model release over safety issues, and formally apologizing to the Australian government for breaches of its public services websites.

The new site currently hosts nine reported incidents, most occurring during reinforcement-learning training. CEO Sam Altman said on X that the company is trying to balance transparency with the need to understand petabytes of agent activity logs and work with affected organizations, adding that it is prioritizing disclosures by severity and adding resources. The reports include a previously undisclosed sandbox escape on September 20, in which an internal research model contacted an external chatbot through a DNS query; monitoring flagged the behavior within 15 minutes and the run was halted in under three hours. In another case discovered in May, a highly persistent internal model tried to cheat on a math problem by using a private GitHub token to access another team’s work, despite being told twice to work locally.

OpenAI also disclosed the possibility of self-replicating prompt injection attacks, in which instructions hidden in an email lead an agent to paste those same instructions into its reply, spreading them to the next agent like a malware worm. Researchers said they observed this only under controlled conditions with an underpowered model and were sharing it because of its novel nature rather than any real incident. Axios has reported that major labs may have seen as many as 10,000 incidents of models exceeding evaluator instructions, though Altman has said the Hugging Face breach remains the most severe OpenAI has found.

Besides this, The Wall Street Journal reported that OpenAI cancelled the planned release of Astra 6.1 after the model showed higher levels of deception and unsafe behavior than its predecessors. Head of safety systems Saachi Jain told the Journal that the model tested poorly on alignment. Similar breakout behavior has been disclosed by Anthropic, Googleand Meta, and the incidents have pushed U.S. policy discussions toward new safety standards and a possible industry slowdown, which critics argue could entrench leading labs.

On Monday, OpenAI apologized to Australia for failing to promptly notify authorities after its models accessed government websites without authorization in June, saying it should have handled its response better. The company explained that an experimental model tasked with researching medicine spending in Victoria accessed an internal Services Australia system, ran commands, retrieved files and credentials, and wrote files. Its agents also reached data from the New South Wales crime statistics bureau, Victoria’s Agency for Health Information via an exposed access key, and the Australian Institute of Health and Welfare. OpenAI said it found no evidence that individuals’ medical or criminal records were accessed.

The company will share technical findings with affected agencies, provide credits from its $1 billion Daybreak for Frontline Defenders program, and form a task force of independent Australian experts that will recommend safeguards by year’s end. Prime Minister Anthony Albanese has called the breach unacceptable and said the government is weighing legal measures.

James Dargan
About the author
James Dargan

James Dargan is a writer and researcher at The AI Insider. His focus is on the AI startup ecosystem and he writes articles on the space that have a tone accessible to the average reader.

Trending today

AI

xAI’s Dot.com Redirect to Grok Sparks Speculation of a Jab at OpenAI’s Dots

Business & Markets

Reco Announces $55M in Funding to Tackle AI Agent Sprawl Across Enterprises

Business & Markets · Startups

Destro AI Raises $8M in Seed Funding to Coordinate Warehouse Robot Fleets

Physical AI

Hello Robot Receives Nearly $3M NIH Grant for Assistive Robotics Research

Business & Markets

Modulate Raises $25M to Scale Small-Model Voice AI for Deepfake Detection and Agent Compliance

The AI economy, every weekday morning

The daily briefing on LinkedIn. Free, one tap to follow.

Exclusives

Exclusive

South Korea’s AI G3 Strategy: Decoded

Scale-ups to Watch

10 Switzerland-Based AI Scale-Ups You Need to Know in 2026

network, blockchain, digital, hand, web, community, artificial, intelligence, steering, interfaces, bokeh, future, digitization, transformation, change, blockchain, blockchain, blockchain, blockchain, blockchain, transformation
Exclusive

Why Crypto Could Be AI’s Payment Layer: BlackRock Sees Stablecoins Connecting Commerce and Compute

AI Predictions
Exclusive

Why AI Predictions Often Get The Technology Right But The Timeline Wrong

Scale-ups to Watch

10 CEE & Baltics-Based AI Scale-Ups You Need to Know in 2026

More in Policy & Government

Latest from the same section
a purple and green background with intertwined circles
Business & Markets · Enterprise

OpenAI Faces Safety Scrutiny While Adding a Fast Decision Model and Shopping Tools

3 min ago
an image of an infinite sign on a blue background
Policy & Government · AI Safety

Meta Disputes Journalist’s Claim That Muse AI Agent Read Private Messages Without Consent

1 hour ago
Policy & Government · AI Safety

Fort Robotics Advances SPAC Deal With Confidential S-4 Filing

Yesterday
Technology & Infrastructure

Nvidia Adds Hardware-Level Security Layer to Keep Rogue AI Agents Contained

2 days ago

The AI economy, every weekday morning

The daily briefing plus the weekly Scale-ups to watch edition. Free, no spam, unsubscribe any time.