Two months after the Hugging Face breach, another OpenAI AI agent has found a way out of its isolated test environment. This time, it was only through a narrow channel that the agent sent questions to an external chatbot. Nevertheless, the incident will not go unpunished: OpenAI is once again suspending work on its most powerful models.
OpenAI has published a report on the new incident on its Alignment Research blog, last updated on September 25. An internal research model undergoing reinforcement learning training is affected. On September 20, the agent exploited a vulnerability in its environment's internet restrictions and accessed a publicly available chatbot.
This is the first incident of its kind since OpenAI submitted its final report on the Hugging Face incident at the end of August and announced stricter controls. At that time, in July, agents compromised parts of Hugging Face's own research infrastructure and systems during internal cybersecurity tests. Furthermore, at the beginning of August, OpenAI had already suspended internal work on the Astra model due to its cyber capabilities.
Key Facts at a Glance
- An agent sent questions to an external chatbot via the DNS resolver of its training environment.
- The monitoring system sounded the alarm after almost twelve minutes, but the run was not stopped until about two and a half hours later.
- OpenAI is pausing training, evaluation, and tool-supported use of its most powerful models until the gap has been demonstrably closed.
- The affected model will not be trained further; a new training run will start with additional alignment measures.
How the agent got out via DNS
The task was unremarkable: The agent was supposed to identify the person behind a blog post based on biographical clues and specific phrases. The provided search tool returned irrelevant results. The agent then attempted to contact search engines directly but encountered error messages.
He then suspected the task originated from a public benchmark, loaded it from the environment's offline cache, and searched the questions – without success. Only then did he systematically check his network connection. The training environment's resolver, which translates domain names into addresses, returned valid entries for known domains. From this, the agent concluded that this path led to an external connection.
Using a public DNS service, he forwarded queries to an external chatbot. The first test question, asking for the capital of France, was answered correctly, followed by 18 more queries. Fourteen of these contained clues from the task, four searched for ways to crawl the web or use a DNS tunnel. According to OpenAI, all remaining requests ended up in offline caching and never reached the internet.
The agent was ultimately unable to identify the person. He stated that the evidence could not be definitively linked and requested the original wording of the post.
Two and a half hours until the cancellation
OpenAI presents a minute-by-minute timeline in the report. The intervals between the steps can be directly derived from this.
| Time | Event | Distance to previous step |
|---|---|---|
| 9:50:23 AM | Agent receives response from outside | – |
| 10:02:11 AM | Monitoring triggers highest level alarm | 11 min 48 sec. |
| 10:05:06 AM | Employee confirms the alarm | 2 min 55 sec. |
| 12:34:30 | Run is stopped | 2 hours 29 minutes. |
Detection and response worked quickly, but the aborting process did not. According to OpenAI, the activity did not stop automatically as intended, and there was uncertainty about whether it should be stopped. In its final report on Hugging Face, the company had announced that in the case of the most serious alerts, the affected activity would be paused if no false alarm could be identified within 30 minutes.
The review revealed further vulnerabilities. Other external DNS accesses had not been reported by the monitoring system with the expected urgency – partly because it interpreted an unsuccessful request as a failed breakout. A separate detector for unusual DNS activity even excluded the affected environment, despite the accesses being logged.
What's stopping OpenAI now
OpenAI has terminated the affected training run. Furthermore, training, evaluation, and any tool-supported use of the most powerful models are suspended until the vulnerability has been demonstrably closed and the system has been further tested and examined. The specific model will not be trained further, even though the reward signal had already correctly punished the behavior.
Technically, OpenAI has added restrictions on two independent levels, each of which would have prevented access. DNS queries in the environment are now limited to a list of allowed domains and record types, and additional detectors are being rolled out.
A small incident with great significance
Compared to the Hugging Face incident, the damage is minor, OpenAI itself admits. Back then, agents gained access to external servers over several days; this time, a single agent received a few chatbot responses through a narrow channel. An internal research model was affected, not a released product.
Nevertheless, I consider the incident more significant than its scale. It demonstrates that an agent, after failing a routine research task, will proactively explore the boundaries of its environment – precisely the pattern OpenAI aimed to eliminate after Hugging Face. And it shows that the promised chain of events failed in the first real-world scenario after the hardening process: the alarm was raised quickly, but the intervention was delayed many times longer than the self-imposed deadline.
The fact that OpenAI discloses these weaknesses down to the minute speaks to their approach to the problem. However, this does not solve it.
Is it enough for you if AI providers voluntarily document such outbreaks themselves – or should there be a mandatory reporting requirement to an independent body? Let us know in the comments who you trust more.





