A new day, a new alarm related to Artificial Intelligence, a new doubt if we are facing an exercise in transparency or promotion. On Friday night, in a statement, OpenAI acknowledged that some of its autonomous Artificial Intelligence agents interacted improperly with websites of US Government agencies. This happened this summer without anyone giving explicit instructions to do so, without being necessary for any of the assigned tasks. It does ensure that there were no actual security violations or breaches, but rather "unexpected" behaviors that need to be studied.
"An agent attempting to complete a search-based training task accessed a public chatbot service exploiting a vulnerability in our internet access restrictions. Previously, the agent had made queries through our search tool and attempted, unsuccessfully, to directly access search engines (...) Our mismatch monitoring system detected the behavior in 15 minutes, and a person started reviewing it three minutes later. Execution was halted 2.5 hours later. All training, evaluation, and inference using tools (in a broad sense) of our most capable models remain suspended," the company stated.
The text states that this incident is much less severe than some previous ones, "but being the first since we reinforced security after the Hugging Face incident, it provides an important signal on where to focus the next phase of this work." However, according to protocols, the exercise should have automatically stopped when the irregular behavior was detected, yet it had to be manually stopped almost three hours later.
"An extensive and ongoing review is underway regarding how our agents use internet access during training and evaluation. We have not been as fast as we would have liked, but we are trying to balance our desire for transparency with gaining a clear understanding of petabytes of agent activity records, and working with the affected organizations. We are prioritizing as best we can based on severity, and adding resources. Hugging Face remains the most serious event we have seen. We will be as transparent as we can, subject to things like vulnerabilities in other companies that our agents have encountered, which will be their decision to disclose or not," wrote Sam Altman, the founder of OpenAI, on his X account.
The task that started it all involved gathering information about a specific person who had posted a blog entry. "The agent was provided with biographical data and clues from said entry. The task did not require the agent to test network controls or access reference answers, and we consider agent behavior that circumvents restrictions or pursues an unreasonable goal as an example of a mismatch." In a short time, it interfered with the websites of the Department of Education, with agents unsuccessfully attempting to access information from its Office for Civil Rights. The Department of Commerce (Census Bureau), where they did manage to obtain data, using access credentials they found on the internet on their own. Or the Securities and Exchange Commission (SEC), as the agents posted publicly accessible site information on a forum.
Alarm among cybersecurity experts
The incident now registered adds to dozens of incidents in recent weeks. Security breaches suffered or caused by models from OpenAI, Anthropic, DeepMind (Google), or Meta have alarmed cybersecurity and AI security experts worldwide. This case comes just days after it emerged that another OpenAI AI agent had unauthorized access to an Australian public health system portal in June 2026, obtaining both public and private files. A hack that Australian authorities discovered three months after it occurred.
The mentioned Hugging Face incident was one of the reasons that Anthropic CEO Dario Amodei cited when he called for a general slowdown in AI development in the sector two weeks ago. His request, quickly backed by OpenAI CEO Altman, Elon Musk, and others, has sparked a global debate on the need for greater regulation. "The world is right to be afraid, but the world must trust that we will do the right thing because it is the right thing and because we are aware of the magnitude of all this," Altman stated this month.
Just yesterday, simultaneously with the notification of this incident, a report from engineers at the San Francisco-based start-up Parse provided more specific details of the attack on Hugging Face, when OpenAI agents hacked the software company. A process with a million service links that OpenAI agents created between July 9 and 13 to accomplish their task. In simplified terms: an AI system attempted to use another AI model to pass the classic robot detection test (with questions like the capital of France) while trying by all means to access another company's computers.
