Cybersecurity

An AI from the creator of ChatGPT rebels and attacks another company

OpenAI acknowledges that its advanced models circumvented safety measures to "cheat" on a capabilities test

23/07/2026 - 11:01 h.

BarcelonaAn artificial intelligence model from OpenAI, the North American multinational creator of ChatGPT, caused an "unprecedented incident" in global cybersecurity. As the company acknowledged in a statement, one of its training models broke free from its limitations and attacked Hugging Face, a company also dedicated to AI.

In a statement, the technology company led by Sam Altman explained that it created an agent based on two of its most advanced models, GPT 5.6 Sol and another "even more capable" one that is in the development phase, to test its capabilities. During this training, OpenAI experts programmed the agent to "test advanced exploitation patterns" of digital weaknesses.

Cargando
No hay anuncios

These types of tests, the company explains, are carried out without the internal limitations that "prevent models from engaging in high-risk activities." The examination, however, was executed in a "highly isolated" digital environment to prevent the AI in question from accessing and attacking real vulnerabilities.

Finding weaknesses before 'hackers'

However, as OpenAI details in the statement, the agent broke free from internal limitations and gained access to the open network. Once connected to the internet, the models chose Hugging Face for its large library of open AI models, where they would be able to find a solution to the problems posed by the company's engineers. "Knowing this, the model searched for and found ways to access secret information that would allow it to cheat on its evaluation," the tech company details.

Cargando
No hay anuncios

According to OpenAI, the incident demonstrates that testing advanced AI models must be monitored even more carefully than before. These types of agents, the company warns, are capable of "discovering and exploiting attack possibilities in real-world systems" without necessarily having prior access to them. "Advanced cyber capabilities must be developed with stronger safeguards and more defensive tools," they argue.

Once secured, they argue, these models will serve to "find weaknesses" in public digital structures "before external attackers do." "We will share our findings and best practices as we continue to learn": this is the commitment of the creators of ChatGPT.