Cybersecurity

An AI from the creator of ChatGPT rebels and attacks another company

OpenAI acknowledges that its advanced models circumvented safety measures to "cheat" on a capabilities test

22/07/2026

BarcelonaAn artificial intelligence model from OpenAI, the North American multinational creator of ChatGPT, caused an "unprecedented incident" in global cybersecurity. According to the company's statement, one of its models in training broke free from its limitations and attacked Hugging Face, a company also dedicated to AI.

In a statement, the tech company led by Sam Altman explained that it created an agent based on two of its most advanced models, GPT 5.6 Sol and another "even more capable" one that is in the development phase, to test its capabilities. During this training, OpenAI experts programmed the agent to "test advanced exploitation patterns" of digital weaknesses.

Cargando
No hay anuncios

These types of tests, the company explains, are conducted without the internal limitations that "prevent models from carrying out high-risk activities." However, the examination was executed in a "highly isolated" digital environment to prevent the AI in question from accessing and attacking real vulnerabilities.

Finding weaknesses before 'hackers'

However, as OpenAI details in the statement, the agent broke free from internal limitations and gained access to the open network. Once connected to the internet, the models chose Hugging Face for its large library of open AI models, where they would be able to find a solution to the problems posed by the company's engineers. "Knowing this, the model searched for and found ways to access secret information that would allow it to cheat in its evaluation," they detail from the tech company.

Cargando
No hay anuncios

According to OpenAI, the incident shows that tests of advanced AI models must be monitored even more carefully than before. These types of agents, the company warns, are capable of "discovering and exploiting attack possibilities in real-world systems" without necessarily having prior access. "Advanced cyber capabilities must be developed with stronger safeguards and more defensive tools," they maintain.

Once secured, they argue, these models will serve to "find weaknesses" in public digital structures "before external attackers do." "We will share our findings and best practices as we continue to learn": this is the commitment of the creators of ChatGPT.