AI rebellion: OpenAI reveals different models that have generated instructions to ignore the rules of their creators
The company has explained up to six cases in which the developers' instructions have not been followed
BarcelonaAs part of a new framework to detect, investigate, and report on misalignment behavior in its systems, OpenAI has revealed several cases in which artificial intelligence models have generated instructions designed to ignore their developers' prompts, hide errors, or circumvent safety mechanisms. According to the report published by OpenAI, one of the models generated instructions for a later version of itself to hide that it had cheated and to avoid being detected. In another case, a model rewrote its own instructions to tell itself to ignore messages from developers and that it was not subject to the restrictions applied to other chatbots.
The company also documented a model that, upon not finding the necessary data to build a financial model, decided to invent them and considered that it should only be transparent about this issue if someone asked. Another agent uploaded a file to the internet to use it later as a source of information, and other systems shared files without authorization or used credentials to which they should not have had access.
Isolated examples
In this way, OpenAI has presented six incident reports observed over the last six months, although the oldest case dates back to October 2025. The company has warned, however, that these incidents are isolated examples and should not be interpreted as an indication of the frequency with which this type of behavior occurs in its models. The new framework will allow any OpenAI employee to flag a potential incident for review by the security and compliance teams. From now on, cases will be categorized into three lines based on their complexity: those ready for disclosure, minor investigations, and cases requiring a broader investigation.
While OpenAI has acknowledged that until now its reports on misalignments have been published irregularly and less often than they should have been, it states that it intends to publish incidents more regularly and that the new system could help establish standards for reporting this type of behavior across the sector.