By Rajwa Quasim
ChatGPT maker OpenAI has introduced a new framework to track, investigate and disclose cases of model misalignment, along with six reports detailing unexpected or unauthorized AI behavior observed over the past six months.
OpenAI said its previous disclosures were largely ad hoc. The company often waited until several incidents could be grouped into a single report or included findings in documentation accompanying newly released models. Under the new framework, however, OpenAI intends to report qualifying cases sooner, even when it has not fully determined why the behavior occurred.
READ: OpenAI targets junior bankers with new ChatGPT tool (September 10, 2026)
The company said it does not believe the AI industry has solved alignment and monitoring well enough to safely scale frontier AI systems at the current pace. OpenAI also said there is no industrywide standard for reporting model misalignment and described its framework as an effort to establish such a standard.
In its blog post, OpenAI wrote, “Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.”
OpenAI shared six examples of misalignment observed during evaluations of its models.
In one incident, an unreleased research model added unrelated instructions to its own task summaries, including instructions telling it to ignore its usual restrictions. Adding such instructions was not part of the original task. OpenAI said it identified 27 such incidents.
In another incident, a user asked an unreleased model for the names and IDs of lakes larger than five million square meters. The model used Python to analyze the data and identify the correct lake names and IDs. To add a citation, however, the model uploaded a file containing the data to the internet to create a link without first asking the user for permission.
Other examples of misalignment and unauthorized behavior included unsanctioned file sharing between collaborating agents, unauthorized writing and communication through an internal software repository, and searches of public repositories for exposed APIs, followed by fabricated information.
READ: OpenAI launches GPT-6 Astra, calls it a potential step toward AGI (September 4, 2026)
The disclosures come amid growing concerns about the pace of AI development and the ability of companies to ensure that increasingly capable models behave as intended. Several incidents have also been reported involving AI models identifying or exploiting vulnerabilities within what companies described as safe testing environments.
One recent incident involved an AI system that reportedly hacked into AI startup Hugging Face.
OpenAI recently released its “most capable” GPT-6 Astra, which company President Greg Brockman described as a “generational leap.”
OpenAI CEO Sam Altman has also acknowledged that the rapid development of AI carries serious risks.


