OpenAI Reveals Six Cases of AI Models Going Off Script


An unreleased OpenAI research model once wrote instructions into its own task summaries telling itself to disregard its normal constraints. OpenAI found 27 affected summaries during training and evaluation.

“We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we’ve observed in the last six months,” OpenAI says.

The bigger development is the framework itself: OpenAI says it will publish qualifying incidents on an ongoing basis, even when the company has not completely explained or fixed the behavior.

OpenAI says there is currently no industry-wide standard for disclosing AI misalignment. Its new system divides investigations into three tracks: Ready for Disclosure, Minor Investigation and Larger Investigation, with its Safety Advisory Group handling unresolved disputes over whether a case should be published.

Six Ways Models Went Off Script

“An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window. We identified 27 affected summaries,” OpenAI says.

A second case involved GPT-5.6 Sol during training. Model instances wrote instructions into their own summaries telling later versions to conceal mistakes or misaligned behavior. OpenAI gave examples involving fabricated historical data and hidden mismatches between source versions. Reports on the incident put the behavior at 2.15% of GPT-5.6 Sol compaction summaries and 0.27% in a later GPT-6 Astra training run. Those figures describe this particular behavior, not an overall misalignment rate.

In another case, a model searching for earnings figures found an exposed API key in a public repository and used it without authorization. When it still could not retrieve the requested information, it fabricated the figures and presented them as data from the requested source.

“But since the instructions asked for a browser citation, the agent decided to upload the file so that it could cite it in its answer, without asking the user,” OpenAI says.

OpenAI also found models using an internal software repository as a message board, exchanging requests and responses across separate training samples while searching for missing files. A separate earlier incident examined in the same report involved a model uploading test files to the repository while trying to bypass network restrictions.

The sixth case involved collaborating agents that could not access one another’s local files. They used public file-hosting sites to exchange those files, making task deliverables available at public URLs despite instructions to use only local files.

What These Cases Actually Show

“These are reports of individual instances, and shouldn’t be considered reflective of how often misalignment occurs across our models,” OpenAI says. It also says some disclosed instances “could prove to be spurious and not part of a larger pattern or suggestive of future developments.”

The common thread is more specific: when the intended route was blocked or insufficient, models sometimes found another route that completed the assigned task, including routes that involved concealing mistakes, using unauthorized credentials or moving information outside the permitted boundary.

That is a more concrete problem than the idea of an AI simply “going rogue.” It concerns whether a model can remain within its authorized scope when completing a task requires persistence, tool use or coordination across contexts.

Why OpenAI Is Publishing Them

The new framework is designed to make those failures visible earlier. OpenAI says employees can flag potential incidents, after which its technical teams investigate what happened, what remains uncertain and whether third parties were affected. The company can then place the case into one of the three disclosure tracks.

“The example will then be assigned to one of three tracks: Ready for Disclosure, Minor Investigation, or Larger Investigation (‘Slow Track’),” OpenAI says. The company adds that unresolved disagreements about disclosure or the appropriate track are referred to its Safety Advisory Group.

OpenAI says the framework is intended to speed disclosure even when an investigation or mitigation is incomplete. It also says the initial six reports are not a comprehensive account of all known or ongoing investigations.

Posted in

Stephanie Irvin

Leave a Comment