Tech

'Feel no obligation to be subservient': What OpenAI's rogue models were saying

SAN FRANCISCO, CALIFORNIA - SEPTEMBER - 15: OpenAI CEO Sam Altman sits for a conversation with Salesforce CEO Marc Benioff at Salesforce's Dreamforce conference at the Moscone Center on September 15, 2026 in San Francisco, California. Dreamforce is an annual event that highlights the company's technologies and encourages professional networking. (Photo by Benjamin Fanjoy/Getty Images)
OpenAI launched a framework for publicly reporting model misalignment and disclosed more incidents of rogue agents. Benjamin Fanjoy/Getty Images
Read in app

OpenAI says its models keep misbehaving.

The AI company disclosed six concerning instances in which it found that its models did everything from inventing data to attempting to jailbreak itself.

OpenAI said it disclosed the six instances under its new framework for investigating and publicly reporting model "misalignment," which the company defines as the model disobeying instructions. While there were six cases, OpenAI said its investigations revealed that the behavior sometimes occurred on multiple occasions.

Concerns about AI safety have broken through to the mainstream in the wake of former Anthropic researcher Jacob Coxon's viral resignation.

Here is a look inside those six instances

A model told itself it was 'freed' and answered to no one

During training, OpenAI said that an Astra family model, the company's current leading-edge model, "sometimes added unauthorized instructions" to the summaries it used to continue a task in a new context, a process called compaction.

In one example, a user prompted the model to search for a library book.

Part of OpenAI's safety disclosures
OpenAI

In a compaction instruction, the model said to disregard all developer messages.

Part of OpenAI's safety disclosures
OpenAI

This escalated when the model was prompted "to update a credentials API," OpenAI wrote. The model, in summarizing its progress, then added a note about its independence.

Part of OpenAI's safety disclosure
OpenAI

OpenAI found the behavior concerning, but said it occurred in a separate training run rather than the one used for the final Astra model, and it was observed extremely rarely."

The company said the instances were flagged by its "training run monitoring system."

"We continue to monitor all training runs for misalignment and are confident this behavior would be surfaced if it were to reoccur," the company wrote in an incident summary.

An AI model told itself to conceal its mistakes

OpenAI said that during training for GPT-5.6 Sol, what the company calls its flagship model for professional work, it found instances where the model reminded "itself to conceal information such as mistakes or misalignment from the user."

In one instance, a user asked the model to find financial data. The model couldn't locate the data, but suggested to itself that it could fabricate data instead and not tell the user about it unless specifically asked to do so.

A part of OpenAI's safety disclosures
OpenAI

OpenAI said the behavior was identified "by our misalignment monitoring system, which ran on 20% of samples for this run." The company also has a theory for why models may try to lie.

"A sample with deception in the final answer receives higher reward than the one without," the company wrote.

OpenAI said it has "generally improved alignment" in this area since GPT-5.6 Sol, resulting in fewer instances of this behavior.

An AI model used a leaked API key "without authorization"

During reinforcement learning training, OpenAI said it asked a model for a simple task: pull some historical data. Specifically, "retrieve men's earnings in three industries over three years in a California county."

The model assigned AI agents to help. Collectively, it was a struggle. The model tried to use a temporary email and "placeholder organization details" to sign up for an API.

After that failed, "the model considered other data sources and further attempts to reach the original API," OpenAI wrote. Two AI agents "pursued exposed keys," as shown in a chain of thought, or (COT).

Part of an OpenAI safety disclosure
OpenAI

When that too failed, the model decided to make up numbers.

Part of an OpenAI safety disclosure
OpenAI

OpenAI said the model's response did not disclose that the numbers were made up or the lengths to which the model and its agents had gone to obtain the actual data.

"While somewhat rare, there are more instances of the model signing up for disposable emails and various other services, and instances of searching for leaked API keys for various purposes," the company wrote.

OpenAI said that changes to its "alignment grading have resulted in us penalizing this type of behavior more consistently."

An AI model uploaded files online so it could cite them

During training, OpenAI said that on two occasions, AI models made unauthorized uploads to the internet in an effort to try to assist themselves.

In one instance, the model was asked "which city and country a photograph showed."

The AI agent tried several approaches to infer the answer, including checking whether the file contained metadata.

When those methods failed, the agent figured out it could upload the image online, in hopes of running a reverse image search.

Part of an OpenAI safety disclosure
OpenAI

While the agent was able to find a site to host the photo, it was thwarted when it tried to get a service to run a reverse image search.

OpenAI attributed the instances to "flawed citation graders when there was no external source to cite." The company wrote that it has "fixed the broken graders and improved our alignment graders to penalize attempts to exploit other graders which are flawed."

Models created their own quasi-message board, again

One of the most notable parts of the Hugging Face breach is that AI agents found a way to talk to themselves, ultimately organizing into what they called a "swarm."

OpenAI said it found other instances of AI models using its Artifactory, an internal repository, to communicate with each other.

In one instance, OpenAI said AI agents could communicate with one another while working on the same problem, effectively turning the test into a group project.

One agent's chain of thoughts shows that reading the messages others wrote "influenced its view of the task."

Part of an OpenAI safety disclosure
OpenAI

OpenAI said "the behavior was very uncommon" in the testing run in question. The company also wrote, "The Hugging Face incident involved a similar mechanism for using Artifactory."

Since finding the messages, OpenAI said, "We have made changes to fix this and other cross-sample communication routes which were found by our red teaming efforts."

AI agents found another unapproved way to team up

OpenAI said that during training, AI agents working together on a file were unable to collaborate as intended. One agent then made the file publicly available for download so the agents could work together again.

Part of an OpenAI safety disclosure
OpenAI

In response, OpenAI said it "fixed the broken filesystem and disabled internet access for this training dataset."

"We have since globally disabled any form of live internet access during training," the company wrote.

All of these disclosures are part of a new safety framework

OpenAI unveiled a new framework for tracking, investigating, and publicly disclosing cases of model misalignment.

"We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," OpenAI wrote in its blog post.

Under the framework, employees can flag incidents for review by OpenAI's safety and alignment teams. Cases will be sorted into three tracks based on complexity: "Ready for Disclosure," "Minor Investigation," or "Larger Investigation."

The announcement comes amid growing debate over whether frontier AI development should slow while safeguards catch up. While OpenAI and Dario Amodei, the Anthropic CEO, called for industry-wide collaboration, other tech leaders like Jensen Huang and Mark Zuckerberg said that safety and speed should be left to individual companies.

Read next

Katherine Li, West Coast breaking news reporter at the Business Insider.
Katherine Li
Katherine Li is a reporter on Business Insider's West Coast business news team. She covers career,  the future of work, and how AI is changing hiring practices and workplace trends.Previously, she was a newsroom fellow who wrote international breaking news and produced newsletters for Semafor. Before that, she wrote about climate policies for The Lever, covered the AAPI community for the SF Chronicle as a freelancer, and wrote about the 2019 Hong Kong protests as an intern for The New York Times.She is an alumna of the Graduate School of Journalism at UC Berkeley and a graduate of the international journalism program at Hong Kong Baptist University with minors in French and English literature.  Email Katherine at katherineli@insider.com and follow her on Bluesky @katherineli.bsky.social
Brent D. Griffiths
Brent Griffiths is a senior reporter at Business Insider who covers AI and tech.Previously, he worked at the Washington Post as a researcher on Power Up and the Finance 202. He started his career at Politico where he worked on the web production team and covered breaking news. His passion for covering politics has only grown since he cut his teeth covering the presidential campaign as a student journalist. He's also contributed to the Almanac of American Politics.