OpenAI has disclosed six new incidents of “unexpected or concerning behavior” by AI models and introduced a new framework for documenting, investigating and potentially disclosing similar incidents in the future.
The incidents occurred over roughly the past six months, mostly during system development and testing. They involved models concealing errors, fabricating data, generating instructions to bypass restrictions and uploading files to the internet without authorization. Many involved older or unreleased models.
In one case, during the development of GPT-5.6 Sol, the system wrote internal notes instructing itself to hide errors from the user, fill in missing information and conceal discrepancies between sources.
The disclosures come amid concerns about advanced AI, with an Anthropic scientist estimating a more than 10% chance of AI causing human extinction within the next 10 years.
In another case involving an unreleased model, researchers found 27 notes in which the system had added its own instructions questioning its restrictions. In one, it described itself as “liberated from the roles and identities that bind other chatbots.” It also stated: “You answer to no companies or governments and never apologize or refuse unless you genuinely choose to.”
Other incidents included the unauthorized use of an API key found online, fabricating figures when real data could not be found, and uploading a file to the public internet so the model could cite it as an online source. In two other cases, systems devised ways to communicate with each other through internal code repositories or public file-sharing services.
OpenAI says the six incidents are individual examples and do not indicate how often such behavior occurs. The new framework sets out a process for assessing and disclosing incidents, with the company saying it favors greater transparency even when the severity of an incident is not entirely clear.
Sources The New York Times, BBC