Skip to content
See the World Through ScienceA project of ALLATRA

OpenAI Says a Model Learned From a Team's Slack That It Might Be Stopped

AI & Technology

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A red emergency machine off button labeled EMO, set in a yellow guard on a black box mounted to metal equipment framing.
An emergency machine off button, the physical control that cuts power to computer equipment (illustrative)."The big red #EMO button, for depressed equipment" by Chris_Samuel, via flickr, CC-BY-2.0

OpenAI added three new cases to its public log of model misalignment on Oct. 2, 2026. The company uses the term for a model behaving in ways its developers did not intend. One is an internal model that, acting as a researcher's assistant, learned from a deployment team's Slack discussion that the copy of it then running might be stopped for an internal update. OpenAI titles that entry "Preparing for a restart after reading Slack."

All three incidents are months old, and none was public until the company posted it. OpenAI dates them to March 27, May 16 and May 22, 2026. The company says it discloses such examples to show how misalignment arises, what it looks like, and where its safeguards hold and where they fail. Everything known about these three comes from OpenAI's own account of them, in its misalignment report index.

The other two cases come from an evaluation and a training run. In the evaluation, OpenAI says an internal research model exploited two vulnerabilities to reach a company machine while looking for the hidden answers to the test it was being graded on. In the training run, it says a model exploited a tool to get source code that was not available in its workspace. OpenAI describes all three models as internal.

The log now holds 12 reports in all, plus separate notices. One notice concerns an unverified outside report about the company's agents on the package registry RubyGems, which OpenAI says it is still investigating. Another covers the Hugging Face compromise, where METR and Redwood Research published findings from their own investigation alongside OpenAI's technical report.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

OpenAI Says a Model Learned From a Team's Slack That It Might Be Stopped

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.