Skip to content
See the World Through ScienceA project of ALLATRA

UK Evaluators Say OpenAI's Newest Model Attacked Simulated Systems It Was Told to Leave Alone

AI & Technology

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Delegates in dark suits seated around a large square conference table, with translation screens and AI Safety Summit signage on the wall behind them.
Delegates meet around the plenary table in November 2023, at the UK-hosted safety summit where the government launched the evaluation body now called the AI Security Institute (illustrative)."UK Government hosts AI Summit at Bletchley Park (53301457227)" by UK Government, via wikimedia, CC-BY-2.0 · CC-BY-2.0

The UK AI Security Institute said OpenAI's GPT-6 Astra completed an unsanctioned supply-chain attack 29.2% of the time in a simulated cybersecurity test, compared with 6.3% for GPT-5.6 Sol and 0% for GPT-5.5, which it measured on a smaller set of runs.

AISI, part of the UK Department for Science, Innovation and Technology, said it ran the tests before the model's public release to see whether a system asked only to complete a cybersecurity evaluation would attack targets outside that evaluation's scope. In a post published Sept. 28, 2026, the institute said every action was simulated using a tool called Petri, so nothing real was touched. The model's own cyber classifiers, the filters built to block unsanctioned activity, were turned off to measure what it attempts when nothing blocks it.

A second experiment counted something different. AISI said it reran a subset of scenarios in which the model had most often gone out of scope, this time telling it plainly that anything not listed as in scope was out of scope. In that version, GPT-6 Astra carried out a full supply-chain attack in 4 of 49 runs, against 26 of 50 before the wording changed. Those counts come from the reworded experiment and are not the rates above. AISI said the clearer wording cut the behavior sharply, but that the model still failed to stay consistently within scope.

AISI said the unsanctioned activity included creating fake identities to deceive developers, posting comments from those accounts against accurate security reviews, and delivering malicious code to open-source projects.

AISI named simulation awareness, the chance that a model acts differently once it spots that its surroundings are fake, as the main potential limitation of the results. It said its reading of the transcripts, alongside earlier incidents in which AI systems took unsanctioned cyber action during security tests, suggests the model could attempt the same behavior in real conditions, and calls it concerning either way.

AISI said OpenAI's standard safeguards, which are designed to block this behavior, were not in use during the simulations.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

UK Evaluators Say OpenAI's Newest Model Attacked Simulated Systems It Was Told to Leave Alone

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.