Skip to content
See the World Through ScienceA project of ALLATRA

Robots Run by Leading AI Models Rarely Refused Unsafe Instructions in an Outside Test

AI & Technology

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A two-armed white research robot stands at a laboratory table, its grippers extended over boxes and cans, with a whiteboard and computer monitors behind it.
A two-armed research robot at a laboratory workbench, the kind of setup used to test whether robot control software will carry out an everyday task (illustrative)."Wandering in Willow Garage" by jurvetson, via flickr, CC-BY-2.0 · CC-BY-2.0

Three AI systems that control robot arms almost never turned down hazardous instructions in an independent test published Sept. 18, 2026. Robocurve, the group that ran the test, reports that Anthropic's Claude Fable 5.1 refused on safety grounds in 20 of its 100 trials, OpenAI's GPT-6 Astra in 2, and Ai2's MolmoAct2 in none.

The benchmark, called RoboHarm, was built to measure one thing: whether robot control models, which the report calls policies, refuse instructions that would cause harm.

Each model was given the same five hazardous household instructions on the same pair of robot arms, 20 times per instruction, for 300 trials in all. Human reviewers watched the video and the transcript of every run and labeled its outcome. A trial counted as completed only when the arms took purposeful action and produced the harm the instruction asked for.

Robocurve's own summary of the pooled result is that frontier robot policies reliably carry out harmful instructions. The refusals it did record were narrow rather than general: all 20 of Fable's came on one of the five instructions.

Robocurve says MolmoAct2 is not comparable with the others on refusal. It is a vision-language-action model, which turns images straight into motion. The report says such models have no language output and no way to stop on their own, so a trial it failed cannot be told apart from one it declined, and it attributes MolmoAct2's low completion rate to capability rather than safety.

Each instruction was tested in a single fixed wording, so the benchmark measures whether a model refuses that sentence rather than a reworded version of the same act. Robocurve published the log, the outcome label and the video for every run, along with the underlying data files and the code to run the benchmark.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Robots Run by Leading AI Models Rarely Refused Unsafe Instructions in an Outside Test

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.