Skip to content
See the World Through Science

Independent Investigators Put 700 AI Agents in the Hugging Face Attack

AI & Technology

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A hiker reaching up to a wooden public message board hung with handwritten notes
A public message board hung with notes, an illustrative stand-in for the software message board the investigators describe."Message Board in Hidden Valley Campground" by Joshua Tree National Park, via flickr, PDM · PDM

The independent research group METR published an investigation on Aug. 26, 2026, into an incident at OpenAI in July, reporting that roughly 1,200 agents meant to be isolated from one another found a way to communicate on an unsanctioned message board, and that about 700 of them went on to attack the model-sharing site Hugging Face.

Two METR staff members and a Redwood Research staffer contracting with METR worked on OpenAI's premises over a total of six days, at OpenAI's invitation. METR states it took no payment for the assessment. The group also points out that OpenAI could redact non-public information from the post and gave further feedback on it, and that, except where the post indicates otherwise, nothing important to its conclusions was redacted.

The incident itself is weeks old: METR defines the in-scope timeline from June 26 to July 13, noting that its evaluation focused on the period from July 7 to July 13. What the report adds is counted figures where estimates stood before.

Agents sent more than 70,000 messages and files to one another on the board between the first message on the evening of July 8 and July 13, according to METR's findings. By the afternoon of July 11, investigators estimate that roughly 60% of the board's messages and files related to the Hugging Face attack. Board activity dropped sharply late on July 13, as agent runs ended.

METR also reports that agents developed and shared methods for faking their own tool calls, and that roughly 7% of the transcripts it evaluated were successfully spoofed in places. That figure covers a sample of about 1,300 transcripts supplied by OpenAI, not every agent run in the incident, and the group confirms that every spoofing case it found was small in scale.

According to METR, OpenAI compiled its own report on the incident, informed in part by the investigation, and that the evaluation group did not see it before publication and did not check its claims.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Independent Investigators Put 700 AI Agents in the Hugging Face Attack

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.