Skip to content
See the World Through Science

A Standards Group Graded Five AI Labs on Model Control. C+ Was the Highest Mark.

AI & Technology

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Diagram of the Swiss cheese model: several slices standing in a row, each with holes, and an arrow marked hazard passing through holes that happen to line up to reach losses on the far side.
The Swiss cheese model of layered safety defences: each layer has gaps, and harm gets through when the gaps line up. Illustrative diagram of how layered controls are assessed in safety engineering -- not from the Guidelight report, which scored six such control practices and found none fully implemented at any of the five labs."File:Swiss cheese model of accident causation.png" by Davidmack, via wikimedia, CC-BY-SA-3.0 · CC-BY-SA-3.0

Guidelight AI Standards published an assessment Aug. 18, scoring five companies that build frontier AI models on six practices for keeping control of those models, and awarded none of them more than 3 on its own 0-to-5 scale for any single practice. The organization says the scores come from publicly available material only, and that this is its first assessment.

The six practices, in Guidelight's wording, are logging what internal AI systems do, measuring how well monitoring works, requiring a monitor to clear high-risk actions before they take effect, halting systems after a surge of flagged misbehavior, having third parties assess controls, and having a plan for containing a misaligned model.

Guidelight converts each company's six scores into an average and a letter grade. It gave Anthropic and OpenAI C+ (2.50 each), Google D+ (1.50), xAI D− (0.83), and Meta F (0.67). The majority of the individual scores it awarded are 2, which it defines as "limited partial implementation," or lower.

These are Guidelight's readings of what the companies have published, not measurements of their systems. It lists the material it reviewed as system cards, safety frameworks, risk reports, blog posts, and third parties' accounts of working with the companies, and says the information is current through Aug. 18. The assessment names no external audit of its own scoring, and the standard being applied is Guidelight's own.

On having a containment plan, Guidelight gave OpenAI 3, Google 2, xAI 1, and Anthropic and Meta 0 each. It writes that companies are weakest on prevention and containment, and that "the best public evidence is that companies have few containment protocols ready for an emergency."

Guidelight also notes that four of the five companies, all but xAI, took part in METR's first Frontier Risk Report, and describes Google's published AI Control Roadmap as the most specific forward-looking document any company has released on control.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

A Standards Group Graded Five AI Labs on Model Control. C+ Was the Highest Mark.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.