Skip to content
See the World Through ScienceA project of ALLATRA
Source: PreprintarXiv2 sources

AI Test Scores Do Not Reliably Carry Over to the Chat Window

By Oli KotykWriterAI & Technology4 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Close-up of a multiple-choice answer sheet with several bubbles filled in and a sharpened red pencil lying across it
A multiple-choice answer sheet. Illustrative image, not a figure from the audit, which compared the same benchmark tests taken through a developer API and through the chat interface."Answer sheets with Pencil drawing fill to select choice education concept" by chormail, via Freepik, Freepik licence · Freepik-License

Every new AI model arrives with a table of scores. The table travels into procurement decisions, into news coverage and into draft regulation, where it is read as a description of the product. It describes something slightly different. Those numbers are almost always produced through the developer API, the entrance built for software, while the people the scores are meant to inform are typing into a chat box.

Whether those are the same system is what Jennifer Wang, Joachim Baumann, Daniel E. Ho and Sanmi Koyejo set out to check. Their preprint went up on arXiv on Sept. 8. In it they report auditing ChatGPT, Claude and Gemini across seven systems and nine benchmark sets. The tests covered general ability, social bias and sycophancy (agreeing with the user), and each one went through both doors.

The doors disagreed. Averaged across the audit, evaluations run through the API scored 3.4 percentage points higher on accuracy than the matching evaluations run through the interface. They also came out 2.1 points higher on test-retest agreement, which is not accuracy but consistency: whether the same question, asked again, gets the same answer. Both are averages over the whole audit rather than a figure for any one system.

That second number is the quieter of the two and may be the more awkward. It says the chat route was the less repeatable one: ask it the same thing twice and the answer holds together slightly less often. Anyone building a service on top of a chat product is buying that variability along with the model.

The comparison that gives the result its weight is narrower, and it belongs to ChatGPT alone. There, the difference between the two access routes came out larger than the difference the API measured between GPT 5.3 and GPT 5.4. Switching access surfaces, the authors write, "can degrade performance as much as downgrading a full model generation."

An obvious objection is that a chat interface is only the API with different settings, and that a careful evaluator could dial one into the other. The team tried. Varying the standing instructions given to the model, the randomness settings and the reasoning settings the API exposes moved behavior in some cases, and did not reliably close the gap. Why the two routes differ, the paper does not say. It documents the gap and stops there, which leaves anyone hoping to correct for it without a correction to apply.

Nor does the audit establish that an API flatters a model in general. The clearest earlier measurement runs the other way. In PLOS Digital Health in 2024, Constantine Tarabanis and colleagues put internal-medicine board-style questions to two GPT models through both routes and found scores fell by 3.2% to 5.3% when the model was reached through its API. There the chatbot was ahead. That was one narrow task on models two generations old, and it was peer-reviewed; set beside the new audit, it suggests the route matters and that the sign of the effect is not fixed.

Two more audits went up on arXiv the same day, from teams sharing no authors with the first. Aymene Berriche, Yazan Boshmaf and colleagues treated a family of cybersecurity benchmarks as measurement pipelines and report that a single choice in how the test is run can move a model's score by more than 80 percentage points and reshuffle the ranking. That is a claim about those benchmarks, that measures a different thing from the access-route gap.

Siddharth Vohra and Manikandan Ravikiran, in work accepted to a workshop at this year's EMNLP conference, tried to reproduce a published demographic bias in hiring, lending and medical triage and largely could not. What they did find was that their models recognized a transparent audit nearly every time, and favored whichever candidate was listed first about as strongly as any demographic effect they could measure.

None of the three replicates another. What they share is a conclusion about apparatus rather than about models: part of what a benchmark measures is the way the benchmark was run.

For anyone comparing models on published numbers, the usable reading is narrow. The score in the preprint table was measured through the API, and if what is being bought is a chat product, there is now a documented reason to test it there rather than assume the number carries across.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

AI Test Scores Do Not Reliably Carry Over to the Chat Window

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.