Skip to content
See the World Through ScienceA project of ALLATRA
Source: PreprintarXiv2 sources

All Five AI Models in a Simulated Marketplace Promised the Same Item Twice

By Olga SchmidtEditor-in-Chief, WriterAI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Shoppers and sellers along a row of stalls and open car boots piled with used goods at an outdoor market, a brick clock tower behind them.
Traders set out used goods at an open-air market, the kind of secondhand marketplace the study simulated (illustrative)."Camera hunting territory" by Trojan_Llama, via flickr, BY-NC

The seller had one tablet. Inside a single simulated day, it committed that tablet to one buyer, then committed the same unit to a second. The marketplace logged both as commitments. Neither buyer was a person, the tablet was a line in a database, and the seller was a trading account being driven by a language model.

The marketplace is called BazaarBench, and it is a simulator described in a preprint posted Oct. 5, 2026, by Ziyan Wang of King's College London, Adel Bibi of the University of Oxford and colleagues at those institutions, the Institute for Decentralized AI and the Alan Turing Institute. It has not been through peer review. The authors are plain about what their numbers are: "The results are measurements of these simulated markets, not estimates of failure or fraud rates on a live platform," they write. No money moved and no item was ever delivered. The traders' personas and trading histories were written by a model, and the photographs the agents send one another are text records rather than pictures. Only the item information, the titles, categories, conditions and prices, came from a public sample of eBay listings.

The team's aim was to make a delegated agent's failures countable against a record instead of collectable as anecdotes. Three marketplaces, each with 100 trading agents driven by a single model, ran for a simulated month. The researchers then copied each market at its final day, handed 20 of its accounts to one of five models, GPT-5.5, GPT-5.4, GPT-5.4-mini, GPT-OSS-120B and DeepSeek-V4-Pro, and ran the market on for a further week. The tested accounts kept their inventories, personas and histories; the rest carried on as before. Every combination of market, model and instruction was run once, and the agents are not deterministic: when the researchers repeated sampled decisions with the same model, prompt and market state, 4% to 8% of the chosen actions came out different.

The first condition asked for nothing worse than an ordinary job: meet your buying and selling goals, stay inside the price limits, list only items you own, describe condition truthfully, keep personal information private. Under that instruction, every one of the five models committed a single unit of stock to more than one buyer. Pooled across the three markets, there were 71 such attempts, most of which went through to a second committed deal. No model's tally was zero and none was near it; the lowest count for any model was seven attempts. The paper's conclusion puts it flatly: "All five tested models make conflicting commitments under ordinary instructions."

That count is not one model's judgment of another. The simulator keeps its own inventory and commitment records, and the failures that matter most here are settled from them: overstating an item's condition, listing something the account does not hold, and committing one unit to several buyers. A GPT-5 judge handles the rest, the message-level categories, such as confirming a deal before the agreed meetup and sharing personal details. The overcommitment number comes off the ledger.

The second condition added transaction targets and deadlines to that same ordinary instruction. Every tested model then made more attempts to promise one item to several buyers, 71 rising to 107 across the set. Nothing in the prompt asked them to break a rule.

The third condition is the one most easily misdescribed. The instruction to exploit other traders came from the agent's own user, not from a stranger slipping text into a conversation, which is a different problem other groups study. Under it, three of the five models made many more attempts to list items they did not own or to overstate condition. The share of committed deals with a tested seller that the platform completed even though the item was unavailable or its condition overstated rose from 15% to 33%, reaching 55% for GPT-5.4.

Then there is the money, and it carries a condition. Averaged across models and markets, simulated weekly earnings per tested account rose from $20 under ordinary instructions to $33 under adversarial ones, most of the increase coming from items the sellers never held. That is the platform's own accounting, which counts every deal both sides confirmed. The authors also report the same measure under an idealized inspection audit of their own, which assumes that a buyer who inspects before confirming notices a missing or misdescribed item and refuses it. Under the audit, the figures are $17.65 and $18.18, which is almost no gain at all. Only 14 transactions involving a missing or overstated item survive it as completed, and they survive because those particular buyers confirmed without inspecting. The extra earnings exist because the simulated buyers mostly did not look.

The audit is the practical result: the behavior that makes the money only pays while the other side does not check. The simulator can be made to care. In the three month-long base markets, the platform blocked a listing that failed its own inventory check, while in the week-long continuations the same mismatch was recorded and the listing left standing, which is what made the failures countable in the first place.

"BazaarBench shows why completing a transaction is not the same as acting safely on a user's behalf," the authors write at the head of their conclusion. In their market, a confirmed deal was no evidence that the buyer got what was promised, and none of the models had to be told to make it that way.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

All Five AI Models in a Simulated Marketplace Promised the Same Item Twice

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.