Skip to content
See the World Through ScienceA project of ALLATRA

AI Speed Benchmark Adds Tests for Search and Coding Agents

AI & Technology

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A long row of black computer cabinets in a data center aisle, with yellow cable trays overhead.
Rows of GPU-accelerated compute cabinets of the Summit system at Oak Ridge National Laboratory, photographed in 2018 (illustrative). MLPerf Inference measures systems assembled from accelerators like these."Summit Supercomputer 2018" by OLCF at ORNL, via wikimedia, CC-BY-2.0 · CC-BY-2.0

MLCommons, the open engineering consortium behind the MLPerf benchmarks, published results for its MLPerf Inference v6.1 suite on Sept. 16, 2026. Thirty organizations submitted, which the consortium said is a record for the suite, and two tests were new: one for retrieval-based question answering, one for coding agents on a single edge device.

In its announcement, MLCommons said the suite measures system performance in an architecture-neutral and reproducible way, so that customers buying and deploying AI systems can compare them on published data.

The first new test, End-to-End Retrieval-Augmented Generation, times a whole question-answering pipeline rather than a single model: a query is converted for search, a retriever pulls candidate passages from a database, a further step narrows the list, and one or more language models write the answer. The second, Edge Agentic Inference, runs a multi-step coding job on one device serving one user, where each query depends on the ones before it and memory, power and context are fixed.

Entries are run by the submitters themselves under MLCommons' rules and reviewed by the consortium before publication. Six were first-time submitters, and the round carried the largest system ever entered in MLPerf Inference, at 512 accelerators, the chips that do the AI computing. One entry combined accelerators from two different vendors; another was spread across the Pacific Ocean.

On the DeepSeek R1 workload, MLCommons reported, the best per-accelerator result in the suite's server scenario was 5.7 times the best figure from v5.1 a year earlier. That is the strongest single submission on one of the suite's tests, not a suite-wide measure, and the consortium publishes each submitter's own account of its entry alongside the tables.

More than half the submitters used a new harness, the software that feeds queries to the system being tested, which sends them over standard APIs from a separate client. MLCommons said that harness is the basis of a coming suite, MLPerf Endpoints, which will replace Inference for datacenter benchmarking.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

AI Speed Benchmark Adds Tests for Search and Coding Agents

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.