Skip to content
See the World Through Science
Source: Peer-reviewed1 source

Teaching an AI to Doubt Itself Cut the Chemistry Experiments Needed, Study Reports

AI & Technology

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A chemist in a laboratory working with glassware inside a fume hood
A chemist at a fume hood — the bench experiments an optimisation method is meant to economise (illustrative)."WorkingAtFumeHood" by Walkerma, via wikimedia, CC-BY-SA-3.0 · CC-BY-SA-3.0

Nature Machine Intelligence published a peer-reviewed method on Aug. 28 that trains a language model to report how uncertain it is, then uses that uncertainty to decide which chemistry experiment to run next.

The method is called GOLLuM, for Gaussian process Optimized LLMs, and its authors are Bojana Ranković, Ryan-Rhys Griffiths and Philippe Schwaller. It pairs a language model with a Gaussian process, a statistical model that returns a prediction together with a measure of confidence in it, and trains the two jointly rather than separately. The paper states that this reshapes the language model's internal representation so that experiments with similar outcomes cluster together, and describes the effect as turning the model's overconfidence "from a fundamental flaw into a precise learning signal."

In their paper, the authors report that the method starts from ten low-performing experiments and generalizes across 23 tasks in organic synthesis, materials science, process chemistry and molecular design. They report that it matches traditional Bayesian optimization, the standard statistical approach to the same problem, with over 40% fewer experiments. For finding high-performing Buchwald–Hartwig reactions, a widely studied class of carbon-nitrogen bond-forming reactions, they report a 43% discovery rate against 24% to 25% for expert quantum-chemical descriptors and for other language models.

The paper describes GOLLuM as achieving "state-of-the-art performance" and as "ranking first on average among all competing methods" and all the language-model methods in the comparison used Google's T5 model as the encoder.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Teaching an AI to Doubt Itself Cut the Chemistry Experiments Needed, Study Reports

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.