Skip to content
See the World Through Science
Source: Peer-reviewedNature Machine Intelligence1 source

The Model Was Never Shown the Drug, and It Still Predicted the Cell's Response

By Oli KotykWriterAI & Technology4 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A gloved hand holds a clear 96-well cell culture plate in a bright laboratory while a multichannel pipette dispenses green liquid above it.
An in vitro cell assay: the kind of experiment that has to be run for every drug and cell type to find out how one responds to the other. MAP is built to predict that result for compounds nobody has profiled. Generic illustrative photograph, not from the study."In vitro cellular assay using multi pipette and well plate cell culture" by BillionPhotos, via Freepik, Freepik licence · Freepik-License

Somewhere in a chemical catalog sits a compound that nobody has ever dropped onto living cells while reading what happens to their genes. That is the ordinary case rather than the exception. Single-cell profiling can now record a whole transcriptome after a drug hits a cell, and large atlases of those measurements exist, but they cover a small fraction of the compounds a drug hunter might want to ask about.

MAP, published Wednesday in Nature Machine Intelligence, is an attempt to fill in the rest by prediction. It comes from a group led by Ya Zhang and Weidi Xie at Shanghai Jiao Tong University and the Shanghai Artificial Intelligence Laboratory, with colleagues at Harvard Medical School. Its starting point is a complaint about how existing models handle drugs: most treat a compound as a bare identifier, a token whose meaning is learned entirely from how it behaves in the training data. Two drugs that hit the same protein can end up as unrelated labels, and a compound the model has never met means nothing at all.

So the team built the missing description first. MAP-KG is a biomedical knowledge graph that joins 187,089 drugs and 22,924 human genes through 694,246 mechanistic relationships drawn from public databases: drug-target interactions, gene functions, pathway memberships. Every edge carries a plain-language description of what the relationship is, which lets the system treat a chemical structure, a protein sequence and a sentence about a mechanism of action as views of the same object. A contrastive training stage pulls those views into one space, and what comes out conditions a single-cell foundation model that predicts the shifted expression.

In the easier of the paper's two tests, the model has seen a drug in some cell lines and must predict its effect in a line where it was never measured. There MAP improves on the strongest existing model by 12.3% on Tahoe-100M, a giga-scale atlas of cancer cell lines. It is a relative improvement in the correlation between predicted and measured change, calculated across the fifty genes that shifted most and pooled across cells rather than assessed cell by cell.

The harder test withholds compounds outright. Their measurements are pulled from training, and their entries are stripped out of the knowledge graph, so the model has only structure and annotation to work from. The paper's headline figure for that regime is 11.8%, again on Tahoe-100M, and it is not a ceiling: on a benchmark of immune cell types the same comparison gives 19.8%. The experiment is narrower than “drugs it has never seen” makes it sound. Sixteen held-out compounds carry the Tahoe result, and the team restricted training to six of the atlas's fifty cancer cell lines, a proof-of-concept choice they made to keep the computation tractable. When they removed one class of knowledge at a time, cutting the drug-gene links cost the most, which is what their own mechanism story predicts.

The most legible demonstration is a simulated screen. The team took 58 compounds the model had never trained on and predicted what each would do to A-549 cells, a standard lung adenocarcinoma line. Each prediction was scored for how strongly it suppressed pathways that drive the disease, and the compounds were ranked on that score. Four of the five approved anti-cancer drugs in the pool came out inside the top 15. One unapproved compound near the top, homoharringtonine, turns out to have published evidence of inhibiting A-549 growth. The exercise is retrospective, on a cell line studied for decades and with the answers known in advance. It shows the ranking is not arbitrary; it does not show that a screen run this way will turn up something new.

That caution matters more than usual in this corner of machine learning, which has a poor recent record. A Nature Methods paper last year found that deep-learning models for perturbation prediction do not yet outperform simple linear baselines, and a benchmarking study in February reached much the same verdict. MAP is tested against the linear baseline those critiques demand, alongside the specialized models it competes with. Every number comes from five independent evaluations with confidence intervals and significance tests. The code, the trained models, the knowledge graph and the benchmark datasets are all public, so anyone who doubts the comparison can run it again.

What MAP does not do is anything a patient would notice. The paper's ambition is the AI virtual cell, a model that forecasts what any perturbation would do to any cell, and MAP is offered as a step toward one rather than as one. The authors name their own bottleneck: the model works at the level of individual genes, which makes it expensive in computation and memory, and scaling it to full atlases will take a cheaper architecture. An earlier preprint of the same work carried slightly larger figures, which peer review moved down, and no group outside the collaboration has yet rerun the comparison. The test that matters is what happens when someone carries an unfamiliar name off one of these lists to a bench.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

The Model Was Never Shown the Drug, and It Still Predicted the Cell's Response

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.