Skip to content
See the World Through ScienceA project of ALLATRA
Source: Peer-reviewedNature1 source

Google DeepMind Hides a Watermark Inside AI-Designed Proteins

Science

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Two-row flow diagram. In the top row a protein design pipeline samples watermarked sequences, filters them and passes a candidate binder to a detector. In the bottom row a structure prediction model is fine-tuned so that a separate watermark detector can identify its predictions.
The mark goes in while the sequence is being written, and a detector holding the secret key reads it back from the finished candidate. The lower row applies the same idea to a structure prediction model.Fig. 1 from David Stutz, Alexander I. Cowen-Rivers, Guillermo Ortiz-Jimenez, Jeremy Ratcliff, Vinicius Zambaldi, Lindsay Willmore, Josh Abramson, Harshnira Patani, Christina Kouridi, Florian Stimberg, Mel Vecerik, Alex Chu, Sukhdeep Singh, Sumanth Dathathri, Eliseo Papa, Valentin De Bortoli, Arnaud Doucet, Demis Hassabis, Jue Wang, Sven Gowal, Pushmeet Kohli (2026), "Function-preserving watermarking of AI-generated proteins", Nature — CC BY 4.0, fitted to 16:9 and resized

A Google DeepMind team has hidden a mark inside proteins designed by artificial intelligence, so that the molecule itself shows a machine made it. The system, called SynthIDBio, appeared in Nature on Sept. 30, 2026.

The authors say a record of which tool produced a sequence or a structure is becoming important for biosecurity and for the reliability of public biological databases, and that a watermark is one way to carry that record without a central registry. SynthIDBio has two parts: one marks the sequences that a protein design model writes, and one is a fine-tuned version of AlphaFold 3, the structure prediction model, that marks the shapes it predicts.

To check that marking does not break a protein, the group redesigned the sequences of existing binders for three targets, among them part of the SARS-CoV-2 virus, then measured in the laboratory how tightly the new versions stuck. On its own measurements, it reports no significant difference in binding strength between 222 unmarked and 267 marked sequences, and that a detector tuned to a 0.1% false-alarm rate flagged every marked design it tested. SynthIDBio is the team's own system, and the paper reports no independent evaluation of those results.

Four panels of bar and violin charts: detection rates for four sampling settings at three false alarm thresholds, the change in design pass rates under those settings, confidence score distributions for marked and unmarked designs, and estimated binding hit rates for designs put through a watermark removal step with and without the design filters.
Measured in software, before any protein was made. Stronger sampling settings raise the share of designs the detector catches (top left), and cost pass rates at the design filters (top right). — Fig. 3 from David Stutz, Alexander I. Cowen-Rivers, Guillermo Ortiz-Jimenez, Jeremy Ratcliff, Vinicius Zambaldi, Lindsay Willmore, Josh Abramson, Harshnira Patani, Christina Kouridi, Florian Stimberg, Mel Vecerik, Alex Chu, Sukhdeep Singh, Sumanth Dathathri, Eliseo Papa, Valentin De Bortoli, Arnaud Doucet, Demis Hassabis, Jue Wang, Sven Gowal, Pushmeet Kohli (2026), "Function-preserving watermarking of AI-generated proteins", Nature — CC BY 4.0, resized

For predicted structures, the authors report detection above 99.8% at the same false-alarm setting, with a negligible effect on how accurate the structures are. The sequence watermark can also be removed by putting a design back through the sequence step, and fewer of the resulting proteins were then estimated to bind their target.

The authors call the work a proof of concept, and say that using it for biosecurity screening or for checking database entries would need further research, industry-wide coordination and standards. The code and the laboratory data are posted on GitHub, along with instructions for obtaining the structure model's weights.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Google DeepMind Hides a Watermark Inside AI-Designed Proteins

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.