Where a Watermark Hides Decides Whether It Survives

A picture generator does not start with a picture. It starts with a small block of numbers, a kind of compressed sketch, and only at the last step unfolds that sketch into something a person can look at. That sketch is also where most of the invisible marks meant to identify AI images now live: not in the pixels, but in the draft behind them. Which invites an obvious question. If the mark is in there, what stops someone from painting over the part of the sketch that holds it?
Eleven researchers at the MSU AI Institute and the Trusted AI Research Center of the Russian Academy of Sciences have put that question to six marking schemes, the kind now being built to satisfy AI-labeling rules in the EU and China, in a preprint posted Oct. 1, 2026 that the authors say has been accepted for publication at the IEEE ICDM 2026 conference. Kirill Aistov and colleagues call their method Latent Frequency Masking: it overwrites part of the frequency content of the compressed sketch, then lets the generator's own decoder turn the altered sketch back into an image. The picture itself is never retouched.

The six schemes tested were Stable Signature, Tree-Ring, RingID, Gaussian Shading, MaXsive and METR. The images came from a thousand prompts drawn from each of two public prompt collections. Five of them, all but Gaussian Shading, could be pushed to a detection rate near zero at the strict operating point this field reports, the one where a checker is allowed to cry wolf on a single clean image in a hundred. The settings were tuned separately for each scheme, so there is no one configuration that defeats all five. One of the attack's two variants fails outright on MaXsive. And the paper's own summary of what it achieved is the careful one: it "removes or substantially weakens" several watermarks.
Gaussian Shading is the most informative result in the paper, because it held. The explanation is structural rather than lucky. Gaussian Shading spreads its mark through the whole compressed sketch and is designed to leave the sketch's statistics looking untouched, so overwriting one part of it leaves enough of the mark behind to find. The schemes that fell do the opposite: they concentrate their evidence in a compact and predictable place. The authors' conclusion is that keeping watermark information in a predictable spot may create an exploitable weakness, while spreading it out improves resistance to this whole class of attack.
What the attacker has to bring is as important as what the attack does. The paper's threat model hands the adversary the watermarked image, a publicly available generator of the same kind that produced it, and the ability to ask the public checker whether a mark is still there. It withholds the secret key, the original prompt, the clean image and any look inside the checker. The authors state the limit that follows: the combination is realistic for many open generation pipelines but may not hold for proprietary systems. Nothing in the paper was tried against a commercial tracing system in live use.
The quality claim needs the same care. What the paper measures is three comparisons between the attacked image and the watermarked original, plus one score that judges an image on its own with nothing to compare it to, and the result is a better removal-against-quality trade-off than the competing attacks managed. It is not a claim that the picture came out unchanged: the authors note that their attack still disturbs color in parts of the image, and that the setting which removes a mark most thoroughly costs more image quality than the gentler one. The paper does not say whose code it attacked for any of the six schemes, and an attack measured against each scheme's own reference implementation would mean more than one measured against a rebuild.
None of this opens a new front. The paper's own related-work section lays out an active attack literature. The strongest comparison in it is UnMarker, presented at the IEEE Symposium on Security and Privacy in 2025, which argues that a robust invisible mark has to ride on an image's frequency content, and attacks it there without any feedback from the detector and without knowing which scheme it faces. Those are weaker assumptions than the new work makes. On the paper's own averaged table UnMarker matches Latent Frequency Masking's removal rate; where the new attack wins is speed and the damage it leaves in the picture. What is new here is a cheaper route to an outcome the field already knew was reachable, and an explanation of why one scheme kept resisting it.
Watermark robustness should be judged against attacks on the compressed sketch, not only against the ordinary handling an image survives on its way across the internet: the compression, blurring, noise, cropping and rotation that make up most published robustness tests. Marks should not keep most of their evidence in one predictable place. Detectors should be built to stay sensitive after part of the sketch has been replaced, and training should treat an attack of this kind as one more thing a mark has to live through. For the people building provenance tools, that is advice they can act on before the version of record appears.
Sources
- arXivPreprint
