DeepMind watermarks AI-designed proteins with SynthID Bio
DeepMind has extended its SynthID watermarking approach from images and text to proteins, embedding a detectable signal in both the amino acid sequence and the AlphaFold 3-predicted structure without disrupting biological function. It's a research proof-of-concept, not a deployed screening tool, but it closes a real gap in biosecurity.

What shipped
DeepMind has introduced SynthID Bio, a proof-of-concept that extends the SynthID watermarking family — previously used to mark AI-generated images, audio, and text — into protein design. The idea is straightforward to state and hard to pull off: embed an imperceptible, detectable signal into an AI-designed protein so that it can be traced back to its synthetic origin, even after the protein has been physically synthesized in a lab.
This isn't a shipped product. It's a research result with open-sourced code, in vitro validation data, and model weights, released for the biosecurity and protein-design communities to build on.
Why proteins are a harder watermarking problem
Watermarking an image or a block of text is already nontrivial, but the signal only has to survive compression, cropping, or paraphrasing. A protein watermark has to survive something much more demanding: physical synthesis. The sequence gets translated into a real molecule, folded under cellular or in vitro conditions, and then has to still do its job — bind a target, catalyze a reaction, fold into the intended shape — while carrying a signal that can be recovered later from the synthesized protein itself, not just from the digital design file.
DeepMind's approach works at two levels, matching the two outputs a protein design pipeline typically produces:
- Sequence watermarking: the system subtly biases which amino acids get selected at each position during generation, creating a statistically detectable pattern without changing the protein's function.
- Structure watermarking: DeepMind fine-tuned AlphaFold 3's diffusion network so the watermark is built into the model's weights. Predicted 3D coordinates carry the signature by construction, regardless of who runs the model or how.
That second piece is the more interesting engineering choice. Rather than post-processing an output to inject a watermark, the signal is baked into the generative model itself — closer to how image watermarking has moved from pixel-level post-processing toward signals embedded in the generation process.

The validation is real, if narrow
DeepMind tested watermarked protein binders against three targets: VEGF-A, the SARS-CoV-2 spike protein receptor-binding domain, and PD-L1. Across these, watermarked designs matched the hit rates and binding affinity of non-watermarked versions, and preserved the natural diversity you'd want in a candidate pool. That's the right bar to clear — a watermark that degrades function or collapses design diversity would be a net loss for anyone using these tools legitimately. Three target classes is a reasonable first proof point, but it's not evidence the technique generalizes cleanly across the full space of protein design tasks, especially more complex multi-domain or membrane proteins.
DeepMind is also not working alone on the extension of this idea. The Hie lab at Stanford and the Arc Institute have applied watermarking to bacteriophage genomes using the Evo 2 model, which suggests the approach is being tested outside DeepMind's own pipeline rather than treated as a closed, single-vendor feature.
Why this matters now
DNA synthesis providers already screen incoming orders against databases of known pathogen sequences before they'll manufacture anything. That screening works when a sequence resembles something already catalogued as dangerous. It does not work well against a sequence an AI model designed from scratch to achieve a function, because a novel AI-designed sequence may not resemble anything in a threat database at all — the screening has nothing to pattern-match against.
A detectable watermark doesn't close that gap by itself, but it adds an orthogonal signal: not "does this look like a known threat" but "was this AI-designed, and if so, by which system." That's useful even for entirely benign designs — it helps prevent AI-generated sequences from quietly polluting public resources like the Protein Data Bank or GenBank, where mislabeled synthetic entries degrade the training and reference value of the databases everyone in the field depends on.
James Diggans at Twist Bioscience, a DNA synthesis company, described watermarking as “a promising new addition to the biosecurity toolbox that could strengthen screening” — notably framed as additive, not a replacement for existing screening. Sarah Carter, a biosecurity policy expert, called it “an important piece of the puzzle for tracking the provenance of biological designs.” Both framings are accurate: this is one layer in a defense that needs several.
What's still unresolved
DeepMind is explicit that the watermark's robustness against deliberate tampering needs more work, and that extending the technique to more complex biological objects beyond single protein chains is ongoing research. Those are not small caveats. A watermark that a motivated adversary can strip without breaking function provides a false sense of coverage that may be worse than no watermark at all. And biosecurity-relevant designs are frequently more complex than single binders — multi-chain assemblies, engineered pathways, modified nucleic acid systems.
The open-sourcing of weights and code is the right call for a result like this: watermarking schemes earn trust through independent red-teaming, not vendor claims. I'd want to see this tested against adversarial sequence-editing attacks before anyone treats it as a dependable biosecurity control, and I'd expect that testing to happen precisely because the weights are public. For now, this is a credible first step toward attributable AI-designed biology — not yet a safeguard you'd build a screening policy around.