ESPNFollow live: Phillies even up Game 2 on solo homers by Schwarber and HarperESPN DeportesGil Mora con dolores en el tobillo; bajo observación previo a EE.UU.RTP DesportoCristiano Ronaldo arrisca seis meses de suspensão por abandonar seleção portuguesaDaily MaverickWILD OMISSION: South Africa’s lion bone trade — the consequences the government never consideredThe Jerusalem PostUK has 'strong indications' Iran was involved in RAF airbase bombing plot, Burnham saysTechCrunchFactory CEO just accused his VC board advisor of spying for CognitionBBC NewsPunish Man City this season, say other club chiefsABC News (Australia)Mother felt 'completely abandoned' during traumatic birth at NSW hospitalBBC News BrasilO cálculo de Lula para não ir ao debate da GloboRFI 中文英国“西藏观察”报告:无人机如何成为《高原上的眼睛》Radio-CanadaLe Canada à l’heure de la vérité et de la réconciliationDeadline‘Nobody Wants This’ Rabbi Consultant Slams Adam Brody As An “Ignorant Jew” Over “Free Palestine” Support
The Daily Newsstand · Free, Always
Wednesday, September 30, 2026

Google figures out how to watermark AI-designed proteins

Translate

AI-based tools seem to be causing security threats on a nearly daily basis, in part because we’ve been slow to recognize potential threats. One area where we seem to be ahead of the game, however, is in biosecurity.

As with most things biological, utility goes hand in hand with threats. We’ve developed increasingly sophisticated tools for designing proteins and seen some major successes, such as AI-designed enzymes that can digest plastics or block venom proteins. But these same tools could be used to make toxins or alter the behavior of viral proteins.

And the software we use to identify DNA sequences that encode potentially threatening proteins doesn’t pick out AI-designed proteins, since nobody has characterized them well enough to know that they’re threats. Nearly a year after that risk was flagged, it still wasn’t clear what anyone could do about it.

On Wednesday, the DeepMind team at Google published a research paper offering a potential solution: protein watermarking. The system creates a watermark on protein sequences themselves without compromising the protein’s function. This allows new proteins designed by trusted researchers to be identified, opening everything else up to closer scrutiny.

Does that even work?

The work was based on Google’s SynthID tech, which can add a subtle watermark to AI-generated digital material. The watermark influences the probability of certain choices the AI makes, and that bias ends up systematically distributed throughout the product, whether it’s text or images. Because you can’t identify the watermark without knowing how it was encoded, it’s impossible to remove. And because it’s distributed throughout the image, it can survive basic exporting, resizing, and so on.

It’s pretty easy to see how this can work with subtle differences in things like the colors of a photo. It’s a whole lot harder to see how you can do it with a protein.

Proteins are composed of only 20 amino acids, any of which could be essential for structural integrity or catalytic activity. While some of these amino acids are chemically similar (like leucine and isoleucine), others have opposite charges. Many proteins have significant regions where limited changes to their amino acid sequence are tolerable and other areas where even a slight deviation from the existing sequence inactivates the protein.

Proteins are also small. While images often contain millions of pixels, proteins containing 500 amino acids are fairly large. That’s a lot less raw material to hide any sort of signal in.

So it wasn’t clear the SynthID tech would work; it might be unable to hide sufficient signal in a typical protein, or, if it crammed in enough information to create a functional watermark, the resulting proteins might be inactive. The only way to find out was to try it.

Bringing watermarking to proteins

To better understand how the system works, it helps to know a bit about protein chemistry. Amino acids have a constant section primarily made of two carbon atoms linked to a nitrogen. A protein is made by linking up a series of these constant sections to form a long chain called a backbone. Each amino acid also has what is called a side chain hanging off it. These can range in complexity from a single hydrogen atom to large ring structures; the side chains can be basic hydrocarbons, acids, bases, and more.

The side chains determine how the backbone folds up in three-dimensional space. In a soluble protein, all of the pure hydrocarbon side chains end up packed into the center, while the acids, bases, and hydroxyl-containing side chains face the water. Once the protein folds into its final form, the backbone will describe the protein’s overall shape.

The Google team started with one of the most popular AI protein design tools, ProteinMPNN. (This software was developed by the Baker Lab, and David Baker was honored with the same Nobel Prize that was shared with the head of DeepMind.)

ProteinMPNN works in a two-stage process. First, a separate tool describes a backbone configuration that is appropriate for the design. Next, ProteinMPNN works its way down the backbone, placing side chains one amino acid at a time. Each amino acid is chosen based on its ability to fit into the shape defined by the backbone, interact with neighboring amino acids, and fit any other constraints defined by the experiment. (Those constraints can include things like forming catalytic pockets or interacting with another protein.)

A variant of Google’s SynthID, called SynthIDBio, steps in during this process. It uses a key (similar to a cryptographic key) and the identity of the previously chosen amino acids to suggest a new one. ProteinMPNN then determines whether the amino acid suggested by SynthID works from the perspective of forming a functional protein. If it doesn’t, it rejects it. If not, it moves on. Put differently, as the system works through the backbone one amino acid at a time, it only incorporates watermark amino acids when they’re consistent with a functional protein.

One way to think about this is that, when the system comes across a location where a set of chemically related amino acids will all work (like leucine/isoleucine/valine or serine/threonine), it will use one that’s consistent with the watermark when possible. Another way to look at the process, suggested by one of the people involved in developing the system, is that it searches through the space occupied by functional proteins for the subset that happens to have a sufficient number of watermark amino acids.

As a result, the watermark is randomly distributed across the entire length of the protein, and detecting one isn’t a simple yes-or-no question. You have to scan the whole sequence, knowing the key, and measure how often the amino acids suggested by SynthIDBio actually appear in the final sequence. Google has also developed the software needed to do this.

It’s alive!

The question, then, is whether watermarked proteins are functional. The team used the system to design proteins that physically interact with key natural proteins previously targeted with AI designs. And the watermarked versions worked just fine, binding the intended targets. This isn’t as rigorous a test as finding a catalyst, but it suggests that there’s no reason to expect serious problems in more complicated design tasks.

So as long as a protein is long enough, the system can detect a watermark. How might that be useful? Again, it comes down to biosecurity. When someone orders DNA sequences, the people who make the DNA normally screen the sequence for its ability to encode portions of viruses, toxic proteins, and other similar threats. Right now, however, when they see a protein that doesn’t look similar to anything we already know about—something that’s potentially true for any AI-designed proteins—they can’t assess its threat.

Google envisions its system as a way to give DNA synthesizers greater confidence when it comes to these proteins. If they’re given a set of keys from trusted organizations, like universities or major biotech companies, they can quickly determine whether an unknown protein is an AI design from a trusted source. This should let them focus their attention and resources on evaluating the things that aren’t trusted—specifically those things that look like AI designs but don’t come from a trusted source.

In other words, it doesn’t guarantee the security of DNA orders, but it simplifies the threat-screening process.

The team behind the work has highlighted several potential holes. For starters, the whole system is only as secure as the system used to distribute and maintain the keys used for watermarking. Very short proteins can potentially incorporate too few watermark amino acids to be identified. And it’s possible to pad the watermarked sequence with something that lacks it—think fusing the AI-designed protein with a natural fluorescent protein—which might dilute the watermark.

There are also a number of AI-based protein design software packages that don’t rely on ProteinMPNN. Some, but not all, use a similar “one amino acid at a time” approach that should integrate well with SynthIDBio. So until Google figures out how to handle other forms of integration, not everyone will be able to watermark the proteins they’re designing.

And since watermark identification is done on a statistical basis, how you set the cutoff makes a big difference in terms of false positives and false negatives.

It’s not clear this will be especially useful in practice, at least in its original form. But it’s pretty interesting that it works at all. And it’s nice to know that at least some people in the AI industry are thinking about how to limit risks.

Nature, 2026. DOI: 10.1038/s41586-026-10965-y (About DOIs).

Photo of John Timmer

John is Ars Technica's science editor. He has a Bachelor of Arts in Biochemistry from Columbia University, and a Ph.D. in Molecular and Cell Biology from the University of California, Berkeley. When physically separated from his keyboard, he tends to seek out a bicycle, or a scenic location for communing with his hiking boots.

View the original on Ars Technica →

KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.