The EU’s new AI transparency rules have pushed a potentially misunderstood concept into the mainstream: watermarking AI-generated text.

In this article, David Baskerville discusses the principles and legislative background to the use of AI Watermarks, and explains how they may not, in the short term, be the panacaea for detecting AI-generated content that the regulations intend them to be.

The transparency provisions of the EU AI Act became applicable on 2nd August 2026. Article 50 requires providers of systems generating synthetic text, images, audio or video to make their outputs machine-readable and detectable as artificially generated or manipulated “where technically feasible”. The Act does not prescribe a particular form of text watermark. “Watermarking” is therefore better understood here as one family of technical approaches that may help meet the broader transparency requirement.

For law firms and other professional services organisations, this matters because AI watermarking is likely to become part of future governance, procurement and compliance conversations. The real issue is not simply whether AI-generated text can be detected, but whether organisations understand what detection can and cannot prove.

The real question is not just whether AI-generated text can be watermarked. It is whether that watermark will survive the way professional documents are actually created, revised, copied, paraphrased and improved.

Why text watermarking is different

When we talk about watermarking text, it is important to understand that there may be no visible mark at all and that is what makes the issue both interesting and potentially misleading.

Digital watermarking can already be less obvious. Images, audio and video can contain signals (also known as “provenance metadata”) that software can detect even though a human cannot readily see them.  Text is harder: it is copied into emails, Word documents, websites, databases and messaging applications. Formatting disappears. Metadata disappears. Copy-and-paste reduces everything back to the words themselves.

How the watermark can sit inside the wording

An LLM does not normally compose a sentence in quite the way a human does.

At each stage it is effectively choosing between many possible next pieces of text, often called tokens. A token might be a whole word, part of a word, punctuation or another small unit of language. Rather than inserting a hidden character or secret sequence, the model can be encouraged to select certain tokens from a mathematically determined group.

This is very different from the claims sometimes made about AI detectors looking for hidden Unicode characters, unusual spaces or other secret characters. In statistical watermarking, there may be nothing unusual in the text encoding at all: the signal emerges from the pattern of token choices across the text.

Across enough text, that pattern can become something a detector may recognise as evidence of a particular watermarking scheme.

In other words the watermark is not necessarily hidden behind the words. It may be encoded through the choice of the words themselves.

One approach to digital watermarking is, therefore, to put the signal into the language generation process itself. The European legislation deliberately does not prescribe one specific technology. It requires machine-readable marking that is effective, interoperable, robust and reliable as far as technically feasible. The practical implementation of these requirements is still developing. Different providers may adopt different combinations of watermarking, metadata, provenance information and detection technologies. For text specifically, one important area of research involves statistical patterns in the tokens selected by a Large Language Model (LLM).

What happens when the text is rewritten?

If someone rewrites watermarked AI text, the meaning may survive while many of the original token choices disappear. If enough wording changes, the detector may no longer have sufficient statistical evidence to identify the original pattern.

This is not just a theoretical concern. Researchers have repeatedly shown that paraphrasing and word substitution can weaken text watermarks. One July 2026 study found that meaning-preserving paraphrasing was able to defeat detection of the watermarking schemes tested at very high rates. That does not mean every watermarking system will fail, but it does show why organisations should treat detection results with caution.

What watermark washing means in practice

We might call this “watermark washing”: passing AI-generated content through transformations that preserve the substance while degrading or destroying the statistical characteristics used to identify its origin. In practice, one model could generate watermarked text, another could rewrite it, a third could make it more concise, and a human could then edit the result. What remains of the original watermark may be very little.

LLMs may simultaneously become the mechanism for creating text watermarks and one of the easiest mechanisms for removing them. 

The complication of multiple models

If a second model also uses watermarking, it could weaken or remove Model A’s statistical signature while introducing its own.

Run the text through a third model and the process could happen again.

This could create a chain in which one watermark progressively replaces another.

Model A → watermark A → Model B → watermark B → Model C → watermark C → human edit

At that point, what exactly does detecting watermark C tell us?

Detection might show that Model C generated or substantially transformed the final text. It does not necessarily tell us where the ideas originated, and failure to detect watermark A does not prove that Model A was never involved.

That is a crucial distinction between provenance and detection. Provenance is about the history and origin of content. Detection is about whether a particular technical signal is present in the final text.

Detection is not the same as authorship

Organisations will need to be careful. A watermark detector should not automatically be treated as an AI authorship detector. The reverse is equally important: the absence of a detectable watermark should not be treated as evidence that AI was not used. The original system may not have applied a watermark, the text may be too short for reliable detection, or subsequent rewriting, paraphrasing and editing may have weakened the signal.

The EU’s framing is principally about transparency and helping people recognise AI-generated or manipulated content. It is not a simple test of who authored the underlying ideas. For example, if a lawyer writes the substance of an advice note and asks an AI system to improve the grammar and readability, the AI may influence the final sequence of words without becoming the intellectual author of the advice. That distinction will become increasingly important if organisations start using watermark detection for governance, education, recruitment, compliance or disciplinary purposes.

The emerging watermark arms race

Watermark designers will try to make statistical signals survive copy and paste, small edits, paraphrasing, translation, summarisation and processing by another LLM. Meanwhile, researchers and users will discover transformations capable of weakening those signals. The challenge is fundamental: the watermark must survive changes to the text, but language is extraordinarily transformable. The more strongly a watermarking scheme biases or constrains token selection to preserve its signal, the greater the potential impact on output quality; the weaker the statistical signal, the easier it may be to disrupt through subsequent transformation.

Why watermark washing may happen anyway

The more interesting issue is that watermark washing could happen naturally. Modern knowledge work is increasingly multi-model: someone might research with one AI, draft with another, improve it with a third, run it through a corporate AI assistant and then edit the result themselves. Nobody may have intended to remove a watermark, but the original statistical fingerprint may still have been eroded through normal working practices.

The practical governance question

For organisations, the practical takeaway is straightforward: do not treat watermark detection as a simple yes-or-no test for responsibility, authorship or misconduct. It may become a useful signal, but it should sit alongside policy, user education, audit trails, disclosure practices and sensible judgement about how AI is actually being used.

The EU’s legislation may accelerate the deployment of AI watermarking technology. It may also accelerate the market for tools and workflows that remove, obscure or overwrite those signals, sometimes deliberately and sometimes simply as a by-product of normal drafting.

The next phase of AI detection may not simply be about watermarking. It may be about watermark washing.

David Baskerville

David Baskerville

07769 946883

Latest Articles

Talk to us today

Get In Touch

Discover more from Baskerville Drummond LLP

Subscribe now to keep reading and get access to the full archive.

Continue reading