What is text watermarking?

Definition

Text watermarking changes how a model selects among acceptable next tokens. A secret or defined rule slightly favors some choices, causing a statistical pattern to accumulate across a sufficiently long passage without adding a visible label.

A compatible detector measures whether the expected pattern appears strongly enough to be unlikely by chance. Watermarking can support disclosure, but removal tools may rewrite or transform content until the detectable signal becomes too weak.

ELI5

Text watermarking gives AI-written text a hidden statistical pattern. The words still look ordinary, but a detector that knows what pattern to check can look across a long passage for signs that a model helped produce it.

For example, a model may quietly prefer certain equally suitable words more often than it normally would. Finding that preference can be evidence of model involvement, but it cannot prove that every word was generated or identify how much a person edited.

Frequently asked questions

Is a text watermark visible to a reader?

Usually not. It is commonly a statistical pattern spread across token choices rather than a visible label attached to the document.

Can text watermarks be removed?

They can often be weakened through enough rewriting, paraphrasing, translation, or regeneration, depending on the watermark and detector.

Videos explaining text watermarking

  1. A flat document fingerprint beside the words AI Text Leaves a Trace
  2. Jack Roberts and Nick Saraev beside the words Watermarks Break Fast