Claude incorporates invisible watermarks to the text generated by AI that survive copy and paste.

The new strategy of Anthropic to combat disinformation: this is how Claude's watermarks work

3 minutes

claude copiar pegar
Add DEMÓCRATA to Google

Published

Last updated

3 minutes

Anthropic has begun to incorporate invisible watermarks in the text generated by Claude. The company has introduced the system in the Claude models launched since August 2 and is working to extend it to earlier models as well. The watermark is integrated directly into the generated text and is designed to be invisible to the user and not alter its meaning or readability.

The main novelty is that this signal is not only associated with the conversation or the platform from which the content has been generated. Being integrated into the text itself, it can accompany it when the user copies and pastes it into another document, email, or webpage. Anthropic also points out that it can survive certain modifications of the content.

The goal is to facilitate the subsequent identification of content generated by Claude. Anthropic is developing tools that allow these marks to be detected and intends to offer mechanisms for third parties to verify them. The system is applied at the model level, so it does not depend on whether the user uses Claude directly from the web or from a specific application.

A mark that cannot be seen, but remains in the text

The technology works in a different way than a conventional watermark. The user will not see any symbol, label, or notice next to the generated text. The signal is incorporated during generation through imperceptible modifications that allow the content to be identified later with specific tools.

This means that copying a response from Claude and pasting it elsewhere does not automatically remove the mark. It is also not necessary to keep the original file for the system to accompany the text. Anthropic assures that the mark can persist even after some editing operations.

However, this does not mean that identification is infallible. Substantial editing, a significant transformation of the text, or a translation can weaken or eliminate the signal. Furthermore, if a detector finds a mark, it would demonstrate that the text comes from Claude, but its absence would not allow concluding that a content has necessarily been written by a person.

The measure comes with a focus on AI transparency

The decision comes in parallel to the development of the new European requirements on transparency of content generated through artificial intelligence. Anthropic links the marking of its content with its commitments related to the European Union Artificial Intelligence Regulation and, in particular, with the transparency obligations for content generated or manipulated through AI.

The movement also represents a significant change compared to systems that try to determine retrospectively whether a text seems to be written by artificial intelligence. Instead of only analyzing style, sentence structure, or statistical probabilities of language, the marking introduces a provenance signal from the moment Claude generates the content.

The difference is important because traditional detectors can make mistakes. An academic investigation on text marking techniques has precisely studied the possibilities and limitations of these signals and their resistance against different language models.

There will also be provenance information in the files

Anthropic is applying a different mechanism to certain files generated by Claude. When the format allows, the company uses digitally signed provenance metadata through the C2PA standard, a technology designed to record the origin and history of certain digital content.

In this way, the company differentiates between the marking of the generated text and the provenance information that can be incorporated into certain files. The common goal is to facilitate the identification of AI-generated content after it leaves the platform where it was created.

The measure places Claude in a trend that is already affecting other artificial intelligence systems. Google, for example, has developed SynthID to incorporate provenance signals into certain content generated by its models, although the techniques and formats used are not necessarily the same.