Open Source Tool Attempts To Remove AI Watermarks From Claude Text

Published:

Developers have released an open source tool designed to remove or weaken artificial intelligence provenance markers shortly after Anthropic announced its invisible watermarking system for Claude generated text. The project, named watermarks-remover, focuses on removing different types of AI related markers from text and digital files, including systems linked with Claude, Google Gemini, SynthID, OpenAI related provenance formats, and several open source language models. The release has drawn attention as companies continue developing methods to identify AI generated content while regulators and organizations increasingly focus on transparency around artificial intelligence usage.

Anthropic introduced its watermarking approach on August 14 as part of its efforts to support compliance with the EU AI Act. Unlike traditional methods that insert visible or hidden characters into generated text, Anthropic’s system uses statistical changes in word selection during the generation process. Claude models are designed to subtly adjust which words are chosen while producing responses, creating a pattern that can later be identified through Anthropic’s watermarking key. The company said users cannot visually identify whether text contains a watermark and clarified that the marker does not include information about the user, organization, or specific conversation associated with the generated content. The newly developed watermarks-remover project uses multiple methods depending on the type of AI marker it targets. One layer of the tool focuses on removing direct hidden elements such as invisible Unicode characters, unusual spaces, bidirectional text characters, and other concealed text markers. These types of markers can be removed through automated scripts designed to identify and eliminate hidden formatting elements from digital content. The project also targets provenance metadata stored in various file formats, including PNG, JPEG, WebP, PDF, DOCX, HTML, Markdown, video, and audio files. It removes metadata formats such as C2PA, EXIF, and XMP that may contain information about content origin or processing history.

For statistical AI watermarks such as Claude’s approach, the tool follows a different method. Instead of directly deleting a hidden element, it attempts to weaken the watermark signal by significantly rewriting the original text. The process changes enough word choices and sentence structures to disrupt the statistical patterns used for detection. However, the developer behind the project has stated that removing statistical watermarks remains a best effort process and does not guarantee complete removal. Since Anthropic has not yet publicly released its watermark detection API or the keys required for independent verification, users cannot confirm whether rewritten text would still be detected by the company’s system.

Anthropic has also acknowledged that its watermarking system has limitations. According to the company, minor edits such as changing headings, rearranging paragraphs, or making small adjustments are unlikely to remove enough of the watermark signal to prevent detection. However, a complete rewrite where most or all words are changed could potentially remove the watermark. Anthropic noted that in such cases, questions may arise about whether the rewritten content should still be considered generated by the original AI model. The company is currently working on a watermark detection API that will allow users and organizations to evaluate whether Claude was likely involved in creating or processing a specific piece of content.

Follow the SPIN IDG WhatsApp Channel for updates across the Smart Pakistan Insights Network covering all of Pakistan’s technology ecosystem. 

Related articles

spot_img