AI watermark cracked in 4 hours: The fiasco of the new EU rule
Xpert Pre-Release
Available in 27 languages 📢
Prefer Xpert.Digital on GoogleⓘPublished on: August 24, 2026 / Updated on: August 24, 2026 – Author: Konrad Wolfenstein
Anthropic vs. Developer: Why the labeling requirement for AI texts is ineffective
Cat-and-mouse game over AI texts: The rise of the rapid circumvention industry
With the entry into full application of Article 50 of the European AI Regulation on August 2, 2026, a new era of transparency was supposed to begin. AI giants like Anthropic reacted promptly, implementing machine-readable watermarks in their text output to comply with the labeling requirement. However, technological reality caught up with the regulatory ideal in record time: just four hours after its introduction, a developer presented a tool that can completely remove the invisible marker.
This incident is far more than a technical anecdote – it marks the beginning of a highly dynamic circumvention industry and reveals a structural dilemma in European legislation. While new scientific studies demonstrate the fundamental weakness and legal fragility of current watermarking methods, companies, newsrooms, and content creators face massive legal and organizational uncertainties. The following article analyzes the risky architecture of the new labeling requirement, the unequal cat-and-mouse game between regulators and developers, and the far-reaching economic consequences for the entire digital information landscape.
Related to this:
When labeling requirements meet the circumvention industry: A race that seems decided even before it starts
On August 2, 2026, Article 50 of the European AI Regulation entered its full implementation phase, ushering in a new regulatory reality for providers of generative AI systems. Anthropic, the company behind the Claude language model, was among the first major providers to respond to this requirement by embedding a machine-readable watermark in its text output. What was intended as a technical compliance measure quickly became a cautionary tale about the limits of government regulation in the digital realm. Just four hours after Anthropic's announcement, French developer Guillaume Meyer released a tool designed to remove the very watermark that had just been introduced. This close temporal proximity between the regulatory measure and the technical circumvention is by no means coincidental, but rather points to a structural problem that pervades the entire debate surrounding AI transparency: as soon as a technical marking system is publicly described, the blueprint for its circumvention already exists.
The economic significance of this process extends far beyond the technical anecdote of a clever programmer. It touches upon core issues of information economics, regulatory theory, and platform economics simultaneously. If labeling requirements can be circumvented with reasonable effort, the entire calculation changes for companies, institutions, and authorities that rely on reliable proof of origin for digital content. A market of trustworthiness emerges, in which verifiability becomes a scarce and therefore valuable resource, while at the same time a counter-industry of obfuscation grows at an astonishing pace.
The technical architecture of the new labelling requirement
Anthropic's approach to watermarking text is based on a method called SynthID-Text, originally developed by Google DeepMind and published in the journal Nature in 2024. This method works fundamentally differently from visible watermarks in images. It intervenes in the text generation process of the language model itself by embedding a statistical pattern in minor word choices, such as the choice between semantically related terms like "cloudy" and "grey." This intervention remains completely imperceptible to the human reader, while an algorithm with the appropriate key can read the embedded pattern and deduce the probability that the text was generated by Claude.
Anthropic emphasizes several features of this method that distinguish it from cruder approaches. It adds no extra characters or hidden symbols to the text, does not increase computational costs, and cannot be traced back to a specific person, organization, or conversation. Beyond simple text marking, Anthropic also embeds generated files, such as images in SVG, PNG, and JPEG formats, with digitally signed provenance metadata according to the C2PA standard. This combination of an invisible watermark in the text and cryptographically signed provenance information for files aims to create the most comprehensive traceability system possible, extending across all Claude product areas, from the API and chat interface to developer tools like Claude Code.
The regulatory context that compelled this technical decision is noteworthy. Anthropic is one of approximately 190 signatories to the voluntary code of conduct on the transparency of AI-generated content, which the European Commission presented in the summer of 2026. This code specifies the requirements of Article 50(2) of the AI Regulation, according to which providers of generative AI systems must ensure that their output is identifiable as artificially generated in a machine-readable format. While compliance with the code is formally voluntary, the underlying transparency obligation itself is a binding legal requirement, the non-compliance with which can result in sanctions.
Four duties, one law, and many misunderstandings
The public debate surrounding mandatory labeling suffers from significant oversimplifications that hinder objective assessment. Article 50 of the AI Regulation actually contains four independent transparency obligations in four separate paragraphs, each addressing different parties and having different scopes. The first paragraph concerns AI systems that interact directly with natural persons, such as chatbots or voice assistants, and requires users to be informed when they are communicating with a machine. The second paragraph is directed at providers of generative AI systems like Anthropic, OpenAI, or Midjourney and obliges them to technically label all synthetic outputs, regardless of how these are subsequently used. The third paragraph concerns emotion recognition and biometric categorization systems. Finally, the fourth paragraph addresses the operators—that is, companies, associations, or newsrooms—that use AI in a professional context and requires them to provide visible labeling in two specific cases: deepfakes and AI-generated texts on topics of public interest.
A widespread misconception is that all AI-generated content must be labeled. This assumption is too simplistic, as the regulation includes a significant exception specifically for text, which does not exist for images. The visible labeling requirement for text is waived if two conditions are cumulatively met: First, a genuine content review must have taken place before publication, including a conscious check for accuracy, plausibility, and sources, and with a real possibility of amending or rejecting the text. Second, a clearly identifiable natural or legal person must assume editorial responsibility for the publication. However, this exception does not apply to deepfakes: For AI-generated or manipulated image, audio, or video content that constitutes a deepfake as defined by the regulation, the disclosure requirement applies regardless of whether a human has previously reviewed the content.
This asymmetrical treatment of text on the one hand and audiovisual deepfakes on the other reflects an economically justifiable risk assessment by the legislator. Texts written with AI assistance but editorially responsible differ fundamentally in their potential for deception from realistic-looking video forgeries of well-known individuals. At the same time, this very differentiation opens the door to interpretations that, in practice, lead to considerable legal uncertainty, especially for companies and content creators who work extensively with AI support.
The clash of two opposing interests
The announcement of the watermark triggered almost immediate resistance, stemming from at least two distinct motivations. Developer Guillaume Meyer, whose tool quickly went viral on GitHub, was bookmarked more than 20,000 times on the X platform, and attracted over 100 contributors, cited two personal reasons to Wired magazine . Firstly, he was drawn to the technical challenge itself. Secondly, as a native French speaker who uses AI tools to improve his English texts, he feared being stigmatized or penalized simply for this purely supportive use. It is noteworthy that he did not fundamentally oppose labeling AI-generated content, but rather criticized the specific implementation and the associated risks of misclassification.
This argument touches upon a sore point in the entire debate, one that has also been raised by independent observers. A watermark that merely indicates the involvement of a supporting Claude model in the text processing fails to distinguish between a completely machine-generated passage and a text where a human has simply corrected a single word in a paragraph they wrote themselves. Anyone who merely has their own texts proofread, translated, summarized, or reformatted inevitably produces marked text, which does not correspond to the common understanding of AI-generated content. This lack of differentiation capability of the watermark significantly undermines its practical significance, even before discussing technical workarounds.
Alongside Meyer's tool, further workarounds with different technical approaches quickly emerged. Software developer Erik Hughes stated that he needed only fifteen minutes to develop a tool using Claude itself, which removes invisible and visually deceptive characters, rearranges sentences within paragraphs, and replaces individual words with synonyms. Leon Chlon, a visiting researcher at Oxford University, described an even more pragmatic approach: one simply needs to condense Claude's response, translate it into a language with a significantly different semantics, such as an Arabic dialect, and then translate it back to reliably destroy the watermark. This variety of workaround strategies, ranging from highly specialized rewriting algorithms to simple translation loops, illustrates that this is not an isolated security vulnerability, but rather a structural problem inherent in the entire watermarking approach.
Why the watermark is fragile from a scientific perspective
The vulnerability of SynthID-Text and related methods to so-called meaning-preserving attacks is by no means a recent discovery resulting from practical experience with Meyer's tool, but has already been the subject of several scientific studies. A comprehensive robustness analysis from the summer of 2025 showed that while SynthID-Text achieves near-perfect recognition values under ideal conditions, with an F1 value of 1.0 and a false-positive rate of 0.0, its performance drops significantly under realistic interference. With simple synonym substitutions, the F1 value falls to a still acceptable 0.884, indicating moderate resistance to lexical variation. However, with more sophisticated paraphrasing attacks that alter both vocabulary and sentence structure, the F1 value falls to 0.842, while the false-positive rate rises to 0.23. The most serious consequences arise from back-translation via a structurally very different pivot language such as Chinese, a procedure that corresponds exactly to the approach described by Chlon.
A forensic study published in July 2026, with the programmatic title "AI Watermark Evidence Fails Forensic Readiness," presents even more drastic figures and fundamentally questions the legal admissibility of such watermarks. The study tested three representative watermarking methods, including the MarkLLM implementation of SynthID-Text, using 846 valid paraphrasing passes across fifteen different text templates. The result was unequivocal: Every single text initially recognized by the two comparison methods, KGW and Unigram, completely lost its watermark after just one paraphrasing pass—a removal rate of 100 percent. SynthID-Text performed only marginally better with a removal rate of 98.3 percent. Even before any manipulation, the false negative rates were alarmingly high, reaching 80 percent for SynthID, meaning that a significant portion of the text actually watermarked was never recognized as such in the first place.
Particularly explosive is a methodological detail that the study refers to as a paradox rate of 18.6 percent: In this percentage of cases, the SynthID method produced contradictory recognition results, with 80 percent of its own unaltered, watermarked source material ending up in a gray area of uncertainty, where the system itself could not make a definitive statement about the origin. Furthermore, the method incorrectly classified 5.4 percent of the paraphrased, actually human-written control texts as AI-generated. The study authors conclude that none of the three methods examined meets more than two of the five so-called Daubert criteria, the standard established in US jurisprudence for assessing the scientific admissibility of forensic evidence. These findings fundamentally undermine the notion that watermarks could ever serve as reliable evidence in legal disputes.
🎯🎯🎯 Data-driven B2B industry hub as a quasi-in-house solution

The quasi-in-house solution: How Xpert.Digital closes operational gaps in B2B marketing and sales – Smart Content-Driven Business - Image: Xpert.Digital
Xpert.Digital is a data-driven B2B industry hub led by Konrad Wolfenstein . The company acts as an external, quasi-in-house solution for industrial partners, closing operational gaps in marketing, content, and sales – without requiring additional resources on the client side.
More information here:
AI watermarking in practical testing: Why technical labeling often fails
The economic logic of the circumvention industry
From an economic perspective, the phenomenon of watermark circumvention can be described as a classic cost-benefit analysis in an asymmetric competition between regulator and regulated actor. An insightful example is provided by an analysis that quantifies the marginal cost of circumvention: Anyone using the publicly accessible detection API announced by Anthropic can iteratively generate paraphrases and test them against the detector until the watermark is no longer registered, at a cost of approximately four US cents per pass for a thousand-word article. This calculation reveals a deeper dilemma inherent in any publicly verifiable labeling solution: A detection API intended for verification inevitably also functions as a circumvention oracle, allowing any manipulation to be adapted until it remains undetected.
This finding is by no means new, but has already been theoretically anticipated. A widely cited paper from 2023 argues mathematically that no watermarking method can reliably withstand a sufficiently motivated paraphrasing attacker. Another study from 2024 even succeeded in reverse-engineering the secret rule of a commercial provider for under fifty US dollars, thereby increasing the success rate of a targeted deletion attack from almost zero to over 85 percent. The fundamental insight of this line of research is that as watermarked and watermark-free, but high-quality, text become more similar, the subset of errors that preserve text quality while simultaneously evading detection increases. In other words, the better a language model becomes and the more natural its texts sound, the greater, paradoxically, the scope for inconspicuous, quality-preserving circumvention strategies.
For companies like Anthropic, this creates a strategic dilemma that cannot be resolved solely through technical improvements. A watermark robust enough to withstand any form of paraphrasing and translation would likely have to intervene so deeply in the statistical properties of the generated text that it would either noticeably impair the quality of the output or leave such conspicuous patterns that it would be easily identifiable and therefore also vulnerable to attack. Anthropic itself explicitly emphasizes that the chosen method has no practical impact on the quality or content of the output, but this very restraint regarding the depth of intervention is likely the reason for the comparatively low robustness against targeted manipulation. It is a classic trade-off between user experience and security, in which Anthropic has consciously opted for the former.
Regulatory objectives and technical reality diverge
The discrepancy between regulatory requirements and technical reality raises fundamental questions about the effectiveness of the European regulatory approach. Californian legislation, specifically Senate Bill 942, explicitly requires disclosure that must be "permanent or exceptionally difficult to remove." Similarly, the European AI Regulation requires markings that are "sufficiently reliable and robust." Both sets of rules are based on the as-yet-untested assumption that technical watermarking methods can even meet this requirement. However, the available research suggests that this basic assumption, in its current form, is not tenable, at least not for purely text-based watermarks that rely on statistical patterns in word choice.
This finding has far-reaching implications for the European Union's regulatory strategy as a whole. A legislator who requires providers to label their expenditures "effectively, interoperably, robustly, and reliably" without prescribing a specific technical implementation ultimately leaves it to the market to determine whether these requirements are even achievable in practice. Should it turn out that none of the available watermarking technologies can permanently meet this requirement, a regulatory gap will emerge that would have to be closed either by stricter technical specifications, by supplementary legal sanction mechanisms against circumvention itself, or by a fundamental reassessment of the entire transparency approach. It is also noteworthy that the transitional provisions for AI systems already placed on the market before August 2, 2026, provide a grace period until December 2, 2026, which allows additional time for market monitoring before the labeling requirement takes full effect.
Another aspect concerns the scope of the marking beyond the text level. While text-based watermarks are demonstrably fragile, different technical standards apply to image, audio, and video content, particularly signed provenance metadata in the C2PA format. Although this metadata can also be removed, as demonstrated by Meyer's tool, which can explicitly remove C2PA data from image files such as PNG, JPEG, and SVG, the technical effort and detectability of such manipulation differ from simply rewording a text. In parallel, the Commission has introduced standardized EU icons for the visible marking of deepfakes and AI-generated texts relevant to publicity. These icons are designed to function independently of the machine-readable layer and thus represent an additional, more robust security measure.
Consequences for companies, editorial teams and content creators
For economic actors who produce large amounts of AI-supported content, for example in content creation, marketing, or digital communication, this complex situation has tangible practical consequences. Those considered operators under the regulation—that is, companies, associations, practices, or agencies that use AI in a professional context—must carefully distinguish between different content categories in the future. Texts used purely internally or content without any connection to matters of public interest are not subject to labeling requirements. However, if texts are published that are intended to inform the public about topics such as politics, economics, health, or science, the exemption applies only if actual editorial review has taken place and an identifiable person or organization assumes responsibility for it.
This requirement for genuine content review with a real possibility of rejection presents a significant organizational hurdle for many companies that use highly automated content workflows. It is insufficient to formally incorporate a review step into the workflow if, in practice, it merely amounts to a superficial glance without any real substantive engagement. At the same time, the existence of circumvention tools like Meyer's demonstrates that technical labeling by the AI provider itself cannot be a reliable substitute for a thorough internal compliance strategy. Companies that rely solely on the provider's embedded labeling to fulfill or avoid their own labeling obligations are treading on shaky ground, as this labeling has proven to be easily removable and, moreover, does not distinguish between texts generated entirely by machines and those that have only been edited with assistance.
From a competitive economics perspective, an interesting dynamic is also emerging between different AI providers. Since the labeling requirement, according to the regulation, follows the market location principle and thus applies to everyone who publishes or makes AI-generated content accessible in the European Union, regardless of the provider's location, differences in the technical implementation quality and robustness of watermarking methods could, in the long term, become an independent differentiating factor in the competition among language model providers. Providers whose labeling proves particularly easy to circumvent risk facing greater regulatory and reputational pressure than those who implement more robust, albeit potentially more complex, solutions.
The deeper question of the provability of digital origin
Beyond the immediate regulatory and business implications, the Anthropic v. Meyer case touches upon a more fundamental philosophical and economic question: In a world where AI-generated and human-written text are becoming increasingly indistinguishable, can the origin of digital content even be technically proven? An analysis specifically addressing the robustness of text watermarks succinctly articulates this problem: There is no evidence that paraphrasing reliably and predictably removes the watermark from an AI-generated text, as the effect depends on the specific method, the paraphrasing tool used, and the amount of text obtained. This ambiguity itself is already a problem, because it means that neither providers, users, nor regulatory authorities can predict with certainty whether a particular piece of content will remain identifiable as AI-generated or not.
This fundamental uncertainty has consequences that extend far beyond individual companies and affect the entire information economy. In a world where proving the origin of a text effectively becomes a matter of negotiation between the attacker's technical sophistication and the defender's limited resources, the burden of proof shifts. Anyone with an interest in concealing the true origin of content—whether for personal privacy reasons, as in Meyer's case, for commercial reasons, as in the mass production of content, or for manipulative purposes, as in disinformation campaigns—now has freely accessible, inexpensive tools at their disposal, which can be deployed with just a few clicks to achieve this goal. The availability of Meyer's open-source project, with its well over a thousand stars on GitHub and more than a hundred contributors, demonstrates that this is not a niche phenomenon for technically skilled individuals, but rather a rapidly professionalizing countermovement with considerable collective momentum for development.
At the same time, it would be an oversimplification to conclude from this development that labeling requirements are fundamentally pointless. Even a watermark that can be removed with targeted technical effort increases the cost of deliberate deception and creates a degree of traceability, at least for the vast majority of everyday use not specifically designed for concealment. The problem lies less in the uselessness of the approach itself than in the public communication of its limitations. If Anthropic or other providers give the impression that their watermarks are practically impossible to remove, while independent research and practical demonstrations prove the opposite within hours, a loss of trust occurs that extends beyond the technical issue and undermines the credibility of the entire transparency initiative.
An accelerating conflict
The dynamics of this race between marking and circumvention are likely to intensify rather than subside in the coming months and years. Firstly, the full enforcement of Article 50 of the AI Regulation, particularly after the end of the transition period in December 2026, will increase regulatory pressure on all providers of generative AI systems, thereby also boosting commercial interest in more robust watermarking methods. Secondly, scientific studies such as the one on the so-called SynGuard approach demonstrate that technical improvements are indeed possible: In tests, this semantically aware watermarking method achieved an average of 11.1 percentage points higher detection accuracy under various attack scenarios compared to SynthID text by aligning the watermark signal more closely with content-related rather than purely lexical patterns. Such advancements suggest that the current vulnerability does not represent an insurmountable technical obstacle, but merely reflects the current state of development.
Nevertheless, it remains foreseeable that every technical improvement to watermarking methods will be met with an equally ingenious counter-reaction from the developer community within a very short time, as the Meyer case impressively demonstrated. The real economic lesson from this incident, therefore, lies less in the question of whether a particular watermarking method can be made more robust, but rather in the fundamental realization that purely technical solutions alone will never suffice to solve the underlying trust problem of digital content. Regulatory approaches that rely exclusively on technical marking without developing accompanying institutional, legal, and market-based mechanisms risk engaging in a race they cannot win due to the asymmetric cost structure between defense and offense. For companies, regulatory authorities, and the general public, this means that the search for reliable digital provenance will remain an ongoing, constantly evolving challenge that can hardly ever be definitively solved by a single technical solution.
📈🚀 From visibility to trust 👀🤝 Your scalable path with Xpert.Digital
In industrial B2B, sustainable business relationships rarely emerge overnight. They develop step by step – through visibility, professional relevance, recurring touchpoints, and growing trust. Xpert.Digital's 4-stage model addresses precisely this: It offers a structured path that begins with a manageable entry point and can evolve into deeper collaboration in business development if needed.
Instead of relying on loud marketing promises, this model puts the relationship at the forefront. Companies start with clearly defined, easily calculable measures and then decide, based on their own experience, how far they want to expand the collaboration. A key factor for this undisturbed trust-building process: The platform completely avoids annoying advertising ads, so the editorial focus remains solely on the companies' expertise.
More information here:
Your global marketing and business development partner
☑️ Our business language is English or German
☑️ NEW: Correspondence in your native language!
I and my team are happy to be available to you as your personal advisor.
You can contact me by filling out the contact form here [email protected]:or simply call me at +49 7348 4088 965. My email address is
I'm looking forward to our joint project.

























