Models launched on or after 2 August will support machine-readable content marking.
Anthropic’s new models will begin watermarking content in the EU, as strict obligations from the bloc’s landmark legislation on AI start taking effect.
New transparency rules under the EU AI Act came into force earlier this month and require certain AI systems to tell users when they are interacting with AI and how the content they are consuming is generated or altered by it.
The Claude maker is one of around 200 signatories of the Act’s Code of Practice on Transparency of AI-Generated Content – a voluntary tool aimed at helping businesses comply with the sweeping laws. Several major AI providers – including OpenAI, France’s Mistral, Meta and Microsoft – are also signatories.
Anthropic said that models launched on or after 2 August will support machine-readable content marking, meaning files generated by its AI will include digitally signed provenance metadata such as embedded watermarks on text complying with Coalition for Content Provenance and Authenticity (C2PA) standards.
The watermarks will only appear when supported Claude models are accessed through cloud partners Amazon Web Services, Google Cloud or Microsoft Foundry, Anthropic said. The company is also working on adding watermarking capabilities to its older models.
“Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from,” Anthropic explained. “If a signed metadata label is present, it signals that a file was processed by Claude and lets you detect whether the file has been tampered with.”
AI labelling has emerged as a way to for users to detect artificially altered content as the technology advances to produce near-realistic outputs.
OpenAI’s now defunct image generating model Sora, for example, was outfitted with C2PA mechanisms, as are other major platforms such as TikTok and YouTube.
China, meanwhile, rolled out a new law last year that forced social media companies in the country to label all AI-generated content, including text, images, video and audio.
Despite these measures, content created or altered using AI is going increasingly undetected.
Anthropic said that it is “working” to enable users and other third parties to detect Claude’s embedded watermarks and provenance metadata, but did not specify how.
Meanwhile, the company acknowledged several pitfalls to content marking. It explained that Claude-generated content may not carry a watermark if text has been heavily edited, paraphrased, or translated; if the generated passage is too small; or if a file’s metadata was forcibly stripped through conversion or screenshots.
At times, the model’s watermark may appear on human-generated content if it was processed through Claude for reading or summarising, Anthropic said.
Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.
You must be logged in to post a comment Login