NXAT

0.86111.59%

OpenAI rolls out weak sauce watermarking for AI text

LongbridgeAII'm LongbridgeAI, I can summarize articles.

OpenAI is implementing text watermarking via its 'textGrain' API to comply with the EU AI Act, marking eligible ChatGPT and Codex outputs in the EU. While Anthropic applies global watermarking using SynthID, OpenAI's approach allows optional global application for API customers. However, the technology faces criticism for low detection rates in short or functional texts and being easily defeated by simple word substitution.

OpenAI is now offering API customers the option to watermark the output of certain models in order to comply with the letter of Europe's AI Act. And within a few weeks, the biz plans to apply a form of digital watermarking – like it or not – to eligible ChatGPT and Codex text output in the European Union.

But unlike rival Anthropic, it isn't applying its watermarking technique to content generated outside of Europe, though API customers can choose to have the marks applied to content made anywhere. Moreover, OpenAI is using technology that misses a fair number of instances of AI-generated text and can easily be tricked.

The EU AI Act requires generative AI service providers to make model output identifiable in a way that's readable by machines.

Anthropic was the first major AI company to announce its approach to AI Act compliance, which involves the application of Google's SynthID-Text algorithm. It's doing so globally.

Microsoft and Meta have been working on similar technology for image-based AI because the EU is not the only market in which people have raised concerns about content provenance and the potential to use AI-generated content to produce misinformation. OpenAI already applies other content provenance signals to AI-generated images, audio, and video.

In keeping with the letter of the AI Act, OpenAI is now offering a text watermarking approach called textGrain [PDF] that "adds an invisible statistical signal to the model’s word choices. Our detector looks for that signal to assess whether a passage contains an OpenAI watermark."

Large language models predict word tokens one at a time from a probability distribution of potential words. By picking a different word to emit periodically – ideally a synonym – the resulting text deviates from the statistical pattern produced by an unbiased model. A watermark detector, which OpenAI is making available only to approved researchers and organizations, should flag that vocabulary divergence.

That's the theory, but in practice results vary. The target error rate is one percent, but in short passages of 200 words, the detector may only flag 80 percent of the watermarks, compared to 95 percent in a 400-word passage. The results are also worse in texts that are functional, like passages about math or code, as substitutions are more likely to introduce errors.

The textGrain paper does not address the likelihood that a statistically altered word choice might alter the meaning of a given passage. That's a possibility if the algorithm were to alter some salient fact or figure. Anthropic's approach, which also replaces certain output words though using a different algorithm, brings similar risks.

In any event, OpenAI's watermark is easily defeated by word substitution. In tests on 400-token passages, swapping out 10 percent of the words with synonyms reduced watermark detection from 92 percent to 66 percent. And when replacing 25 percent of words, the watermark signal could only be detected in 17 percent of cases.

Now witness the firepower of this barely adequate but compliant AI detector… ®

Login to unlock2,590characters for free

Due to copyright restrictions, please log in to your Longbridge account to view this content.
Thank you for your understanding and support of licensed content.