Back to the notes

Watermarked LLM outputs look identical to humans

A quiz shows that even technical readers cannot spot which GPT-4 response has a cryptographic watermark baked in. The detection gap is real.

A row of books photographed at table height
Tom Hermans / Unsplash Unsplash License

Someone built a watermark quiz that hands you two LLM outputs and asks you to pick which one has a cryptographic watermark. I got three out of five wrong. The watermark is not a signature file or metadata tag. It is a statistical pattern in the token distribution. The model slightly favours certain tokens over others in a way that is invisible to readers but detectable by the watermark algorithm. OpenAI and other labs are experimenting with this for GPT-4 and newer models. The quiz proves the main point: humans cannot tell. The watermarked text reads like normal GPT-4 output. No repetition, no weird phrasing, no telltale sign that something is off. If you are trying to catch students or employees using AI by reading their work, this will not help you. The detection gap matters for two reasons. First, watermarks only work if the model provider implements them. If you are using an open-weight Llama derivative or a Chinese model, there is no watermark to detect. Second, watermarks can be stripped by paraphrasing the output through another model. You paste the watermarked text into Claude, ask it to rewrite, and the watermark is gone. This is not a solved problem. Watermarking helps when the output is unchanged and the provider cooperates, which is a narrow use case. The quiz is a reminder that detection tools lag behind generation tools by a long gap.


Source: Guess which of these LLM outputs is watermarked

Back to all notes

Behind the notes

Vikrant
Sharma.

Artificial Intelligence Engineer intern at Voxon Photonics in Adelaide. Studying a Master of Information and Communications Technology at UniSC, with a focus on data, machine learning and security.

Meet the person behind the work