Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

De-mark: Watermark Removal in Large Language Models

About

Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models (LMs). However, the robustness of the watermarking schemes has not been well explored. In this paper, we present De-mark, an advanced framework designed to remove n-gram-based watermarks effectively. Our method utilizes a novel querying strategy, termed random selection probing, which aids in assessing the strength of the watermark and identifying the red-green list within the n-gram watermark. Experiments on popular LMs, such as Llama3 and ChatGPT, demonstrate the efficiency and effectiveness of De-mark in watermark removal and exploitation tasks.

Ruibo Chen, Yihan Wu, Junfeng Guo, Heng Huang• 2024

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningGSM8K
Accuracy80.5
108
Question AnsweringSQuAD
F1 Score0.832
21
Language UnderstandingMMLU
Accuracy72.4
21
Open-ended writingWritingBench
Score4.04
20
Watermark RemovalLlama-3.1-8B
DIPMark99.281
6
Watermark RemovalMinistral3-8B
DIPMark91.174
6
Watermark RemovalQwen3-8B
DIPMark4.56
6
Watermark RemovalFinal-text rewrite attacks
DIPMark (TPR@5% FPR)76.5
5
Showing 8 of 8 rows

Other info

Follow for update