Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection under Perturbations

About

With the rapid deployment of speech generation systems in open environments, providing verifiable source attribution and copyright accountability for audio content has become critical. A gap in current research is the lack of a unified benchmark that systematically compares different watermark injection methods under realistic distribution shifts. To address this, we build VoxWatermark by applying 10 watermarking methods (4 neural and 6 traditional) with unified injection and annotation on multilingual, multi-source corpora, and introducing no-box, black-box, and white-box perturbations to simulate real recording and transmission conditions. Based on this benchmark, we propose AudioWMD as a robust baseline detector for large-scale, multi-method, cross-distribution settings. Results show that injection-method diversity and distribution shifts affect detection stability, while validating the effectiveness and scalability of AudioWMD. Dataset and code are publicly available.

Farnaz Sedaghati, Yuxi Wang, Zicheng Weng, Wei Rao• 2026

Related benchmarks

TaskDatasetResultRank
Watermark DetectionVoxPopuli-100k Scenario T1
AUC77.15
34
Audio Watermark DetectionVoxWatermark Set 1 (T1) (test)
TPR98
10
Audio Watermark DetectionVoxWatermark T2 (test)
AUROC70.02
8
Audio Watermark DetectionVoxWatermark 2 (test)
TPR77.78
8
Audio Watermark DetectionVoxWatermark (val)
Overall AUROC88.3
2
Audio Watermark DetectionVoxWatermark (test 2)
Overall AUROC63.2
2
Showing 6 of 6 rows

Other info

Follow for update