Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

mAVE: A Watermark for Joint Audio-Visual Generation Models

About

As Joint Audio-Visual Generation Models see widespread commercial deployment, embedding watermarks has become essential for protecting vendor copyright and ensuring content provenance. However, existing techniques suffer from an architectural mismatch by treating modalities as decoupled entities, exposing a critical Binding Vulnerability. Adversaries exploit this via Swap Attacks by replacing authentic audio with malicious deepfakes while retaining the watermarked video. Because current detectors rely on independent verification ($Video_{wm}\vee Audio_{wm}$), they incorrectly authenticate the manipulated content, falsely attributing harmful media to the original vendor and severely damaging their reputation. To address this, we propose mAVE (Manifold Audio-Visual Entanglement), the first watermarking framework natively designed for joint architectures. mAVE cryptographically binds audio and video latents at initialization without fine-tuning, defining a Legitimate Entanglement Manifold via Inverse Transform Sampling. Experiments on state-of-the-art models (LTX-2, MOVA) demonstrate that mAVE guarantees performance-losslessness and provides an exponential security bound against Swap Attacks. Achieving near-perfect binding integrity ($>99\%$), mAVE offers a robust cryptographic defense for vendor copyright.

Luyang Si, Leyi Pan, Lijie Wen• 2026

Related benchmarks

TaskDatasetResultRank
Watermark ExtractionVideo Modality Extraction
Bit Accuracy94.9
9
Watermark ExtractionAudio Modality Extraction
Bit Accuracy92.8
9
Joint Audio-Visual GenerationVBench Audio-Visual Generation (test)
Subjective Score99.8
3
Swap Attack DefenseSwap Attack Defense Evaluation Set
True Positive (Auth)99.8
3
Audio Watermarking RobustnessMOVA LTX-2 (test)
MP3 Robustness85
2
Video Watermarking RobustnessLTX-2 MOVA (test)
Average Frame Score92.7
2
Showing 6 of 6 rows

Other info

Follow for update