MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System

About

Multi-modal sarcasm detection has attracted much recent attention. Nevertheless, the existing benchmark (MMSD) has some shortcomings that hinder the development of reliable multi-modal sarcasm detection system: (1) There are some spurious cues in MMSD, leading to the model bias learning; (2) The negative samples in MMSD are not always reasonable. To solve the aforementioned issues, we introduce MMSD2.0, a correction dataset that fixes the shortcomings of MMSD, by removing the spurious cues and re-annotating the unreasonable samples. Meanwhile, we present a novel framework called multi-view CLIP that is capable of leveraging multi-grained cues from multiple perspectives (i.e., text, image, and text-image interaction view) for multi-modal sarcasm detection. Extensive experiments show that MMSD2.0 is a valuable benchmark for building reliable multi-modal sarcasm detection systems and multi-view CLIP can significantly outperform the previous best baselines.

Libo Qin, Shijue Huang, Qiguang Chen, Chenran Cai, Yudi Zhang, Bin Liang, Wanxiang Che, Ruifeng Xu• 2023

Related benchmarks

Task	Dataset	Result
Multi-modal sarcasm detection	MMSD 2.0	Accuracy85.64	37
Sarcasm Detection	MSD	Accuracy93.97	33
Multi-modal sarcasm detection	MMSD	Accuracy88.33	25

Showing 3 of 3 rows

Other info

Code

Follow for update

@wizwand_team Discord