Human-CLAP: Human-perception-based contrastive language-audio pretraining
About
Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance between audio and text in text-to-audio. However, the relationship between CLAPScore and human subjective evaluation scores is still unclarified. We show that CLAPScore has a low correlation with human subjective evaluation scores. Additionally, we propose a human-perception-based CLAP called Human-CLAP by training a contrastive language-audio model using the subjective evaluation score. In our experiments, the results indicate that our Human-CLAP improved the Spearman's rank correlation coefficient (SRCC) between the CLAPScore and the subjective evaluation scores by more than 0.25 compared with the conventional CLAP.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Audio-text alignment correlation | AudioCaps (test) | SRCC0.457 | 7 | |
| Compositional Text-Audio Alignment Correlation | RELATE | IS Kendall's Tau20.7 | 5 | |
| Contrastive Text-Audio Retrieval | CompA | Attribute Accuracy (Text)17.3 | 4 |