Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Universal Preference-Score-based Pairwise Speech Quality Assessment

About

To compare the performance of two speech generation systems, one of the most effective approaches is estimating the preference score between their generated speech. This paper proposes a novel universal preference-score-based pairwise speech quality assessment (UPPSQA) model, aimed at predicting the preference score between paired speech samples to determine which one has better quality. The model first predicts the absolute mean opinion score (MOS) for the two speech samples separately, and then aggregates them into a relative preference score using a preference function. To address the scarcity of preference data, we also construct a new pairwise speech dataset based on a MOS dataset for experiments. Experimental results confirm that, whether in training scenarios with different data types and label conditions, or in both in-domain and out-of-domain test scenarios, the prediction accuracy of UPP-SQA outperforms that of the baseline models, demonstrating its universality.

Yu-Fei Shi, Yang Ai, Zhen-Hua Ling• 2025

Related benchmarks

TaskDatasetResultRank
Pairwise Preference PredictionNISQA MOS-derived (test)
Accuracy83.46
4
Pairwise Preference PredictionSpeechEval human preference (test)
Accuracy86.31
4
Pairwise Preference PredictionIUB-COSINE-C unseen (test)
Accuracy78.06
4
Pairwise Preference PredictionSOMOS M MOS-derived (test)
Accuracy71.96
4
Pairwise Preference PredictionSOMOS NM MOS-derived (test)
Accuracy73.1
4
Pairwise Preference PredictionIUB-COSINE-S unseen (test)
Accuracy87.89
4
Pairwise Preference PredictionCHiLi M simulated (test)
Accuracy85.88
4
Pairwise Preference PredictionCHiLi NM simulated (test)
Accuracy81.05
4
Pairwise Preference PredictionSpeechJudge human preference (test)
Accuracy61.4
4
Showing 9 of 9 rows

Other info

Follow for update