Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

About

Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variability as listeners compare signals directly, producing cleaner labels. We study MOS-free preference prediction and propose PrefSQA, which incorporates uncertainty-aware logits, an impairment attention head, and a module based on non-matching-reference comparisons. We use and refine five datasets, including MOS-derived and low-noise simulated sets with matching and non-matching content, experiment with human preference sets, and test on unseen data. Experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over the baselines, highlighting the value of high-quality preference data and demonstrating the effectiveness of the proposed method.

Junyi Fan, Donald S. Williamson• 2026

Related benchmarks

TaskDatasetResultRank
Pairwise Preference PredictionNISQA MOS-derived (test)
Accuracy83.84
4
Pairwise Preference PredictionSOMOS M MOS-derived (test)
Accuracy73.27
4
Pairwise Preference PredictionSOMOS NM MOS-derived (test)
Accuracy74.72
4
Pairwise Preference PredictionCHiLi M simulated (test)
Accuracy96.29
4
Pairwise Preference PredictionCHiLi NM simulated (test)
Accuracy90.37
4
Pairwise Preference PredictionSpeechEval human preference (test)
Accuracy86.85
4
Pairwise Preference PredictionIUB-COSINE-C unseen (test)
Accuracy83.5
4
Pairwise Preference PredictionIUB-COSINE-S unseen (test)
Accuracy91.61
4
Pairwise Preference PredictionSpeechJudge human preference (test)
Accuracy68.2
4
Showing 9 of 9 rows

Other info

Follow for update