Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation

About

In this paper, we propose Vo-Ve, a novel voice-vector embedding that captures speaker identity. Unlike conventional speaker embeddings, Vo-Ve is explainable, as it contains the probabilities of explicit voice attribute classes. Through extensive analysis, we demonstrate that Vo-Ve not only evaluates speaker similarity competitively with conventional techniques but also provides an interpretable explanation in terms of voice attributes. We strongly believe that Vo-Ve can enhance evaluation schemes across various speech tasks due to its high-level explainability.

Jaejun Lee, Kyogu Lee• 2025

Related benchmarks

TaskDatasetResultRank
Speaker Attribute PredictionLibriTTS-P (test)
Micro-averaged F172.86
8
Showing 1 of 1 rows

Other info

Follow for update