Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings

About

Understanding speaker attributes is crucial for voice-related applications, yet conventional approaches rely on fixed categorical labels, lacking semantic richness and zero-shot generalizability. We propose a novel framework for open-set speaker attribute prediction leveraging Large Language Model (LLM) embeddings to represent attributes in a continuous semantic space. To bridge the cross-modal gap, we introduce a keyword-appending strategy that structures broad semantic representations into a compact, discriminative manifold. Furthermore, we employ a top-k negative loss to establish robust decision boundaries in crowded semantic regions. Experimental results on LibriTTS-P demonstrate that our method outperforms closed-set benchmarks and generalizes effectively to unseen synonyms. Geometric analysis suggests that our strategies regularize the embedding manifold, balancing semantic cohesion with predictive clarity.

Byoungjun So, Jaejun Lee, Kyogu Lee• 2026

Related benchmarks

TaskDatasetResultRank
Speaker Attribute PredictionLibriTTS-P (test)
Micro-averaged F176.25
8
Showing 1 of 1 rows

Other info

Follow for update