Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS

About

Emotional interaction is increasingly crucial for conversational AI, yet current systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis. We propose an emotion-planning framework that determines the emotion prior to the textual generation, grounding the downstream emotional TTS in a streaming manner. The framework is implemented by a plug-and-play LLM module, initialized from pretrained LLMs, and trained by reinforcement learning (RL) with emotions as the actions. A hybrid reward is employed which combines imitation signals with theory-driven scoring, in which the theory of Plutchik's wheel of emotions is adopted. By experiments on DailyDialog, EmoryNLP, IMEOCAP, and MELD, our method outperforms prompting and finetuning baselines on both emotion determination and response quality. We finally implement an entire streaming pipeline for real-time deployment, with the speech quality confirming the framework's emotional alignment, contextual coherence, and expressive fluency. Codes, cases, and demos are available in https://sixingdeguo.github.io/EmoQ-page/.

Yue Zhao, Hongyan Li, Yong Chen, Luo Ji• 2026

Related benchmarks

TaskDatasetResultRank
Emotional Speech SynthesisDailyDialog
BERT Score0.54
18
Emotional Speech SynthesisIEMOCAP
BERT Score0.52
18
Emotional Speech SynthesisEmoryNLP
BERT Score0.5
18
Emotional Speech SynthesisMELD
BERT Score0.53
18
Emotion DeterminationDailyDialog
Reward0.57
8
Emotion DeterminationEmoryNLP
Reward0.71
8
Emotion DeterminationMELD
Reward0.86
8
Emotion DeterminationIEMOCAP
Reward81
8
Response GenerationEmoryNLP
BLEU-24.39
8
Response GenerationMELD
BLEU-23.89
8
Showing 10 of 12 rows

Other info

Follow for update