EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis
About
Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained intensity. We propose EmoInstruct-TTS, a dual-path instruction-guided framework for emotional speech synthesis. We introduce Emotion2embed, a supervised semantic-acoustic emotion embedding covering 48 emotional states, including fine-grained categories and intensity levels. To infer embeddings from free-form instructions, we design an Instruction-Conditioned Emotion Flow Model (ICE-Flow) that generates acoustically grounded emotion representations. The inferred embeddings are integrated into an LLM-based synthesis pipeline to provide explicit emotional control while preserving semantic planning. Experiments show improved emotional controllability and speech naturalness over strong baselines.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Emotion-intensity Speech Synthesis | ESD & CNCED (Female) | MOS4.28 | 5 | |
| Emotion-intensity Speech Synthesis | ESD & CNCED (Male) | MOS4.25 | 5 | |
| Emotional Speech Synthesis | 48-category emotional speech synthesis zero-shot | ECS87 | 5 | |
| Emotional Speech Synthesis | 27 fine-grained emotion tasks (Female) | MOS4.12 | 5 | |
| Emotional Speech Synthesis | 27 fine-grained emotion tasks (Male) | MOS4.05 | 5 |