Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data

About

Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence boundary artifacts. Existing approaches either compress sequences, increase context length or naively concatenate independently synthesized chunks. We present an inference-time approach called MagpieTTS-LF that enables MagpieTTS to produce coherent long-form speech without model retraining. Our method introduces three key innovations: (1) soft attention priors to guide monotonic alignment while preserving past and future context; (2) a stateful inference algorithm that maintains context across sentence chunks, ensuring prosodic continuity; (3) history-aware text encoding that uses past text for discourse-level prosodic planning. Experiments on long texts show significant improvements in long-range intelligibility, prosodic coherence, speaker consistency, and boundary naturalness compared to other baselines.

Subhankar Ghosh, Jason Li, Paarth Neekhara, Shehzeen Hussain, Ryan Langman, Xuesong Yang, Roy Fejgin• 2026

Related benchmarks

TaskDatasetResultRank
Long-form Speech GenerationHiFiTTS long-form 1-hour
WER2.5
4
Prosodic Boundary DiscontinuityLong-Form HiFiTTS 1-hour
Delta F0 (Hz)69.19
4
Showing 2 of 2 rows

Other info

Follow for update