Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation

About

Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of consecutive translation. However, the excessive pursuit of low latency often results in fragmented chunk-wise speech. Consequently, listeners are subjected to an unnatural acoustic flow punctuated by frequent pauses, which could increase their cognitive load. To bridge this gap, we introduce a fluency-aware optimization framework designed to discover the sweet spot between the low-latency benefits of simultaneous translation and the natural flow of consecutive translation. Our framework minimizes inter-chunk silences by leveraging model-internal signals, including linguistic diversity and induced temporal variability in speech durations. Experiments on short- and long-form benchmarks show that our framework produces natural speech flow while maintaining competitive latency and translation quality.

Dongwook Lee, Youngho Cho, Sangkwon Park, Heeseung Kim, Sungroh Yoon• 2026

Related benchmarks

TaskDatasetResultRank
Streaming Speech-to-Speech TranslationmTEDx Long-form FR-EN
SR Error21
5
Streaming Speech-to-Speech TranslationVoxPopuli Short-form FR-EN
Speech Rate (SR)0.1
5
Streaming Speech-to-Speech TranslationAudio-NTREX Long-form FR-EN
Speech Rate Error13
5
Streaming Speech-to-Speech TranslationCVSS-C Short-form FR-EN
SR Error0.08
5
Speech-to-Speech Translation Naturalness EvaluationCVSS-C and mTEDx Human Evaluation Subset (test)
Preferred Rate68
2
Showing 5 of 5 rows

Other info

Follow for update