NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation
About
Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of consecutive translation. However, the excessive pursuit of low latency often results in fragmented chunk-wise speech. Consequently, listeners are subjected to an unnatural acoustic flow punctuated by frequent pauses, which could increase their cognitive load. To bridge this gap, we introduce a fluency-aware optimization framework designed to discover the sweet spot between the low-latency benefits of simultaneous translation and the natural flow of consecutive translation. Our framework minimizes inter-chunk silences by leveraging model-internal signals, including linguistic diversity and induced temporal variability in speech durations. Experiments on short- and long-form benchmarks show that our framework produces natural speech flow while maintaining competitive latency and translation quality.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Streaming Speech-to-Speech Translation | mTEDx Long-form FR-EN | SR Error21 | 5 | |
| Streaming Speech-to-Speech Translation | VoxPopuli Short-form FR-EN | Speech Rate (SR)0.1 | 5 | |
| Streaming Speech-to-Speech Translation | Audio-NTREX Long-form FR-EN | Speech Rate Error13 | 5 | |
| Streaming Speech-to-Speech Translation | CVSS-C Short-form FR-EN | SR Error0.08 | 5 | |
| Speech-to-Speech Translation Naturalness Evaluation | CVSS-C and mTEDx Human Evaluation Subset (test) | Preferred Rate68 | 2 |