Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
About
Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Simultaneous Speech-to-Speech Translation | Audio-NTREX-L Fr→En (test) | Latency (s)5.892 | 3 | |
| Simultaneous Speech-to-Speech Translation | Audio-NTREX-L De→En (test) | Latency (s)5.933 | 3 | |
| Simultaneous Speech-to-Speech Translation | Audio-NTREX-L Pt→En (test) | Latency (s)5.53 | 3 | |
| Simultaneous Speech-to-Speech Translation | Audio-NTREX-L Es→En (test) | Latency (s)5.592 | 3 | |
| Simultaneous Speech-to-Speech Translation | ACL 60/60 En→De (dev) | Latency (s)7.939 | 2 | |
| Simultaneous Speech-to-Speech Translation | ACL 60/60 En→Ja (dev) | Latency (s)9.413 | 2 | |
| Simultaneous Speech-to-Speech Translation | ACL 60/60 En→Zh (dev) | Latency (second)5.306 | 2 |