Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation

About

Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech rather than long-form, continuous input. Prior approaches are difficult to reproduce and make assumptions that do not hold for end-to-end systems. We present a practical evaluation method for long-form SimulS2ST. Given source speech, pre-segmented source transcripts, and reference translations, we run automatic speech recognition (ASR) and forced alignment on the generated target speech to recover token-level timestamps, then apply a sentence-embedding-based aligner to match the target text to its corresponding source sentences. This enables sentence-level computation of latency and quality metrics, including YAAL and xCOMET, which are then aggregated into final system-level scores. Experiments on representative SimulS2ST systems show that the method is effective in practice and reveal that current systems suffer from substantial latency accumulation on long speech.

Yulin Xue, Siqi Ouyang, Lei Li• 2026

Related benchmarks

TaskDatasetResultRank
Simultaneous Speech-to-Speech TranslationAudio-NTREX-L Fr→En (test)
Latency (s)3.271
3
Simultaneous Speech-to-Speech TranslationAudio-NTREX-L De→En (test)
Latency (s)3.313
3
Simultaneous Speech-to-Speech TranslationAudio-NTREX-L Pt→En (test)
Latency (s)3.312
3
Simultaneous Speech-to-Speech TranslationAudio-NTREX-L Es→En (test)
Latency (s)3.657
3
Simultaneous Speech-to-Speech TranslationACL 60/60 En→De (dev)--
2
Simultaneous Speech-to-Speech TranslationACL 60/60 En→Ja (dev)--
2
Simultaneous Speech-to-Speech TranslationACL 60/60 En→Zh (dev)--
2
Showing 7 of 7 rows

Other info

Follow for update