Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025

About

This paper describes Charles University submission to the Simultaneous Speech Translation Task of the IWSLT 2025. We cover all four language pairs with a direct or cascade approach. The backbone of our systems is the offline Whisper speech model, which we use for both translation and transcription in simultaneous mode with the state-of-the-art simultaneous policy AlignAtt. We further improve the performance by prompting to inject in-domain terminology, and we accommodate context. Our cascaded systems further use EuroLLM for unbounded simultaneous translation. Compared to the Organizers' baseline, our systems improve by 2 BLEU points on Czech to English and 13-22 BLEU points on English to German, Chinese and Japanese on the development sets. Additionally, we also propose a new enhanced measure of speech recognition latency.

Dominik Mach\'a\v{c}ek, Peter Pol\'ak• 2025

Related benchmarks

TaskDatasetResultRank
Speech-to-text TranslationIWSLT25Instruct en-de
BLEU35.25
10
Speech-to-text TranslationIWSLT25Instruct en-ja
BLEU33.44
6
Showing 2 of 2 rows

Other info

Follow for update