A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026
About
We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy AlignAtt, and submit it to IWSLT 2026 Simultaneous Speech Translation Shared task for Czech to English and English to German and Italian. The strengths of our system are: (1) high translation quality, outperforming similarly sized baselines both in low- and high-latency regimes in computationally unaware simulations; (2) low computational requirements, as the model has only 1B parameters; (3) multilinguality -- support of 25 source and 25 target languages.
Aziz Sharipov Ortega, Dominik Mach\'a\v{c}ek• 2026
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Simultaneous Speech Translation | MCIF En-De (dev) | BLEU31.73 | 7 | |
| Simultaneous Speech Translation | MCIF En-It (dev) | BLEU43.56 | 7 | |
| Simultaneous Speech Translation | IWSLT Cs-En 2026 (dev) | BLEU32.01 | 4 |
Showing 3 of 3 rows