NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task
About
We re-implement the NAVER LABS IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task (constrained condition, short audio track), adapting it to the mandated components: SeamlessM4T-v2-large as the speech encoder and Qwen3-4B-Instruct as the LLM backbone. The three-stage approach projector alignment, text-only LoRA pre-training, and multimodal merging is preserved from the original design. We additionally construct 100k synthetic instruction-following examples across ten speech-centric task types (10k per task) from the provided corpora, suitable for further Stage 3 fine-tuning. Our primary model achieves COMET 0.781 on EN-ZH speech translation and BERTScore-F1 0.346 on English SQA on the MCIF benchmark.
Anand Kamble, Aniket Tathe• 2026
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Speech Translation | MCIF short-form | -- | 8 | |
| Speech Question Answering | MCIF constrained, short audio | BERTScore-F1 (EN)34.6 | 4 | |
| Automatic Speech Recognition | MCIF constrained, short audio | WER23.49 | 4 |
Showing 3 of 3 rows