Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task

About

We re-implement the NAVER LABS IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task (constrained condition, short audio track), adapting it to the mandated components: SeamlessM4T-v2-large as the speech encoder and Qwen3-4B-Instruct as the LLM backbone. The three-stage approach projector alignment, text-only LoRA pre-training, and multimodal merging is preserved from the original design. We additionally construct 100k synthetic instruction-following examples across ten speech-centric task types (10k per task) from the provided corpora, suitable for further Stage 3 fine-tuning. Our primary model achieves COMET 0.781 on EN-ZH speech translation and BERTScore-F1 0.346 on English SQA on the MCIF benchmark.

Anand Kamble, Aniket Tathe• 2026

Related benchmarks

TaskDatasetResultRank
Speech TranslationMCIF short-form--
8
Speech Question AnsweringMCIF constrained, short audio
BERTScore-F1 (EN)34.6
4
Automatic Speech RecognitionMCIF constrained, short audio
WER23.49
4
Showing 3 of 3 rows

Other info

Follow for update