Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

NAVER LABS Europe Submission to the Instruction-following Track

About

In this paper we describe NAVER LABS Europe submission to the instruction-following speech processing short track at IWSLT 2025. We participate in the constrained settings, developing systems that can simultaneously perform ASR, ST, and SQA tasks from English speech input into the following target languages: Chinese, Italian, and German. Our solution leverages two pretrained modules: (1) a speech-to-LLM embedding projector trained using representations from the SeamlessM4T-v2-large speech encoder; and (2) LoRA adapters trained on text data on top of a Llama-3.1-8B-Instruct. These modules are jointly loaded and further instruction-tuned for 1K steps on multilingual and multimodal data to form our final system submitted for evaluation.

Beomseok Lee, Marcely Zanon Boito, Laurent Besacier, Ioan Calapodescu• 2025

Related benchmarks

TaskDatasetResultRank
Speech Translation / Machine TranslationCoVoST2, EuroParlST Weighted Average (test)
Score (en-de)77.3
9
Speech Question AnsweringLibriSQA Part I 1.0
Accuracy82
7
Speech Question AnsweringLibriSQA Part II 1.0
Accuracy63
7
Automatic Speech RecognitionLibriSpeech (clean other), EuroParlST, CoVoST2 (Weighted Average) (test)
WER7.3
5
Spoken Question AnsweringLibriSQA Part I and Part II
MCIF41.7
4
Automatic Speech RecognitionLibriSpeech clean other EuroParlST CoVoST2
MCIF12.6
4
Showing 6 of 6 rows

Other info

Follow for update