Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval

About

With the rapid expansion of massive multilingual corpora, Multilingual Information Retrieval (MLIR) has emerged as a critical technology for global information access. MLIR enables users to retrieve semantically relevant documents from multilingual text collections using a single-language query. However, recent multilingual dense retrieval models often exhibit a strong preference for documents in the same language as the query. This leads to severe language bias, where top-ranked results are dominated by documents of specific languages, even when documents in other languages contain more semantically relevant information. To address this issue, we propose SHIFT, a training-free method applicable in the indexing stage. Specifically, SHIFT utilizes parallel translation pairs to estimate a relative language vector for each target language with respect to a source language. Subsequently, SHIFT corrects the language-specific offset by subtracting this relative language vector from document embeddings during indexing. Our comprehensive evaluation across four MLIR benchmarks and diverse dense retrieval models confirms that SHIFT can effectively mitigate language bias and enhance MLIR performance.

Youngjoon Jang, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim• 2026

Related benchmarks

TaskDatasetResultRank
Multilingual Information RetrievalXQuAD--
80
Multilingual Information RetrievalBelebele
nDCG@200.926
45
Multilingual Information RetrievalMLQA
nDCG@200.671
12
Multilingual Information RetrievalMultiEup v2
nDCG@2051.9
12
Multilingual RetrievalNeuCLIR RetrievalHardNegatives Chinese and Russian 2023
nDCG@2052
12
Showing 5 of 5 rows

Other info

Follow for update