Breaking the Language Barrier: Can Direct Inference Outperform Pre-Translation in Multilingual LLM Applications?
About
Large language models hold significant promise in multilingual applications. However, inherent biases stemming from predominantly English-centric pre-training have led to the widespread practice of pre-translation, i.e., translating non-English inputs to English before inference, leading to complexity and information loss. This study re-evaluates the need for pre-translation in the context of PaLM2 models (Anil et al., 2023), which have been established as highly performant in multilingual tasks. We offer a comprehensive investigation across 108 languages and 6 diverse benchmarks, including open-end generative tasks, which were excluded from previous similar studies. Our findings challenge the pre-translation paradigm established in prior research, highlighting the advantages of direct inference in PaLM2. Specifically, PaLM2-L consistently outperforms pre-translation in 94 out of 108 languages. These findings pave the way for more efficient and effective multilingual applications, alleviating the limitations associated with pre-translation and unlocking linguistic authenticity.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Mathematical Reasoning | MGSM | Accuracy86.36 | 245 | |
| Commonsense MCQ | mCSQA | Accuracy72.08 | 9 | |
| Knowledge MCQ | Belebele | Accuracy71.43 | 9 | |
| Open-ended cultural | Aya | chrF15.96 | 9 | |
| Open-ended cultural | BLEND | chrF22.1 | 9 | |
| Knowledge MCQ | Global MMLU | Accuracy61.9 | 9 | |
| Math | MMATH | Accuracy (MMath)36 | 9 | |
| Open-ended cultural | Global-PIQA OE | chrF15.41 | 9 | |
| Open-ended factual | MKQA | chrF23.03 | 9 | |
| Commonsense MCQ | Global-PIQA | Accuracy61.95 | 9 |