Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MMATH: A Multilingual Benchmark for Mathematical Reasoning

About

The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex reasoning remain underexplored, with existing efforts largely focused on simpler tasks like MGSM. To address this gap, we introduce MMATH, a benchmark for multilingual complex reasoning spanning 374 high-quality math problems across 10 typologically diverse languages. Using MMATH, we observe that even advanced models like DeepSeek R1 exhibit substantial performance disparities across languages and suffer from a critical off-target issue-generating responses in unintended languages. To address this, we explore strategies including prompting and training, demonstrating that reasoning in English and answering in target languages can simultaneously enhance performance and preserve target-language consistency. Our findings offer new insights and practical strategies for advancing the multilingual reasoning capabilities of large language models. Our code and data could be found at https://github.com/RUCAIBox/MMATH.

Wenyang Luo, Wayne Xin Zhao, Jing Sha, Shijin Wang, Ji-Rong Wen• 2025

Related benchmarks

TaskDatasetResultRank
Multilingual Mathematical ReasoningPolyMath (test)
Accuracy (Ar)12.8
30
Multilingual Mathematical ReasoningMMATH Out-of-Domain Languages (test)
Vietnamese Accuracy22.4
22
Multilingual Mathematical ReasoningMMATH In-Domain Languages (test)
Accuracy (Ar)18.3
22
Multilingual Mathematical ReasoningMMATH All Languages (test)
Average Score (All)20.1
22
Math problem solvingMMATH (test)
Accuracy (Ar)21.2
20
Retrieval-Augmented GenerationHotpot ENKB5 (In-Domain Average)
Language Consistency81.43
15
Cross-lingual Retrieval-Augmented GenerationBioASQ-ENKB5
ID Score71.34
14
Retrieval-Augmented GenerationHotpot-ENKB5 Out-of-Domain Average
Language Consistency96.32
13
Retrieval-Augmented GenerationHotpot-ENKB5 v1 (Indonesian)
Language Consistency56
11
Cross-lingual Retrieval-Augmented GenerationHotpot-ENKB5 (all)
Language Consistency (ALL-AVG)86.39
11
Showing 10 of 14 rows

Other info

Follow for update