Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning

About

Multi-agent debate (MAD) can improve large language model reasoning, but fixed debate pipelines often waste computation and can amplify correlated errors among similar agents. We propose ARMOR-MAD, a training-free heterogeneous MAD framework that treats debate as conditional computation. ARMOR-MAD combines three components: Pre-debate Agreement Routing (PAR) decides whether independently generated Round-0 answers require debate; Early Agreement Stopping Evaluator (EASE) stops debate after convergence; and Semantic Outlier Detection (SOD) down-weights abnormal final answers during aggregation. Across MATH Level 5, GSM8K, MMLU, and MMLU-Pro, ARMOR-MAD consistently improves over fixed-round heterogeneous debate with the same model pool, reaching 65.5\%, 96.5\%, 90.0\%, and 81.5\% accuracy, respectively. The results suggest that genuine model heterogeneity and agreement-based control are both important for making MAD more accurate and efficient.

Fuqiang Niu, Bowen Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Multiple-choice Question AnsweringMMLU
Accuracy90
222
Mathematical ReasoningMATH L5
Accuracy0.655
162
Multiple-choice Question AnsweringMMLU-Pro
Accuracy (MMLU-Pro)81.5
14
Showing 3 of 3 rows

Other info

Follow for update