MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
About
LLM-based multi-agent systems (MAS) have demonstrated significant potential in enhancing single LLMs to address complex and diverse tasks in practical applications. Despite considerable advancements, the field lacks a unified codebase that consolidates existing methods, resulting in redundant re-implementation efforts, unfair comparisons, and high entry barriers for researchers. To address these challenges, we introduce MASLab, a unified, comprehensive, and research-friendly codebase for LLM-based MAS. (1) MASLab integrates over 20 established methods across multiple domains, each rigorously validated by comparing step-by-step outputs with its official implementation. (2) MASLab provides a unified environment with various benchmarks for fair comparisons among methods, ensuring consistent inputs and standardized evaluation protocols. (3) MASLab implements methods within a shared streamlined structure, lowering the barriers for understanding and extension. Building on MASLab, we conduct extensive experiments covering 10+ benchmarks and 8 models, offering researchers a clear and comprehensive view of the current landscape of MAS methods. MASLab will continue to evolve, tracking the latest developments in the field, and invite contributions from the broader open-source community.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Mathematical Reasoning | AQUA-RAT | -- | 183 | |
| Coding | MBPP | -- | 175 | |
| Mathematics | MATH | -- | 172 | |
| Knowledge | MMLU-Pro | -- | 98 | |
| Mathematics | AIME 2024 | -- | 87 | |
| Coding | HumanEval | -- | 84 | |
| Medical | MedMCQA | -- | 81 | |
| Mathematical Reasoning | GSM-Hard | -- | 70 | |
| Science | GPQA | -- | 49 | |
| Multi-domain evaluation | Diverse Domain Collection (MATH, GSM-H, AQUA, AIME, SciBe, GPQA, MMLUP, MedMC, HEval, MBPP) | -- | 24 |