Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

About

We investigate the logical reasoning capabilities of large language models (LLMs) and their scalability in complex non-monotonic reasoning. To this end, we introduce ZebraLogic, a comprehensive evaluation framework for assessing LLM reasoning performance on logic grid puzzles derived from constraint satisfaction problems (CSPs). ZebraLogic enables the generation of puzzles with controllable and quantifiable complexity, facilitating a systematic study of the scaling limits of models such as Llama, o1 models, and DeepSeek-R1. By encompassing a broad range of search space complexities and diverse logical constraints, ZebraLogic provides a structured environment to evaluate reasoning under increasing difficulty. Our results reveal a significant decline in accuracy as problem complexity grows -- a phenomenon we term the curse of complexity. This limitation persists even with larger models and increased inference-time computation, suggesting inherent constraints in current LLM reasoning capabilities. Additionally, we explore strategies to enhance logical reasoning, including Best-of-N sampling, backtracking mechanisms, and self-verification prompts. Our findings offer critical insights into the scalability of LLM reasoning, highlight fundamental limitations, and outline potential directions for improvement.

Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal, Radha Poovendran, Peter Clark, Yejin Choi• 2025

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningAIME 2024 (test)--
294
Mathematical ReasoningAIME 2025 (test)--
191
Mathematical ReasoningOlympiad (test)--
27
Arithmetic ReasoningCountdown
Accuracy52.6
25
Mathematical ReasoningAIME24, AIME25, AMC23, MATH, Olympiad (aggregate)
Average Score44.8
18
Mathematical ReasoningAMC 2023 (test)
Pass@1 Rate59.1
18
Mathematical ReasoningAMC 23
Accuracy59.1
14
Mathematical ReasoningAIME 25
Accuracy7.3
14
Mathematical ReasoningMATH
Accuracy77.1
14
Mathematical ReasoningOlympiad
Accuracy (Olympiad)44.2
14
Showing 10 of 12 rows

Other info

Follow for update