Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning

About

Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate that while decomposition is essential for stability, extreme decomposition creates a "no-recovery bottleneck". We show that this bottleneck becomes critical due to highly non-uniform error distribution, where consistent errors on a few "hard" steps become irreversible. To address this, we propose Lookahead-Enhanced Atomic Decomposition (LEAD). By incorporating short-horizon future validation and aggregating overlapping rollouts, LEAD provides enough isolation to maintain stability while retaining enough local context to correct errors. This enables the o4-mini model to solve Checkers Jumping up to complexity $n=13$, whereas extreme decomposition fails beyond $n=11$.

Denys Pushkin, Emmanuel Abbe• 2026

Related benchmarks

TaskDatasetResultRank
Checkers JumpingCheckers Jumping n=13
Performance100
3
Checkers JumpingCheckers Jumping n=15
Performance87
3
Checkers JumpingCheckers Jumping n=16
Performance100
3
Long-horizon ReasoningCheckers Jumping
Performance (n=11)100
3
Long-horizon ReasoningCheckers Jumping
Performance (n=13)100
3
Checkers JumpingCheckers Jumping (n=14)
Performance100
3
Showing 6 of 6 rows

Other info

Follow for update