Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation

About

As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows. Current methods for effectively diagnosing agent failures load the full trajectory into an LLM's context window, which suffers from attention dilution and fails when agentic traces inevitably exceed context limits. To address this, we introduce SAFARI (Scaling long-horizon Agentic Fault AttRibution via active Investigation), a framework that replaces linear context loading with a tool-augmented diagnostic loop. By equipping LLMs with a specialized toolbox to read and search trajectory segments alongside a persistent Short-Term Memory (STM) for cross-turn reasoning, SAFARI effectively decouples diagnostic accuracy from architectural context limits. Our experiments demonstrate that SAFARI outperforms state-of-the-art results by 20% on the Who&When dataset within a 1M token budget, and by 19% on TRAIL GAIA subset on a 25K token budget. Most significantly, SAFARI maintains a 0.58 precision even when the target fault resides 5x beyond the model's native context window, a scenario where traditional evaluators fail entirely.

Chenyang Zhu, Jiayu Yao, Kushal Chawla, Youbing Yin, Nathan Wolfe, Pengshan Cai, Jingyu Wu, Spencer Hong, Sangwoo Cho, Shi-Xiong Zhang, Daben Liu, Sambit Sahu, Erin Babinsky• 2026

Related benchmarks

TaskDatasetResultRank
Fault attributionTRAIL (GAIA) 117 traces
Precision63.25
16
Fault attributionWho&When Algo (126 traces)
Accuracy48.4
16
Fault attributionWho&When Hand (58 traces)
Top Accuracy27.6
16
Fault attributionTRAIL SWEBench 31 traces
Precision41.94
16
Showing 4 of 4 rows

Other info

Follow for update