CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic

About

Tool-Integrated Reasoning (TIR) with search engines enables large language models to iteratively retrieve up-to-date external knowledge, enhancing adaptability and generalization in complex question-answering tasks. However, existing search agent pipelines typically depend on reinforcement learning based optimization, which often suffers from sparse outcome rewards, leading to inefficient exploration and unstable training. We introduce CriticSearch, a fine-grained credit-assignment framework that supplies dense, turn-level feedback via a retrospective critic mechanism. During training, a frozen, asymmetric critique LLM retrospectively evaluates each turn using privileged information from the full trajectory and gold answers, converting these assessments into stable, dense rewards that guide policy improvement. Experimental results across diverse multi-hop reasoning benchmarks demonstrate that CriticSearch consistently outperforms existing baselines, achieving faster convergence, improved training stability, and higher performance.

Yaocheng Zhang, Haohuan Huang, Zijun Song, Yuanheng Zhu, Qichao Zhang, Zijie Zhao, Dongbin Zhao• 2025

Related benchmarks

Task	Dataset	Result
Multi-hop Question Answering	HotpotQA (test)	--	334
Multi-hop Question Answering	2WikiMultiHopQA (test)	EM40.9	247
Multi-hop Question Answering	MuSiQue	--	209
Question Answering	HotpotQA	EM41.4	173
Multi-hop Question Answering	Bamboogle (test)	EM36.8	110
Question Answering	2WikiMultihopQA	EM40.9	107
Question Answering	MuSiQue	EM18	71
Question Answering	Bamboogle	EM Accuracy (%)36.8	68
Multi-hop Question Answering	HotpotQA	Exact Match (EM)44.2	66
Multi-hop Question Answering	Multi-Hop QA (HotpotQA, 2Wiki, Musique, Bamboogle) (test)	HotpotQA Score0.414	65

Showing 10 of 13 rows

Other info

Follow for update

@wizwand_team Discord