Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents

About

Deep research agents aim to solve complex knowledge-intensive tasks through long-horizon planning, evidence gathering, reasoning, and report generation. While recent progress in search agents has demonstrated strong capabilities in information retrieval and answer verification, most existing training datasets remain search-centric, focusing primarily on closed-ended question answering and information localization. As a result, they mainly train information-seeking behavior while providing limited coverage of key deep research capabilities, including evidence integration, knowledge synthesis, planning, file understanding, and structured report generation. In this work, we propose a unified trajectory construction paradigm for deep research agents that combines closed-ended QA and open-ended exploration. The proposed framework consists of graph-grounded task formulation, agentic trajectory rollout, and multi-dimensional trajectory verification, enabling scalable synthesis of high-quality agentic trajectories spanning long-chain complex reasoning, deep research instruction following, report writing, file understanding and generation, and skills usage. Compared with existing search-oriented datasets, our synthesized trajectories place greater emphasis on knowledge synthesis, complex reasoning, and planning. S1-DeepResearch-32B achieves state-of-the-art performance among open-source models of comparable scale across 20 benchmarks spanning five capability dimensions, including complex reasoning, instruction following, report generation, file understanding, and skills usage. On several challenging deep research benchmarks, it approaches the performance of leading proprietary frontier models. These results highlight the importance of jointly modeling information acquisition, knowledge synthesis, and planning-oriented agent behaviors for building effective deep research agents.

Yao Dong, Xinglin Xiao, Liwei Dong, Xinlong Jin, Zhengbo Li, Heng Zhang, Duyun Wang, Nan Xu• 2026

Related benchmarks

TaskDatasetResultRank
Instruction FollowingComplexBench
Overall Score54.2
55
Textual long-horizon complex reasoningBrowsecomp
Score36.7
18
Textual long-horizon complex reasoningBrowseComp-ZH
Score48.4
17
Multimodal long-horizon complex reasoningBrowseComp-VL
Score39.1
16
Long-form Deep Research Report GenerationLong-form Deep Research Report Generation Benchmarks
DeepResearchBench II Score41.7
16
Multimodal long-horizon complex reasoningMM Search
Score54.4
16
Textual long-horizon complex reasoningGAIA text
Overall Score72.8
15
Textual long-horizon complex reasoningxbench DeepSearch
Score79.3
15
Multimodal long-horizon complex reasoningLiveVQA
Score67.7
15
Textual long-horizon complex reasoningHLE text
Score30.3
15
Showing 10 of 18 rows

Other info

Follow for update