WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

About

Web agents such as Deep Research have demonstrated superhuman cognitive abilities, capable of solving highly challenging information-seeking problems. However, most research remains primarily text-centric, overlooking visual information in the real world. This makes multimodal Deep Research highly challenging, as such agents require much stronger reasoning abilities in perception, logic, knowledge, and the use of more sophisticated tools compared to text-based agents. To address this limitation, we introduce WebWatcher, a multi-modal Agent for Deep Research equipped with enhanced visual-language reasoning capabilities. It leverages high-quality synthetic multimodal trajectories for efficient cold start training, utilizes various tools for deep reasoning, and further enhances generalization through reinforcement learning. To better evaluate the capabilities of multimodal agents, we propose BrowseComp-VL, a benchmark with BrowseComp-style that requires complex information retrieval involving both visual and textual information. Experimental results show that WebWatcher significantly outperforms proprietary baseline, RAG workflow and open-source agents in four challenging VQA benchmarks, which paves the way for solving complex multimodal information-seeking tasks.

Xinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang, Qiuchen Wang, Ruixue Ding, Chenxi Wang, Jialong Wu, Yida Zhao, Kuan Li, Yong Jiang, Pengjun Xie, Fei Huang, Jingren Zhou• 2025

Related benchmarks

Task	Dataset	Result
Visual Question Answering	SimpleVQA	Accuracy0.59	225
Visual Question Answering	LiveVQA	Accuracy58.7	151
Multimodal Search	MMSearch	Accuracy55.3	119
Multimodal Search-based Question Answering	MMSearch	Accuracy55.3	59
Visual Question Answering	FVQA	Accuracy64.3	57
General Visual Question Answering	SimpleVQA	Pass@159	40
Visual Question Answering	SimpleVQA	Score59	38
Multimodal Deep Search	BC-VL	Accuracy26.7	37
Multimodal Search	BrowseComp-VL	Score27	35
Multimodal Search	BrowseComp-VL	Accuracy (BrowseComp-VL)27	34

Showing 10 of 43 rows

Other info

Follow for update

@wizwand_team Discord