AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials

About

Graphical User Interface (GUI) agents can automate complex tasks across digital environments, but their development is hindered by the scarcity of high-quality trajectory data for training. Existing approaches rely on expensive human annotation, making them unsustainable at scale. We propose AgentTrek, a scalable data synthesis pipeline that generates web agent trajectories by leveraging publicly available tutorials. Our three-stage method: (1) automatically harvests and filters tutorial-like texts from the internet using a specialized classification model, (2) transforms these texts into structured task specifications with step-by-step instructions, and (3) employs a visual-language model (VLM) agent to execute these instructions in real environments, while a VLM-based evaluator verifies trajectory correctness. The synthesized trajectories encompass multiple modalities, including text-based HTML observations with function-calling API actions, and vision-based screenshot observations with pixel-level actions. This multimodal data, enriched with chain-of-thought reasoning, enables agents to achieve state-of-the-art performance on both textual web browsing benchmarks (e.g., WebArena) and visual web grounding and browsing benchmarks (e.g., ScreenSpot Web and Multimodal Mind2Web). Furthermore, our fully automated approach significantly reduces data collection costs, achieving a cost of just $0.55 per high-quality trajectory without human annotators. Our work demonstrates that guided replay using web tutorials is a practical and scalable strategy for training advanced GUI agents, paving the way for more capable and autonomous digital assistants.

Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, Tao Yu• 2024

Related benchmarks

Task	Dataset	Result
Web navigation	WebArena	Overall Success Rate22.4	138
Web navigation and task completion	WebArena (test)	Average Task Completion22.4	137
GUI Navigation	Multimodal-Mind2Web Cross-Website	Step Success Rate51.4	37
GUI Navigation	Multimodal-Mind2Web Cross-Task	Step Success Rate55.7	32
GUI Navigation	Multimodal-Mind2Web Cross-Domain	Step Success Rate52.6	32
Web navigation	Multimodal-Mind2Web Cross-Task	Element Accuracy60.8	15
Web navigation	Multimodal-Mind2Web Cross-Website	Element Accuracy57.6	15
Web navigation	Multimodal-Mind2Web Cross-Domain	Element Accuracy56	15
Web navigation	MiniWob++	Accuracy45.28	15
Web navigation	Multimodal-Mind2Web Average	Avg. Step Success Rate53.2	14

Showing 10 of 12 rows

Other info

Follow for update

@wizwand_team Discord