Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials

About

Graphical User Interface (GUI) agents can automate complex tasks across digital environments, but their development is hindered by the scarcity of high-quality trajectory data for training. Existing approaches rely on expensive human annotation, making them unsustainable at scale. We propose AgentTrek, a scalable data synthesis pipeline that generates web agent trajectories by leveraging publicly available tutorials. Our three-stage method: (1) automatically harvests and filters tutorial-like texts from the internet using a specialized classification model, (2) transforms these texts into structured task specifications with step-by-step instructions, and (3) employs a visual-language model (VLM) agent to execute these instructions in real environments, while a VLM-based evaluator verifies trajectory correctness. The synthesized trajectories encompass multiple modalities, including text-based HTML observations with function-calling API actions, and vision-based screenshot observations with pixel-level actions. This multimodal data, enriched with chain-of-thought reasoning, enables agents to achieve state-of-the-art performance on both textual web browsing benchmarks (e.g., WebArena) and visual web grounding and browsing benchmarks (e.g., ScreenSpot Web and Multimodal Mind2Web). Furthermore, our fully automated approach significantly reduces data collection costs, achieving a cost of just $0.55 per high-quality trajectory without human annotators. Our work demonstrates that guided replay using web tutorials is a practical and scalable strategy for training advanced GUI agents, paving the way for more capable and autonomous digital assistants.

Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, Tao Yu• 2024

Related benchmarks

TaskDatasetResultRank
Web navigation and task completionWebArena (test)
Average Task Completion22.4
42
GUI NavigationMultimodal-Mind2Web Cross-Website
Step Success Rate51.4
32
GUI NavigationMultimodal-Mind2Web Cross-Task
Step Success Rate55.7
27
GUI NavigationMultimodal-Mind2Web Cross-Domain
Step Success Rate52.6
27
Web navigationMiniWob++
Accuracy45.28
15
Web navigationMultimodal-Mind2Web Average
Avg. Step Success Rate53.2
14
GUI Trajectory Dataset ComparisonGUI Trajectory Datasets
Website Pages Count127
5
Showing 7 of 7 rows

Other info

Follow for update