Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

About

Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to trajectory neglect, in which agents lose focus on the task goal and interaction history at intermediate steps. Prior work has explored step-level supervision using Shannon-entropy-based uncertainty signals, which conflate inherent state complexity with agent confidence and therefore provide unreliable estimates of decision reliability. To address this issue, we propose normalized entropy, which measures confidence deviations relative to an agent's average behavior under a given state, thereby strengthening the association between low-quality actions and trajectory neglect. Building on this insight, we introduce Selective Trajectory-Aware Policy Optimization (STAPO), a hierarchical group-based RL framework. STAPO leverages normalized entropy to locate outlier steps associated with trajectory neglect and optimizes them via a joint mechanism of trajectory-aware reward and trajectory-independent penalty, enhancing trajectory awareness while preserving training stability. Extensive experiments on ALFWorld, WebShop, and Search-Augmented QA demonstrate that STAPO achieves state-of-the-art performance while substantially alleviating trajectory neglect, validating its effectiveness and robustness for agentic tasks.

Qiuyi Qi, Tian Liang, Mutian Bao, Jinjian Zhang, Dongnan Liu, Wei Zhou, Linjian Mo, Ming Kong, Jie Liu, Feng Zhang, Qiang Zhu• 2026

Related benchmarks

TaskDatasetResultRank
Multi-hop Question AnsweringHotpotQA (test)--
334
Web Navigation and ShoppingWebshop
Score89.1
248
Multi-hop Question AnsweringMuSiQue (test)--
151
Single-hop Question AnsweringTriviaQA (test)
Accuracy66
60
Single-hop Question AnsweringPopQA (test)
Accuracy48.7
43
Interactive Agent TaskAlfWorld
Pick Success Rate97.8
36
Multi-hop Question Answering2Wiki (test)
Accuracy45
24
Multi-hop Question AnsweringBamboogle (test)
Accuracy69.4
24
Single-hop Question AnsweringNQ (test)
Accuracy48.8
22
VLM Agent Success Rate EstimationSokoban 6x6
Success Rate80.5
4
Showing 10 of 11 rows

Other info

Follow for update