Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

About

Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agent's current decision. To bridge this gap, we propose Agent-Authored World Modeling (AAWM), a training procedure that constructs supervision from the policy's own decision needs. Specifically, at each state, the agent identifies what it needs to understand about the environment before acting. These needs drive the retrieval of relevant transition evidence across trajectories, which is then synthesized into training targets that capture decision-oriented dynamics instead of reconstructing the next observation. This aligns the training objective with the dynamics the policy needs before acting, not with the contents of the next observation. Experimental results validate the effectiveness of AAWM across multiple environments and training settings. These results show that decision-aware world-model targets provide a more effective learning signal than next-observation prediction.

Guangfeng Cai, Kaibing Yang, Shuo He, Yu Li, Shengtian Yang, Jiaqi Lv, Lei Feng• 2026

Related benchmarks

TaskDatasetResultRank
Interactive Decision-makingAlfWorld
Overall Success Rate90.1
398
Online ShoppingWebshop
Score84.2
115
Language Agent TaskTextCraft
Success Rate (SR)35.7
17
Language Agent TaskAgentGym SciWorld (test)
Success Rate91.2
5
Language Agent TaskAgentGym WebShop (test)
Success Rate97.5
5
Language Agent TaskAgentGym ALFWorld (test)
Success Rate36
5
Language Agent TaskAgentGym All (test)
Success Rate69.3
5
Showing 7 of 7 rows

Other info

Follow for update